Information processing system and information processing method

WO2026204072A1PCT designated stage Publication Date: 2026-10-01SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/007079
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2026-02-26
Publication Date
2026-10-01

Smart Images

  • Figure JP2026007079_01102026_PF_FP_ABST
    Figure JP2026007079_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an information processing system and an information processing method that make it possible to reflect a user's appearance in a virtual space. On the basis of a plurality of images captured by a plurality of cameras in a head-mounted display (HMD) worn by a user and a plurality of cameras of a plurality of trackers worn by the user, a three-dimensional model of the user is constructed, and the three-dimensional model is used to generate an avatar of the user. The present invention can be applied to the presentation of virtual space images using an HMD.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system and information processing method

[0001] The present disclosure relates to an information processing system and an information processing method, and particularly relates to an information processing system and an information processing method capable of reflecting a user's appearance in a virtual space.

[0002] So-called Mixed Reality (MR) technology is generally becoming widespread, which estimates in real time changes in the position and posture (hereinafter also referred to as pose) of an HMD (Head Mounted Display) worn on a user's head, captures the estimated pose changes as movement of the head, and reflects the changes in an image drawn on a display to present the image to the user.

[0003] With this MR technology, a user can view and listen to an image in a virtual space presented on an HMD, thereby having a feeling as if the user themself has entered the drawn virtual space.

[0004] Note that the MR technology referred to herein also includes Virtual Reality (VR) technology, Augmented Reality (AR) technology, and the like.

[0005] Furthermore, by applying this MR technology, attaching trackers (or controllers) to the hands and feet to estimate the poses of the hands and feet, and reflecting the movements of the hands and feet in addition to the movement of the head (HMD) in the virtual space, it is possible to make the system recognize actions using the hands and feet (such as grasping an object or kicking) in the virtual space, and to implement an intuitive user interface.

[0006] To implement the above-described MR technology, pose estimation of the head (HMD) and the hands and feet (trackers) is essential. Conventionally, the Outside-In method, which performs pose estimation by capturing the HMD and trackers with a camera installed on the environment side, has been mainstream. However, the Outside-In method requires complicated work for installing and preparing the camera, and can only be implemented in a place where the camera is installed.

[0007] Therefore, an Inside-Out method has been proposed, which involves equipping the HMD with a camera to estimate the poses of the head (HMD), hands, and feet (trackers) from images captured by the HMD's camera.

[0008] In this Inside-Out method, for example, tracker pose estimation is achieved using images of the tracker captured by the HMD's camera. Therefore, in the Inside-Out method, there is no need to install a camera in the environment, eliminating the need for cumbersome camera installation work, and making it usable in any location. (See Patent Document 1).

[0009] More recently, a method has been proposed in which cameras are mounted on both the HMD and the tracker, and the poses of both the HMD and the tracker are estimated by mutually utilizing images captured by both cameras. In this method, not only can the user's pose be reflected in the rendering using only the device worn by the user (HMD, tracker, etc.), but the pose of the tracker (hands and feet) can also be estimated even in positions that are not visible to the HMD's camera.

[0010] International Publication No. 2020 / 110659

[0011] Incidentally, a technology called telepresence has been proposed that recreates and presents a person in a remote location in a virtual space as if they were actually there, allowing for visual recognition.

[0012] By applying the aforementioned technology to this telepresence, it is conceivable that the pose of the HMD or tracker could be estimated from images captured by the camera, and the user's movements could be reflected and displayed in the composition of images in the virtual space as seen from the user's perspective, as well as in the avatar that represents the user as seen from the perspective of other users in the virtual space.

[0013] However, since telepresence is a technology that allows users to visually perceive people in remote locations as if they were present, it is sometimes important that the rendered avatar reflects not only the user's movements but also the user's appearance.

[0014] However, avatars are often computer graphics (CG) characters that do not generally reflect the user's appearance, and even with the application of the aforementioned technologies, it is generally not possible to reflect the user's appearance, which is important in telepresence.

[0015] This disclosure is made in light of these circumstances, and in particular, reflects the user's appearance within the virtual space.

[0016] One aspect of the information processing system of this disclosure is an information processing system comprising: an HMD (Head Mounted Display) having a plurality of cameras worn by a user; a plurality of trackers having a plurality of cameras worn by the user; a 3D model configuration unit that constructs a 3D model of the user based on a plurality of images captured by the plurality of cameras of the HMD and the plurality of cameras of the plurality of trackers; and an avatar generation unit that generates an avatar of the user using the 3D model.

[0017] One aspect of the information processing method of this disclosure is an information processing method that includes: a three-dimensional model configuration process that constitutes a three-dimensional model of the user based on a plurality of images captured by a plurality of cameras of an HMD (Head Mounted Display) having a plurality of cameras worn by the user and a plurality of cameras of a plurality of trackers having a plurality of cameras worn by the user; and an avatar generation process that generates an avatar of the user using the three-dimensional model.

[0018] In one aspect of this disclosure, a three-dimensional model of the user is constructed based on multiple images captured by multiple cameras of a Head-Mounted Display (HMD) worn by the user and multiple cameras of a tracker worn by the user, and an avatar of the user is generated using the three-dimensional model.

[0019] This is a diagram illustrating the overview of this disclosure. This is a diagram illustrating the information processing system of this disclosure. This is a diagram illustrating an example of the external configuration of the HMD and tracker in Figure 2. This is a diagram illustrating an example of the hardware configuration of the HMD. This is a diagram illustrating an example of the hardware configuration of the tracker. This is a diagram illustrating an example of the hardware configuration of the information processing device. This is a diagram illustrating the functions realized by the information processing system in Figure 2. This is a diagram illustrating a method for generating a 3D model of the user, who is the subject. This is a diagram illustrating an example of presenting a blind spot area. This is a diagram illustrating an example of reducing the blind spot area. This is a flowchart illustrating the display process of a virtual space image. This is a diagram illustrating an example of the configuration of a general-purpose computer.

[0020] Preferred embodiments of this disclosure will be described in detail below with reference to the attached drawings. In this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions will be omitted.

[0021] The following describes the configurations for implementing this technology. The explanation will proceed in the following order.

[0022] 1. Overview of this Disclosure 2. Preferred Embodiments 3. Description of a Computer to which this Technology is Applied

[0023] <<1. Overview of this Disclosure>> This disclosure, in particular, enables the user's appearance to be reflected in the virtual space. Therefore, first, an overview of this disclosure will be explained with reference to Figure 1.

[0024] The left side of Figure 1 shows an example configuration that outlines the information processing system of this disclosure. The information processing system 1 in Figure 1 consists of an HMD (Head Mounted Display) 21 equipped with a camera that is worn on the head of the user 11, and trackers 22-1 to 22-5 equipped with cameras that are worn on both hands, the torso, and both feet. Hereafter, when it is not necessary to distinguish between trackers 22-1 to 22-5, they will simply be referred to as tracker 22, and the other components will be referred to similarly.

[0025] More specifically, the HMD 21 in Figure 1 is equipped with cameras 21C-1 to 21C-4, each capturing images of the area around the HMD 21 (around the user's head 11) at different angles of view. Based on the images captured by cameras 21C-1 to 21C-4, the pose (position and orientation) of the HMD 21 (and the user's head wearing it) is estimated.

[0026] The tracker 22 is attached to both hands, the torso, and both feet of the user 11, and each is equipped with two cameras 22C that capture images with different fields of view.

[0027] Specifically, trackers 22-1 and 22-2 are attached to both hands of the user 11 and are equipped with cameras 22C-1-1, 22C-1-2, and cameras 22C-2-1, 22C-2-2, respectively, and capture images of the area around the user 11's hands. Based on the images captured by these cameras, the pose (position and orientation) of both hands is estimated.

[0028] The tracker 22-3 is attached to the torso of user 11 and is equipped with cameras 22C-3-1 and 22C-3-2, which capture images of the area around user 11's torso. Based on the images captured by cameras 22C-3-1 and 22C-3-2, the pose (position and orientation) of the torso is estimated.

[0029] Trackers 22-4 and 22-5 are attached to both feet of user 11, and are equipped with cameras 22C-4-1, 22C-4-2, and cameras 22C-5-1, 22C-5-2, respectively, and capture images of the area around user 11's feet. Based on the images captured by cameras 22C-4-1, 22C-4-2, and cameras 22C-5-1, 22C-5-2, the pose (position and orientation) of both feet is estimated.

[0030] The surface shape and texture of user 11 are determined by matching the images captured from the poses (position and orientation) of the HMD 21 (head) and trackers 22-1 to 22-5 (both arms, torso, and both legs) estimated in this way. Based on this surface shape and texture of user 11, a 3D model 11M is reconstructed that reflects the user 11's visual information (information about the clothing worn and body shape), as shown in the center of Figure 1.

[0031] In this way, the user's movements, obtained from the poses (positions and orientations) of the HMD 21 and trackers 22-1 to 22-4, are reflected in the reconstructed 3D model 11M within the virtual space, thereby generating an avatar that represents the user.

[0032] This makes it possible to display an avatar within a three-dimensional virtual space that reflects not only the user's movements but also their appearance.

[0033] In other words, for example, when generating a first-person perspective virtual space image from the viewpoint of user 11, as shown in the upper right of Figure 1, it becomes possible to generate an image V1 in the virtual space that represents an avatar (arm) VR 11 that reflects the movements and appearance of user 11, based on a 3D model 11M. This makes it possible to realize a UI (User Interface) image, as shown in image V1, using the avatar (arm) VR 11, in the HMD 21.

[0034] Furthermore, for example, when generating a third-person perspective virtual space image from the viewpoint of a user other than user 11, it becomes possible to generate an image V2 in the virtual space that represents an avatar VR 12 that reflects the movements and appearance of user 11, based on the 3D model 11M, as shown in the lower right of Figure 1. This makes it possible for other users' HMDs 21 to display an avatar that looks as if it were user 11 themselves, similar to telepresence.

[0035] As a result, based on the images captured by the HMD 21 and the cameras 21C of the trackers 22-1 to 22-5, it becomes possible to appropriately reflect the user's movements and appearance in the real world in the avatar, which is the user's digital counterpart in the virtual space image.

[0036] <<2. Preferred Embodiments>> <Information Processing System> Next, with reference to Figure 2, an example of the configuration of an information processing system that can reflect and represent the user's movements and appearance in the real world in an avatar, which is a representation of the user in a virtual space image, will be described.

[0037] The information processing system 31 in Figure 2 consists of HMDs (Head Mounted Displays) 51 worn by each of the users 41-1 to 41-n, trackers 52-1 to 52-5, and an information processing device 53, and is configured to exchange data with each other via a network 50, such as the Internet.

[0038] In the following, when it is not necessary to distinguish between trackers 52-1 to 52-5, they will simply be referred to as tracker 52, and the other components will be referred to similarly. Also, HMD 51 and tracker 52 correspond to HMD 21 and tracker 22 in Figure 1.

[0039] The HMD (Head Mounted Display) 51 is configured to be worn, for example, by wrapping around the user's head, and estimates the position and posture (hereinafter also referred to as pose) of the head and transmits it to the information processing device 53. Based on the transmitted pose information, the HMD 51 acquires and presents images in a virtual space that are visible to the user 41 wearing the HMD 51 according to the pose of their own head, which are generated by the information processing device 53.

[0040] Trackers 52-1 to 52-5 are configured to be attached to the user 41's hands, torso, and feet, for example, by wrapping around them, and estimate the poses of the hands, torso, and feet and transmit them to the information processing device 53.

[0041] Furthermore, although FIG. 2 shows an example in which a total of five trackers 52 are attached to both hands, the torso, and both feet of the user 41, the present invention is not limited thereto. The number may be smaller than this, or may be larger than this. For example, the attachment of the tracker 52 to the torso may be omitted. Further, for example, for a foot, a total of two trackers 52 may be attached: one to a portion below the knee (e.g., the ankle) and one to a portion above the knee (e.g., the thigh).

[0042] However, in order to appropriately reflect the movement of both hands and both feet, it is preferable that approximately five trackers 52 in total as referred to in FIG. 1 are attached.

[0043] The information processing device 53 generates a virtual space image including avatars that reflect the pose relationship of both hands, the torso, and both feet of each of the plurality of users 41-1 to 41-n, in accordance with the viewpoint of each of the users 41-1 to 41-n, based on the respective pose estimation results supplied from the HMDs 51 and the trackers 52-1 to 52-5 of the plurality of users 41-1 to 41-n, and supplies the generated virtual space image to each HMD 51.

[0044] That is, based on the respective pose estimation results supplied from the HMDs 51 and the trackers 52-1 to 52-5 of the plurality of users 41-1 to 41-n, the information processing device 53 generates, for example, a virtual space image for the user 41-1 such that the user 41-1 has a first-person viewpoint, and the other users 41-2 to 41-n have a third-person viewpoint. Further, for example, the information processing device 53 generates a virtual space image for the user 41-2 such that the user 41-2 has a first-person viewpoint, and the other users 41-1 and 41-3 to 41-n have a third-person viewpoint.

[0045] <Example of External Appearance Configuration of HMD and Tracker> Next, an example of the external appearance configuration of the HMD 51 and the tracker 52 will be described with reference to FIG. 3.

[0046] The left part of FIG. 3 shows an example of the external appearance configuration of the HMD 51, and the right part of FIG. 3 shows an example of the external appearance configuration of the tracker 52.

[0047] The HMD 51 is worn by wrapping it around the head, such as a band or belt. The left part of FIG. 2 shows an external view of the HMD 51 when worn on the face of user 41, with the face of user 41 facing the plane of the drawing. The HMD 51 is provided with cameras 71-1 to 71-4 at four corners respectively, which capture images of the external environment with different angles of view. The HMD 51 uses a plurality of images captured by the cameras 71-1 to 71-4 to estimate the pose (position and orientation) of the HMD 51 (the head of user 41) by technologies such as Visual SLAM (Simultaneous Localization And Mapping) and Visual Positioning System, for example.

[0048] In addition, apart from the camera 71 that captures images of the external environment, the HMD 51 may also be provided with an inward-facing camera that captures images of the user's face. This camera is used to recognize the user's line of sight direction and estimate facial expressions to be reflected on an avatar. It should be noted that if there is no need to reflect the line of sight direction and facial expressions in real time during operation, capturing an image of the face in advance using the camera of the tracker 52 when the user is not wearing the HMD 51 can be used as a substitute. Therefore, the inward-facing camera is not an essential component in the configuration of the present disclosure, and may be omitted if it is necessary to simplify the device configuration, reduce weight, or cut costs.

[0049] The HMD 51 is internally provided with a motion sensor 72 (dotted line frame) composed of an IMU (Inertial Measurement Unit) and the like, and estimates the pose of the HMD 51 (the head of user 41) based on the 6-axis sensing result consisting of 3-axis acceleration and 3-axis angular velocity detected by the motion sensor 72.

[0050] The HMD 51 improves the accuracy of the estimation result by, for example, combining the pose estimation result based on the plurality of images captured by the cameras 71-1 to 71-4 and the pose estimation result based on the 6-axis sensing result of the motion sensor 72 as needed, and transmits the estimated pose of the HMD 51 (the head of user 41) to the information processing device 53.

[0051] The MHD 51 has a display unit 73 (indicated by a dashed line frame) on its rear side when viewed from the plane of the paper in Figure 3, that is, facing the user 41's eyes. Based on the transmitted pose estimation result, it acquires and displays a virtual space image generated by the information processing device 53.

[0052] The tracker 52 is a band or belt-like device that is wrapped around both hands, the torso, and both feet. The tracker 52 is equipped with two cameras 81-1 and 81-2, each capturing images in different directions. The tracker 52 uses multiple images captured by cameras 81-1 and 81-2 to estimate the pose (position and posture) of the tracker 52 (the hands, torso, and feet of the user 41 to whom the tracker 52 is attached) using technologies such as Visual SLAM (Simultaneous Localization And Mapping) or Visual Positioning System.

[0053] The tracker 52 is equipped with a motion sensor 82 consisting of an IMU (Inertial Measurement Unit) and the like. Based on the 6-axis sensing results, which consist of 3-axis acceleration and 3-axis angular velocity detected by the motion sensor 82, the pose of the tracker 52 (both hands, torso, and both feet of the user 41 to whom the tracker 52 is attached) is estimated.

[0054] The tracker 52 improves the accuracy of the estimation result by combining, as necessary, the pose estimation result based on multiple images captured by cameras 81-1 to 81-2 and the pose estimation result based on the 6-axis sensing result of the motion sensor 82, and transmits the estimated pose of the tracker 52 (both hands, torso, and both feet of the user 41 to whom the tracker 52 is attached) to the information processing device 53.

[0055] In this example, we will describe a case in which the HMD 51 and tracker 52 perform both pose estimation using multiple images captured by cameras 71 and 81, and pose estimation using motion sensors 72 and 82, respectively.

[0056] However, the HMD 51 and tracker 52 will primarily utilize pose estimation results using multiple images captured by cameras 71 and 81, and pose estimation based on sensing results from motion sensors 72 and 82 will be used in combination with the image-based pose estimation results, or used supplementarily in situations where the images are unclear and cannot be used for pose estimation, in order to improve accuracy.

[0057] In other words, since the motion sensors 72 and 82 are intended to improve the accuracy of pose estimation using images, they are not essential components of the configuration of this disclosure and may be omitted if necessary to simplify the device configuration, reduce weight, or lower costs.

[0058] <Example of HMD Hardware Configuration> Next, with reference to Figure 4, an example of the hardware configuration of the HMD 51 in Figure 2 will be described.

[0059] The HMD 51 consists of a processing circuit 101, an input unit 102, an output unit 103, a storage unit 104, a communication unit 105, a drive 106, a removable media 107, a camera 71, and a motion sensor 72, which are connected to each other via a bus 108, enabling the transmission and reception of data and programs.

[0060] The processing circuit 101 consists of a processor and memory, and controls the overall operation of the HMD 51. The processing circuit 101 also includes an environment map creation unit 121, a pose estimation unit 122, and a display control unit 123.

[0061] Furthermore, the environmental map creation unit 121, the pose estimation unit 122, and the display control unit 123 will be described later with reference to Figure 7, along with a description of the functions realized by the information processing system 31 in Figure 2.

[0062] The input unit 102 consists of input devices such as a keyboard, mouse, and touch panel for inputting various types of information, and supplies the input information and corresponding signals to the processing circuit 101. The input unit 102 may also consist of a microphone that receives the user's speech content related to conversations with other avatars in the virtual space as voice input, and supplies the voice information related to the received voice input to the processing circuit 101.

[0063] The output unit 103 is controlled by the processing circuit 101 and includes, for example, a display unit 73 which displays various processing results and input content, and an audio output unit (not shown) which includes a speaker which generates conversations with other avatars in the virtual space, music, and various warning sounds.

[0064] The storage unit 104 consists of an HDD (Hard Disk Drive), an SSD (Solid State Drive), or semiconductor memory, and is controlled by the processing circuit 101 to write or read various data and programs.

[0065] The communication unit 105 is controlled by the processing circuit 101 and enables communication via wired or wireless means, such as LAN (Local Area Network) or Bluetooth (registered trademark), to send and receive various data and programs between the information processing device 53, other HMDs 51, and trackers 52 via the network 50 as needed.

[0066] The drive 106 reads and writes data to removable media 107 such as magnetic disks (including flexible disks), optical disks (including CD-ROMs (Compact Disc-Read Only Memory) and DVDs (Digital Versatile Discs)), magneto-optical disks (including MDs (Mini Discs)), or semiconductor memory.

[0067] The camera 71 consists of a CMOS (Complementary Metal Oxide Semiconductor) image sensor or the like, and captures an image within a predetermined field of view, supplying the capture result to the processing circuit 101, or storing it in the storage unit 104. In Figure 4, only one camera 71 is provided, but in this specification, as shown on the left side of Figure 3, in reality, four cameras 71-1 to 71-4, or any other number of cameras, are connected.

[0068] The motion sensor 72 consists of an IMU and other components, and measures information from a total of six axes, consisting of the acceleration in three axes and the angular velocity in three axes of the HMD 51 (the user's head), and outputs the estimation results to the processing circuit 101.

[0069] <Example of Tracker Hardware Configuration> Next, with reference to Figure 5, an example of the hardware configuration of the tracker 52 in Figure 2 will be described.

[0070] The tracker 52 consists of a processing circuit 151, an input unit 152, an output unit 153, a storage unit 154, a communication unit 155, a drive 156, a removable media 157, a camera 81, and a motion sensor 82, which are connected to each other via a bus 158, and can send and receive data and programs.

[0071] Furthermore, the processing circuit 151, input unit 152, output unit 153, storage unit 154, communication unit 155, drive 156, removable media 157, bus 158, camera 81, and motion sensor 82 have configurations corresponding to the processing circuit 101, input unit 102, output unit 103, storage unit 104, communication unit 105, drive 106, removable media 107, bus 108, camera 71, and motion sensor 72 in Figure 3, so their explanation is omitted.

[0072] However, while the HMD 51 is mounted on the head, the tracker 52 is wrapped around both hands, the torso, and both feet, and needs to be small enough not to hinder the user's movements. Therefore, although the electrical configurations of both hardware components are similar, the tracker 52 is smaller than the HMD 51. Also, in Figure 5, only one camera 81 is shown to be connected, but in this specification, as shown in the right side of Figure 3, in reality, two cameras 81-1 and 81-2, or any other number of cameras, are connected.

[0073] The processing circuit 151 consists of a processor and memory, and controls the overall operation of the tracker 52. The processing circuit 151 also includes an environment map creation unit 171 and a pose estimation unit 172.

[0074] Furthermore, the environmental map creation unit 171 and the pose estimation unit 172 have configurations corresponding to the environmental map creation unit 121 and the pose estimation unit 122, and will be described later in conjunction with the explanation of the functions realized by the information processing system 31 of Figure 2, which will be explained with reference to Figure 6.

[0075] <Example of Hardware Configuration of Information Processing Device> Next, with reference to Figure 6, an example of the hardware configuration of the information processing device 53 in Figure 2 will be described.

[0076] The information processing device 53 consists of a processing circuit 191, an input unit 192, an output unit 193, a storage unit 194, a communication unit 195, a drive 196, and a removable media 197, which are connected to each other via a bus 198, and can send and receive data and programs.

[0077] Furthermore, the processing circuit 191, input unit 192, output unit 193, storage unit 194, communication unit 195, drive 196, removable media 197, and bus 198 have configurations corresponding to the processing circuit 101, input unit 102, output unit 103, storage unit 104, communication unit 105, drive 106, removable media 107, and bus 108 in Figure 3, so their explanation is omitted.

[0078] The processing circuit 191 consists of a processor and memory and controls the overall operation of the tracker 52. The processing circuit 191 also includes an environmental map management unit 201, a 3D reconstruction unit 202, and a display image generation unit 203. Of these, the 3D reconstruction unit 202 further includes a matching processing unit 211 and a surface 3D shape estimation unit 212.

[0079] The memory unit 194 stores a shared environment map 231 that is generated based on environment maps supplied by the HMD 51 and tracker 52, and image data 232 used to generate images in the virtual space.

[0080] Furthermore, the environmental map management unit 201, the 3D reconstruction unit 202, and the display image generation unit 203, as well as the matching processing unit 211 and the surface 3D shape estimation unit 212 of the 3D reconstruction unit 202, and the shared environmental map 231 and image data 232 will be described later in conjunction with the explanation of the functions realized by the information processing system 31 of Figure 2, which will be explained with reference to Figure 6.

[0081] Furthermore, since the information processing device 53 acquires the poses of the head, both hands, torso, and both feet from the HMD 51 and tracker 52 worn by multiple users 41-1 to 41-n and generates images in the virtual space, it is desirable that the processor and memory constituting its processing circuit 191 have higher processing capabilities than the processing circuits 101 and 151 of the HMD 51 and tracker 52.

[0082] <Functions realized by the information processing system> Next, referring to the functional block diagram in Figure 7, the functions realized by the information processing system 31 in Figure 2 will be explained.

[0083] The cameras 71-1 to 71-4 of the HMD 51 each capture images of the area around the HMD 51 (images of the area around the head on which the HMD 51 is worn), each with a different field of view, and output these images to the environment map creation unit 121, the pose estimation unit 122, and the information processing device 53.

[0084] More specifically, cameras 71-1 to 71-4 capture images of the surroundings of the HMD 51 and supply them to the environment map creation unit 121 and the pose estimation unit 122 of the processing circuit 101. The processing circuit 101 controls the communication unit 105 to transmit the images supplied by cameras 71-1 to 71-4 to the information processing device 53 via the network 50.

[0085] The motion sensor 72 measures the acceleration and angular velocity of the HMD 51 in three axes and outputs the six-axis measurement results to the environment map creation unit 121 and the pose estimation unit 122.

[0086] The environmental mapping unit 121 estimates the self-position and orientation (pose) of the HMD 51 (user 41's head) based on images captured by cameras 71-1 to 71-4, using technologies such as Visual SLAM (Simultaneous Localization And Mapping) and Visual Positioning System, and simultaneously creates an environmental map of the surrounding area.

[0087] Furthermore, the environmental map creation unit 121 creates an environmental map based on the six-axis sensing results of the HMD 51 (acceleration in three axes and angular velocity in three axes) measured by the motion sensor 72.

[0088] The environmental map creation unit 121 then combines the environmental map created based on the image with the environmental map based on the six-axis sensing results, and controls the communication unit 105 to supply it to the information processing device 53.

[0089] Furthermore, the environmental map creation unit 121 primarily uses an environmental map created based on images, and uses the 6-axis sensing results as a supplement to improve the accuracy of the environmental map. For this reason, the motion sensor 72 is not an essential component and may be omitted if necessary. In this case, the environmental map creation unit 121 outputs the environmental map created based on images to the information processing device 53.

[0090] Similarly, the cameras 81-1 and 81-2 of the tracker 52 capture images of the area around the tracker 52 (images of the hands, torso, and feet of the user 41 to whom the tracker 52 is attached) with different fields of view, and output these images to the environment map creation unit 171, the pose estimation unit 172, and the information processing device 53.

[0091] More specifically, cameras 81-1 and 81-2 capture images of the area around the tracker 52 and supply them to the environment map creation unit 171 and the pose estimation unit 172 of the processing circuit 151. The processing circuit 151 controls the communication unit 155 to transmit the images supplied by cameras 81-1 and 81-2 to the information processing device 53 via the network 50.

[0092] The motion sensor 82 measures the acceleration and angular velocity of the tracker 52 in three axes and outputs the six-axis measurement results to the environment map creation unit 171 and the pose estimation unit 172.

[0093] The environmental mapping unit 171 of the tracker 52 estimates the self-position and posture (pose) of the tracker 52 (both hands, torso, and both feet of the user 41 to whom the tracker 52 is attached) using technologies such as Visual SLAM (Simultaneous Localization And Mapping) and Visual Positioning System, based on images captured by cameras 81-1 and 81-2, and simultaneously creates an environmental map of the surrounding area.

[0094] Furthermore, the environmental map creation unit 171 creates an environmental map based on the six-axis sensing results of the tracker 52 (acceleration in three axes and angular velocity in three axes) estimated by the motion sensor 82.

[0095] The environmental map creation unit 171 then combines the environmental map created based on the image with the environmental map based on the sensing results in the six axes, and controls the communication unit 155 to supply it to the information processing device 53.

[0096] Furthermore, the environmental map creation unit 171 primarily uses an environmental map created based on images, and uses the 6-axis sensing results as a supplement to improve the accuracy of the environmental map. For this reason, the motion sensor 82 is not an essential component and may be omitted if necessary. In this case, the environmental map creation unit 171 outputs the environmental map created based on images to the information processing device 53.

[0097] The environmental map management unit 201 of the information processing device 53 generates a shared environmental map 231 by integrally managing the individual environmental maps supplied by the HMD 51 and the tracker 52, and stores it in the storage unit 194.

[0098] In other words, the shared environment map 231 is an environment map that is integrated and managed by connecting individually generated environment maps in the HMD 51 and tracker 52. For this reason, the shared environment map 231 can be considered a massive environment map formed by mutually complementing areas that cannot be fully captured by individual environment maps.

[0099] The environmental map management unit 201 controls the communication unit 195 to transmit the shared environmental map 231 to the HMD 51 and the tracker 52, respectively.

[0100] The pose estimation unit 122 of the HMD 51 estimates the pose of the HMD 51 (the head of the user 41 wearing it) based on the shared environment map 231 supplied by the information processing device 53, the images supplied by cameras 71-1 to 71-4, and the 6-axis sensing results supplied by the motion sensor 72, and supplies the pose estimation result to the information processing device 53.

[0101] Similarly, the pose estimation unit 172 of the tracker 52 estimates the pose of the tracker 52 (both hands, torso, and feet of the user 41 wearing it) based on the shared environment map 231 supplied by the information processing device 53, the images supplied by cameras 81-1 and 81-2 respectively, and the 6-axis sensing results supplied by the motion sensor 82, and supplies the pose estimation result to the information processing device 53.

[0102] Furthermore, in both the pose estimation unit 122 of the HMD 51 and the pose estimation unit 172 of the tracker 52, if the motion sensors 72 and 82 are omitted as described above, the pose is estimated based on the shared environment map 231 and the images captured by the respective cameras 71 and 81.

[0103] The 3D reconstruction unit 202 of the information processing device 53 reconstructs a 3D model of the user 41 based on the images supplied from the HMD 51 and trackers 52-1 to 52-5, as well as the pose estimation results, and supplies it to the display image generation unit 203. The 3D reconstruction unit 202 performs the same processing for multiple users 41, reconstructing a 3D model for each user and supplying it to the display image generation unit 203.

[0104] More specifically, the 3D reconstruction unit 202 includes a matching processing unit 211 and a surface 3D shape estimation unit 212.

[0105] The matching processing unit 211 searches for images containing the same subject based on the images supplied from the HMD 51 and trackers 52-1 to 52-5, as well as the pose estimation results, and then matches the corresponding positions within the images.

[0106] The surface 3D shape estimation unit 212 uses the processing results from the matching processing unit 211 to reconstruct a 3D model of the user 41, which is the subject, for example, using SfM (Structure from Motion) or deep learning, and supplies the reconstructed 3D model to the display image generation unit 203.

[0107] More specifically, the surface 3D shape estimation unit 212 uses the processing results from the matching processing unit 211 to determine the correspondence between identical positions that constitute the surface of the user 41, which is the same subject in multiple images, and the estimated poses of the cameras 71 or 81 that captured each image. Using the principle of triangulation, the unit determines the 3D coordinates of each point that constitutes the surface of the user 41, which is the subject in the image, thereby reconstructing a 3D model of the user 41.

[0108] That is, for example, consider the case where the user 41, who is the subject of the photograph, is captured as images V11 and V12 by cameras 81-11 and 81-12 of trackers 52-11 and 52-12 attached to the user's left hand 41AL and right hand 41AR, respectively.

[0109] For example, in Figure 8, the matching processing unit 211 searches for images V11 and V12 in which the same subject, user 41, is captured (contained within the images) by matching from among multiple images. The matching processing unit 211 also determines the correspondence between each point within the region constituting the surface of user 41, which is the subject, in each of images V11 and V12. That is, in Figure 8, for example, within the region constituting the surface of user 41, points Pc11 and Pc12, Ps11 and Ps12, and Pt11 and Pt12 correspond, so the matching processing unit 211 determines the correspondence between each point between these images.

[0110] Then, the surface 3D shape estimation unit 212 determines the positions of corresponding points Pc, Ps, and Pt on the surface of the user 41 in real space using the principle of triangulation, based on the images V11 and V12 retrieved by the matching processing unit 211, the correspondence between each point in the region constituting the surface of the user 41 within each image, and the pose estimation results of the cameras 81-11 and 81-12 that captured the images V11 and V12.

[0111] The surface 3D shape estimation unit 212 reconstructs a 3D model of the user 41 by determining each point on the surface of the user 41's head, both hands, torso, and both feet using a method similar to the method described with reference to Figure 8.

[0112] In this case, by using images, it becomes possible to appropriately reproduce the surface shape and texture at each point that makes up the surface of the reproduced 3D model. Furthermore, it is not always necessary for both the surface shape and texture to be reproduced in the 3D model; at least one of them may be adopted to reduce the processing load. In addition, although Figure 8 has illustrated an example of using the principle of triangulation with two images, the surface shape and texture may be determined with higher accuracy by using three or more images of the same subject, the user 41.

[0113] Furthermore, by combining the pre-estimated poses of the HMD 51 and tracker 52 using images captured by cameras 71 and 81, with the pose estimation results obtained using motion sensors 72 and 82, the number of unknowns in the calculation is reduced, thereby lowering the burden on the computation process. This makes it possible to reconstruct the user's 3D model faster and with higher accuracy. In particular, by increasing the processing speed, real-time processing is expected to be realized.

[0114] Furthermore, although it is assumed that user 41 will wear the HMD 51, it is also possible for user 41 to refrain from wearing the HMD 51 and use only the tracker 52 until the 3D model of the head, especially the face, is reconstructed.

[0115] As a result, the 3D model of user 41 will reflect user 41's current appearance, including face shape, clothing, and body shape (surface shape and texture).

[0116] The display image generation unit 203 uses the 3D model for each user 41 supplied by the 3D reconstruction unit 202, the poses from the HMD 51 and tracker 52 respectively, and the image material data 232 necessary for the virtual space image construction to generate a virtual space image that includes an avatar reflecting the user's movements and appearance, and transmits it to each user 41's HMD 51.

[0117] When the display control unit 123 of the HMD 51 acquires the image of the virtual space transmitted from the information processing apparatus 53, it displays the image on the display unit 73.

[0118] With the above configuration, when expressing an avatar that is the alter ego of the user in an image of a virtual space, it is possible not only to reflect the user's movements, but also to reflect the user's appearance in the virtual space.

[0119] In particular, by applying a three-dimensional model to an avatar in a virtual world, it is also possible to reproduce the face shape, clothing, and body shape of the user 41 on an image in the virtual space. At this time, the face shape, clothing, and shape of the user 41 may be drawn so as to be faithfully reproduced on the image in the virtual space (when emphasis is placed on a sense of reality), or may be deformed and drawn so as to match the worldview in the virtual space.

[0120] <Device for Generating Three-Dimensional Model> In the above processing, appropriate three-dimensional model reconstruction is achieved by imaging the surface shape of the user 41 serving as a subject from various viewpoints without omission. Therefore, in portions where the imaging result is unclear and the surface shape cannot be estimated, or in missing portions (portions that have not been imaged), there is a risk that an incomplete state such as holes being formed in the three-dimensional model may occur.

[0121] In the present disclosure, since a plurality of cameras 81 are provided not only on the plurality of cameras 71 of the HMD 51 but also on the trackers 52, for example, by changing the positions and orientations of the trackers 52 mounted on both hands and both feet, the configuration is such that a full-body image of the user 41 can be captured relatively easily. Therefore, the configuration is such that a three-dimensional model of the entire body of the user 41 can be generated easily and with high accuracy.

[0122] However, if an explicit work period for reconstructing the three-dimensional model is not provided and the three-dimensional model is reconstructed during natural movement, depending on the user's movement, blind spots (regions such as the back and the back of the head that cannot be imaged by any of the cameras 71 and 81 through only natural movement) may occur.

[0123] While it is likely that interpolation can be performed from the visible portion in many cases, if cameras 71 and 81 are to capture images of the actual object, a UI may be presented to the user that guides cameras 71 and 81 to reduce blind spots based on the estimated position information of cameras 71 and 81.

[0124] For example, when the display image generation unit 203 generates an avatar based on the user 41's 3D model, it may generate an image showing the reconstructible range and proportion of the body surface of the reconstructed 3D model, as shown in Figure 9, and supply it to the HMD 51 to display on the display unit 73. In this way, the user 41 can understand which areas are blind spots, which areas the cameras 71 and 81 should be pointed at, and what the current level of completion is.

[0125] The display image generation unit 203 may also generate an image showing the current position of the cameras on both hands and both feet, and the position to which they should be moved to fill the blind spot area, and supply this image to the HMD 51 for display, or it may indicate the direction to move using arrows or the like.

[0126] In Figure 9, the left side represents the front view 331F of the user 41's 3D model, and the right side represents the back view 331B of the user 41's 3D model. The captured area is represented as a gray area, and the blind spot area is represented as a white area.

[0127] Specifically, the left side of Figure 9 shows that in the front view 331F of the user 41's 3D model, the left foot, the left fingertips, and the left side of the face are not being captured. The right side of Figure 9 shows that in the back view 331B of the user 41's 3D model, almost the entire back, the left foot, the left fingertips, and the left side of the back of the head are not being captured. With a presentation like that shown in Figure 9, the user 41 can relatively easily reduce the blind spot by, for example, adjusting the orientation of the camera 71 of the HMD 51 or the camera 81 of the tracker 52 to capture the areas represented by the white areas.

[0128] Furthermore, as shown in Figure 10, a mirror 351 may be used to reduce the blind spot area.

[0129] For example, the left side of Figure 10 shows an example where the user 41 is facing the mirror 351 directly (state 41F). The user 41FM reflected in the mirror 351 may be captured by the camera 71 of the HMD 51 or the camera 81 of the tracker 52.

[0130] Furthermore, as shown in the center of Figure 10, an example is shown where the user 41 is facing left relative to the mirror 351 (state 41S). In this case, the user 41SM reflected in the mirror 351 may be imaged by the camera 71 of the HMD 51 or the camera 81 of the tracker 52.

[0131] Furthermore, as shown in the right part of Figure 10, an example is shown where the user 41 is facing away from the mirror 351 (state 41B), and the camera 81 of the tracker 52 may capture the user 41BM reflected in the mirror 351. However, in the case shown in the right part of Figure 10, it is practically impossible for the camera 71 of the HMD 51 to capture the user 41BM.

[0132] In the case of Figure 10, for example, cameras 71 and 81 may capture two images: one directly captured in real space and another reflected by the mirror 351. In such cases, the image that is further away from cameras 71 and 81 is considered to be the image reflected by the mirror 351. Therefore, the image that is further away from cameras 71 and 81 may be selected.

[0133] Furthermore, by marking each of the HMD 51 and tracker 52 individually, the cameras 71 and 81 can recognize that they are imaging themselves, depending on their positional relationship with the mirror 351. In this case, the positions of the cameras 71 and 81 and the mirror 351 can be determined from their relationship to their own positions in the images captured by the cameras 71 and 81. In such cases, the part of the user 41 in the captured image can be identified based on the positional relationship between the mirror 351 and the cameras 71 and 81, and used for matching.

[0134] <Display Processing of Virtual Space Images> Next, referring to the flowchart in Figure 11, the display processing of virtual space images by the information processing system 31 in Figure 2 will be explained. Note that, apart from acquiring and displaying virtual space images, the processing of the HMD 51 and the tracker 52 is basically the same, so the processing of the HMD 51 and the tracker 52 will be explained together.

[0135] In steps S31 and S51, the motion sensor 72 of the HMD 51 and the motion sensor 82 of the tracker 52 detect acceleration in three axes and angular velocity in three axes, respectively, and output the six-axis sensing results to the environment map creation units 121 and 171.

[0136] In steps S32 and S52, the camera 71 of the HMD 51 and the camera 81 of the tracker 52 capture images and output them to the environmental map creation units 121 and 171, respectively.

[0137] In steps S33 and S53, the environmental map creation unit 121 of the HMD 51 and the environmental map creation unit 171 of the tracker 52 each create an environmental map based on the 6-axis sensing results and multiple images.

[0138] In steps S34 and S54, the environmental map creation unit 121 of the HMD 51 and the environmental map creation unit 171 of the tracker 52 output the created environmental map and image to the information processing device 53.

[0139] In step S71, the environmental map management unit 201 of the information processing device 53 acquires the environmental map transmitted from the HMD 51 and the environmental map transmitted from the tracker 52. At this time, the matching processing unit 211 acquires the image transmitted along with the environmental map.

[0140] In step S72, the environment map management unit 201 reads the environment map transmitted from the HMD 51 and the environment map transmitted from the tracker 52, as well as the shared environment map 231 stored in the storage unit 194, and updates the shared environment map 231 by combining them. In the case of a new shared environment map 231 that is not stored in the storage unit 194, a new shared environment map 231 is generated using the environment map transmitted from the HMD 51 and the environment map transmitted from the tracker 52, and stored in the storage unit 194.

[0141] In step S73, the environmental map management unit 201 reads the shared environmental map 231 stored in the memory unit 194 and transmits it to the HMD 51 and the tracker 52.

[0142] In steps S35 and S55, the pose estimation unit 122 of the HMD 51 and the pose estimation unit 172 of the tracker 52 each acquire the shared environment map 231.

[0143] In steps S36 and S56, the pose estimation unit 122 of the HMD 51 and the pose estimation unit 172 of the tracker 52 estimate the pose of the HMD 51 and the pose of the tracker 52, respectively, based on the 6-axis sensing results, multiple images, and the shared environment map 231.

[0144] In steps S37 and S57, the pose estimation unit 122 of the HMD 51 and the pose estimation unit 172 of the tracker 52 transmit the estimated poses of the HMD 51 and the tracker 52, respectively, to the information processing device 53.

[0145] In step S74, the matching processing unit 211 and the display image generation unit 203 in the three-dimensional reconstruction unit 202 of the information processing device 53 acquire the respective poses supplied from the HMD 51 and the tracker 52.

[0146] In step S75, the matching processing unit 211 matches the correspondence of the surface positions of the user 41, who is the subject, based on the poses and images supplied from the HMD 51 and the tracker 52, and outputs the resulting image and the correspondence of each point on the image to the surface 3D shape estimation unit 212.

[0147] In step S76, the surface 3D shape estimation unit 212 estimates a 3D model of the user 41, which is the subject of the experiment, based on the matching result image supplied by the matching processing unit 211 and the correspondence between each point on the image, and supplies the estimated 3D model to the display image generation unit 203.

[0148] In step S77, the display image generation unit 203 determines whether the user 41's 3D model is sufficiently complete. That is, the display image generation unit 203 may determine whether the 3D model is sufficiently complete based, for example, on whether the blind spot area, as explained with reference to Figure 9, is higher than a predetermined proportion of the entire 3D model.

[0149] Furthermore, if the blind spot area is higher than a predetermined proportion of the entire 3D model and is determined to be insufficient in its completeness, an image showing the location of the blind spot area, as shown in Figure 9, may be generated, supplied to the HMD 51, and displayed on the display unit 73 for the user to see.

[0150] If, in step S77, it is determined that the user 41's 3D model is not sufficiently complete, the process returns to step S71, and the subsequent processes are repeated. That is, the processes from steps S71 to S77 are repeated until the poses of the HMD 51 and tracker 52 change in accordance with the user 41's movements, new images are captured, the blind spots are reduced, and as a result, it is determined that the 3D model is sufficiently complete.

[0151] Then, if it is determined in step S77 that the user 41's 3D model is sufficiently complete, the process proceeds to step S78.

[0152] In step S78, the display image generation unit 203 generates an avatar image that represents user 41, based on user 41's 3D model and pose, and reflects user 41's movements and appearance.

[0153] In step S79, the display image generation unit 203 generates a virtual space image based on the user 41's avatar image, the avatar images of other users, and a background image based on the image material data 232.

[0154] In step S80, the display image generation unit 203 transmits the generated virtual space image to the HMD 51.

[0155] In step S38, the display control unit 123 of the HMD 51 acquires the virtual space image supplied from the information processing device 53.

[0156] In step S39, the display control unit 123 displays the acquired virtual space image on the display unit 73.

[0157] In steps S40, S58, and S81, it is determined whether or not the termination of the process has been instructed. If termination is not instructed, the process returns to steps S31, S51, and S71, respectively, and the subsequent processes are repeated.

[0158] Then, if the termination of the process is instructed in steps S40, S58, or S81, the process terminates.

[0159] In other words, through the above process, when generating a virtual space image, it becomes possible to reflect not only the movements of user 41, but also the appearance of user 41, including the clothing and body shape they are currently wearing, in the avatar, which is a representation of user 41.

[0160] In the above, we have described an example in which a virtual space image reflecting the movements and appearance of user 41 is not generated and displayed until it is determined that the 3D model of user 41 is sufficiently complete. However, it is also possible to have the avatar reflect the user's movements and appearance and present it as a virtual space image regardless of whether the 3D model is sufficiently complete or not.

[0161] In this case, the virtual space image reflected in the avatar will be displayed with blind spots in the 3D model. However, the user 41 can recognize the blind spots by looking at the avatar presented in an incomplete state, and can recognize the actions necessary to reduce the blind spots while looking at the avatar displayed in the virtual space image.

[0162] As described above, this disclosure makes it possible to create an avatar in a three-dimensional virtual space that reflects not only the user's movements but also their appearance.

[0163] As a result, as explained with reference to Figure 1, for example, in the case of a first-person perspective from the user's point of view, it becomes possible to represent it in a way similar to a UI (User Interface) image.

[0164] Furthermore, for example, in the case of a third-person perspective, which is the viewpoint of another user, it becomes possible to create an effect on the other user's HMD 51 that makes it appear as if the user themselves is present in the virtual space, similar to telepresence.

[0165] As a result, based on the images captured by the cameras 71 and 81 of the HMD 51 and trackers 52-1 to 52-5, it becomes possible to appropriately reflect the user's movements and appearance in the avatar, which is the user's digital counterpart in the virtual space.

[0166] Furthermore, while the above describes an example in which the information processing device 53 manages the environment map, images, and poses supplied from the HMD 51 and tracker 52 to generate a virtual space image to be displayed on the HMD 51, the functions of the information processing device 53 may also be implemented by the HMD 51. That is, the information processing system 31 may be configured using only the HMD 51 and the tracker 52. In this case, the HMD 51 acquires information on the environment map, images, and poses from the tracker 52, generates and manages the shared environment map 231, reconstructs the user 41's 3D model, and generates and displays a virtual space image. Alternatively, one of the trackers 52 may function as the information processing device 53. Furthermore, the information processing device 53 may be implemented by cloud computing consisting of multiple computers on a network.

[0167] <<3. Description of a computer using this technology>>

[0168] The series of processes described above can be executed by hardware or by software. When the series of processes are executed by software, the programs that make up that software are installed on a computer. Here, "computer" includes computers built into dedicated hardware, as well as general-purpose personal computers, for example, that can perform various functions by installing various programs.

[0169] Figure 12 is a block diagram showing an example of the hardware configuration of a computer that executes the series of processes described above by a program.

[0170] In a computer, the processing circuit 1001, ROM (Read Only Memory) 1002, and RAM (Random Access Memory) 1003 are interconnected by a bus 1004.

[0171] An input / output interface 1005 is further connected to the bus 1004. An input / output interface 1005 is connected to an input unit 1006, an output unit 1007, a storage unit 1008, a communication unit 1009, and a drive 1010.

[0172] The input unit 1006 may include physical or virtual operating means that the user operates to input information, such as a keyboard, mouse, or touch panel, as well as means that the user inputs information through voice, eye gaze, etc. Furthermore, the input unit 1006 may include sensors for inputting various physical quantities into the computer. For example, the input unit 1006 may include sensors that acquire physical quantities such as light (including infrared light other than visible light) or sound, such as a camera or microphone. Also, for example, the input unit 1006 may include sensors that acquire other physical quantities such as temperature, moisture content, acceleration, and distance. The output unit 1007 may include means that present information to the user by stimulating the user's perception, such as a display, speaker, or haptic device. The storage unit 1008 is composed of a hard disk, non-volatile or volatile memory, etc., and stores various information (including programs). The communication unit 1009 is a network interface, etc., and performs wired or wireless communication with the outside. The drive 1010 drives removable media 1011 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.

[0173] The processing circuit 1001 includes a processor that executes programs such as a CPU (Central Processing Unit) and a DSP (Digital Signal Processor). The processing circuit 1001 (its processor) performs the series of processes described above by loading the program stored in the memory unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing it. The processing circuit 1001 can output the processing results of the series of processes from the output unit 1007, for example, via the bus 1004 and the input / output interface 1005, as needed. The processing circuit 1001 can also store the processing results in the memory unit 1008 or transmit them from the communication unit 1009.

[0174] The program executed by the computer (processing circuit 1001) can be provided by recording it on a removable medium 1011, such as a package medium. The program can also be provided via wired or wireless transmission media, such as a local area network, the internet, or digital satellite broadcasting.

[0175] In a computer, a program can be installed in the storage unit 1008 via the input / output interface 1005 by inserting the removable media 1011 into the drive 1010. Alternatively, a program can be received by the communication unit 1009 from another device, such as a server, via a wired or wireless transmission medium, and installed in the storage unit 1008. Furthermore, programs can be pre-installed in the ROM 1002 or the storage unit 1008.

[0176] The programs executed by the computer may be programs that are processed chronologically in the order described herein, or they may be programs that are processed in parallel or at necessary times, such as when a call is made.

[0177] The processes that a computer performs according to a program do not necessarily have to follow the order described in the flowchart. In other words, the processes that a computer performs according to a program include processes that are executed in parallel or individually (e.g., parallel processing and object-based processing).

[0178] The program may be processed by a single computer (processor), or it may be processed in a distributed manner by multiple computers. Furthermore, the program may be transferred to a remote computer and executed there.

[0179] When the computer executes a program to perform the above-described series of processes, the processing circuit 1001 (its processor) executes the program to function as the environment map creation unit 121, pose estimation unit 122, and display control unit 123 in the processing circuit 101 of Figure 4, the environment map creation unit 171 and pose estimation unit 172 in the processing circuit 151 of Figure 5, and the environment map management unit 201, 3D reconstruction unit 202, and display image generation unit 203 in the processing circuit 191 of Figure 6.

[0180] In this specification, a system means one component or a collection of multiple components (devices, modules (parts), etc.). Therefore, one or more components of a computer, for example, only the processor, or a combination of the processor and memory, for example, only the processing circuit 1001, or a combination of the processing circuit 1001 to the bus 1004, etc., constitute a system. Regarding a collection of multiple components, it is not necessary whether all components reside in the same enclosure or not. Therefore, multiple devices housed in separate enclosures and connected via a network, or a single device containing multiple modules within a single enclosure, are all systems. Furthermore, for example, the entire computer, or a combination of a computer and other devices such as a server (not shown), also constitute a system.

[0181] Furthermore, for example, each step of a flowchart may be executed by one device, or it may be divided among multiple devices. Additionally, if a single step includes multiple processes, these processes may be executed by one device, or they may be divided among multiple devices. In other words, multiple processes included in a single step can be executed as multiple steps. Conversely, processes described as multiple steps can be combined and executed as a single step.

[0182] Furthermore, for example, a program executed by a computer may be structured so that the steps of the program are executed chronologically in the order described herein, or they may be executed in parallel or individually at necessary times, such as when a call is made. In other words, the steps may be executed in an order different from the order described above, as long as no inconsistencies arise. Moreover, the steps of this program may be executed in parallel with the processing of other programs, or in combination with the processing of other programs.

[0183] Furthermore, for example, the various technologies relating to this disclosure can be implemented independently, as long as they do not conflict with each other. Of course, any combination of the disclosures can also be implemented. For example, some or all of the disclosures described in one embodiment can be combined with some or all of the disclosures described in another embodiment. Also, some or all of the aforementioned disclosures can be implemented in combination with other technologies not described above.

[0184] Furthermore, this disclosure may also take the following configurations: <1> An information processing system comprising: an HMD (Head Mounted Display) having a plurality of cameras worn by a user; a plurality of trackers having a plurality of cameras worn by the user; a 3D model configuration unit that constructs a 3D model of the user based on a plurality of images captured by the plurality of cameras of the HMD and the plurality of cameras of the plurality of trackers; and an avatar generation unit that generates an avatar of the user using the 3D model. <2> The information processing system according to <1>, wherein the 3D model configuration unit matches an image of the user being captured from the plurality of images and constructs the 3D model of the user based on the matched image. <3> The information processing system according to <2>, wherein the 3D model configuration unit matches an image of the user being captured from the plurality of images and constructs the 3D model of the user based on the correspondence between each point in the area of ​​the user being captured between the matched images. <4> The information processing system according to <3>, further comprising: an HMD pose estimation unit that estimates the pose of the HMD, consisting of the position and orientation of the HMD, based on images captured by a plurality of cameras of the HMD; and a tracker pose estimation unit that estimates the pose of the tracker, consisting of the position and orientation of the tracker, from a plurality of images captured by a plurality of cameras of the tracker, wherein the 3D model configuration unit configures the 3D model of the user based on the correspondence between each point in the region in which the user is being captured between the matched images, and at least one of the pose of the HMD and the pose of the tracker in the matched images.<5> The information processing system according to <4>, further comprising: an HMD environment map generation unit that generates an environment map of the surroundings of the HMD based on images captured by multiple cameras of the HMD; a tracker environment map generation unit that generates an environment map of the surroundings of the tracker from multiple images captured by multiple cameras of the tracker; and a shared environment map management unit that integrates the environment map of the surroundings of the HMD and the environment map of the surroundings of the tracker, manages them as a shared environment map, and supplies them to the HMD and the tracker, wherein the HMD pose estimation unit estimates the pose of the HMD based on the shared environment map and images captured by multiple cameras of the HMD, and the tracker pose estimation unit estimates the pose of the tracker from the shared environment map and multiple images captured by multiple cameras of the tracker. <6> The information processing system according to <4>, wherein the three-dimensional model component constructs the three-dimensional model of the user using the principle of triangulation based on the correspondence between each point in the area where the user is being imaged in the matched images and at least one of the pose of the HMD and the pose of the tracker in the matched images. <7> The information processing system according to <3>, wherein the three-dimensional model component constructs the three-dimensional model of the user based on at least one of the surface shape and texture of each point, and the avatar generation unit generates the user's avatar based on at least one of the surface shape and texture of each point constituting the three-dimensional model. <8> The information processing system according to <1>, wherein the avatar generation unit determines whether the three-dimensional model constructed by the three-dimensional model component is complete, and if the completeness is sufficient, generates the user's avatar based on the three-dimensional model. <9> The information processing system described in <8>, wherein the avatar generation unit determines whether the completeness of the three-dimensional model is sufficient based on the proportion of the camera's blind spot area in the three-dimensional model where the configuration based on the image is not made.<10> The information processing system according to <9>, wherein the avatar generation unit determines that the level of completion of the three-dimensional model is insufficient, and generates an image indicating the position of the blind spot region of the camera. <11> An information processing method comprising: a three-dimensional model configuration process that constitutes a three-dimensional model of the user based on a plurality of images captured by a plurality of cameras of an HMD (Head Mounted Display) having a plurality of cameras worn by the user and a plurality of cameras of a plurality of trackers having a plurality of cameras worn by the user; and an avatar generation process that generates an avatar of the user using the three-dimensional model. <12> The information processing method according to <11>, wherein the three-dimensional model configuration process matches an image of the user being captured from the plurality of images and constitutes the three-dimensional model of the user based on the matched image. <13> The information processing method according to <12>, wherein the three-dimensional model configuration process matches an image of the user being captured from the plurality of images and constitutes the three-dimensional model of the user based on the correspondence between each point in the region of the user being captured between the matched images. <14> The information processing method according to <13>, further comprising: an HMD pose estimation process that estimates the pose of the HMD, consisting of the position and orientation of the HMD, based on images captured by a plurality of cameras of the HMD; and a tracker pose estimation process that estimates the pose of the tracker, consisting of the position and orientation of the tracker, from a plurality of images captured by a plurality of cameras of the tracker, wherein the 3D model configuration process configures the 3D model of the user based on the correspondence between each point in the region in which the user is being captured in the matched images, and at least one of the pose of the HMD and the pose of the tracker in the matched images.<15> The information processing method according to <14>, further comprising: an HMD environment map generation process that generates an environment map of the surroundings of the HMD based on images captured by multiple cameras of the HMD; a tracker environment map generation process that generates an environment map of the surroundings of the tracker from multiple images captured by multiple cameras of the tracker; and a shared environment map management process that integrates the environment map of the surroundings of the HMD and the environment map of the surroundings of the tracker, manages them as a shared environment map, and supplies them to the HMD and the tracker, wherein the HMD pose estimation process estimates the pose of the HMD based on the shared environment map and images captured by multiple cameras of the HMD; and the tracker pose estimation process estimates the pose of the tracker from the shared environment map and multiple images captured by multiple cameras of the tracker. <16> The 3D model construction process constructs the user's 3D model using the principle of triangulation based on the correspondence between each point in the area where the user is being imaged in the matched images, and at least one of the pose of the HMD and the pose of the tracker in the matched images. The information processing method according to <14>. <17> The 3D model construction process constructs the user's 3D model based on at least one of the surface shape and texture of each point, and the avatar generation process generates the user's avatar based on at least one of the surface shape and texture of each point constituting the 3D model. The information processing method according to <13>. <18> The avatar generation process determines whether the 3D model constructed by the 3D model construction process is complete, and if it is complete, generates the user's avatar based on the 3D model. The information processing method according to <11>. <19> The avatar generation process determines whether the completeness of the three-dimensional model is sufficient based on the proportion of the camera's blind spot area in the three-dimensional model where the configuration based on the image is not made. The information processing method described in <18>.<20> The information processing method according to <19>, which, when the avatar generation process determines that the level of completion of the three-dimensional model is insufficient, generates an image that shows the position of the blind spot region of the camera.

[0185] 31 Information processing system, 51 HMD, 52, 52-1 to 52-5 Tracker, 53 Information processing device, 71, 71-1 to 71-4 Camera, 72 Motion sensor, 73 Display unit, 81, 81-1, 81-2 Camera, 82 Motion sensor, 121 Environment map creation unit, 122 Pose estimation unit, 123 Display control unit, 171 Environment map creation unit, 172 Pose estimation unit, 201 Environment map management unit, 202 3D reconstruction unit, 203 Display image generation unit, 211 Matching processing unit, 212 Surface 3D shape estimation unit, 231 Shared environment map

Claims

1. An information processing system comprising: an HMD (Head Mounted Display) having multiple cameras worn by a user; multiple trackers having multiple cameras worn by the user; a 3D model constructing unit that constructs a 3D model of the user based on multiple images captured by the multiple cameras of the HMD and the multiple cameras of the multiple trackers; and an avatar generation unit that generates an avatar of the user using the 3D model.

2. The information processing system according to claim 1, wherein the three-dimensional model component matches an image captured by the user from the plurality of images, and constructs the user's three-dimensional model based on the matched image.

3. The information processing system according to claim 2, wherein the three-dimensional model component matches the image captured by the user from the plurality of images, and constructs the three-dimensional model of the user based on the correspondence between each point in the area captured by the user between the matched images.

4. The information processing system according to claim 3, further comprising: an HMD pose estimation unit that estimates the pose of the HMD, consisting of the position and orientation of the HMD, based on images captured by a plurality of cameras of the HMD; and a tracker pose estimation unit that estimates the pose of the tracker, consisting of the position and orientation of the tracker, from a plurality of images captured by a plurality of cameras of the tracker, wherein the 3D model configuration unit configures the 3D model of the user based on the correspondence between each point in the region in which the user is being captured between the matched images, and at least one of the pose of the HMD and the pose of the tracker in the matched images.

5. The information processing system according to claim 4, further comprising: an HMD environment map generation unit that generates an environment map of the surroundings of the HMD based on images captured by multiple cameras of the HMD; a tracker environment map generation unit that generates an environment map of the surroundings of the tracker from multiple images captured by multiple cameras of the tracker; and a shared environment map management unit that integrates the environment map of the surroundings of the HMD and the environment map of the surroundings of the tracker, manages them as a shared environment map, and supplies them to the HMD and the tracker, wherein the HMD pose estimation unit estimates the pose of the HMD based on the shared environment map and images captured by multiple cameras of the HMD; and the tracker pose estimation unit estimates the pose of the tracker from the shared environment map and multiple images captured by multiple cameras of the tracker.

6. The information processing system according to claim 4, wherein the three-dimensional model component constructs the three-dimensional model of the user using the principle of triangulation, based on the correspondence between each point in the area in which the user is being imaged in the matched images, and at least one of the pose of the HMD and the pose of the tracker in the matched images.

7. The information processing system according to claim 3, wherein the three-dimensional model component constitutes the user's three-dimensional model based on at least one of the surface shape and texture of each point, and the avatar generation component generates the user's avatar based on at least one of the surface shape and texture of each point constituting the three-dimensional model.

8. The information processing system according to claim 1, wherein the avatar generation unit determines whether the three-dimensional model constructed by the three-dimensional model building unit is sufficiently complete, and if the degree of completion is sufficient, generates an avatar of the user based on the three-dimensional model.

9. The information processing system according to claim 8, wherein the avatar generation unit determines whether the level of completion of the three-dimensional model is sufficient based on the proportion of the camera's blind spot area in the three-dimensional model where the configuration based on the image is not made.

10. The information processing system according to claim 9, wherein the avatar generation unit determines that the level of completion of the three-dimensional model is insufficient, and generates an image indicating the position of the blind spot region of the camera.

11. An information processing method comprising: a three-dimensional model configuration process that constructs a three-dimensional model of a user based on multiple images captured by multiple cameras of a Head-Mounted Display (HMD) having multiple cameras worn by the user and multiple cameras of a tracker having multiple cameras worn by the user; and an avatar generation process that generates an avatar of the user using the three-dimensional model.

12. The information processing method according to claim 11, wherein the three-dimensional model construction process matches an image captured by the user from the plurality of images, and constructs the user's three-dimensional model based on the matched image.

13. The information processing method according to claim 12, wherein the three-dimensional model construction process matches an image captured by the user from the plurality of images, and constructs the user's three-dimensional model based on the correspondence between each point in the area captured by the user between the matched images.

14. The information processing method according to claim 13, further comprising: an HMD pose estimation process that estimates the pose of the HMD, consisting of the position and orientation of the HMD, based on images captured by a plurality of cameras of the HMD; and a tracker pose estimation process that estimates the pose of the tracker, consisting of the position and orientation of the tracker, based on a plurality of images captured by a plurality of cameras of the tracker, wherein the three-dimensional model configuration process configures the three-dimensional model of the user based on the correspondence between each point in the region in which the user is being captured between the matched images, and at least one of the pose of the HMD and the pose of the tracker in the matched images.

15. The information processing method according to claim 14, further comprising: an HMD environment map generation process that generates an environment map of the surroundings of the HMD based on images captured by multiple cameras of the HMD; a tracker environment map generation process that generates an environment map of the surroundings of the tracker from multiple images captured by multiple cameras of the tracker; and a shared environment map management process that integrates the environment map of the surroundings of the HMD and the environment map of the surroundings of the tracker, manages them as a shared environment map, and supplies them to the HMD and the tracker, wherein the HMD pose estimation process estimates the pose of the HMD based on the shared environment map and images captured by multiple cameras of the HMD; and the tracker pose estimation process estimates the pose of the tracker from the shared environment map and multiple images captured by multiple cameras of the tracker.

16. The information processing method according to claim 14, wherein the three-dimensional model construction process constructs the three-dimensional model of the user using the principle of triangulation, based on the correspondence between each point in the region in which the user is being imaged in the matched images, and at least one of the pose of the HMD and the pose of the tracker in the matched images.

17. The information processing method according to claim 13, wherein the three-dimensional model construction process constructs the user's three-dimensional model based on at least one of the surface shape and texture of each point, and the avatar generation process generates the user's avatar based on at least one of the surface shape and texture of each point constituting the three-dimensional model.

18. The information processing method according to claim 11, wherein the avatar generation process determines whether the three-dimensional model constructed by the three-dimensional model construction process is sufficiently complete, and if the degree of completion is sufficient, generates the user's avatar based on the three-dimensional model.

19. The information processing method according to claim 18, wherein the avatar generation process determines whether the completeness of the three-dimensional model is sufficient based on the proportion of the camera's blind spot area in the three-dimensional model where the configuration based on the image is not made.

20. The information processing method according to claim 19, wherein the avatar generation process determines that the level of completion of the three-dimensional model is insufficient, and generates an image indicating the position of the blind spot region of the camera.