Virtual space providing device
The virtual space providing device addresses GPU load and data transmission delays by selectively acquiring and supplementing user video data, ensuring smooth and natural avatar movements for effective user communication.
Patent Information
- Application Number
- JP2024516215
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-04-22
- Filing Date
- 2023-04-11
- Publication Date
- 2025-09-03
- Estimated Expiration
- 2043-04-11
AI Technical Summary
Existing systems for generating virtual spaces with multiple user avatars face increased GPU load and data transmission delays due to the need for real-time reflection of full-body 3D content, leading to jerky movements and hindered communication.
A virtual space providing device that selectively acquires and transmits only a portion of user video data during certain periods, supplementing missing data with previous captures to generate avatars, reducing data transmission and processing delays.
Facilitates smooth communication between users by minimizing data transmission and processing delays, ensuring avatars move naturally and reducing the burden on GPU resources.
Smart Images

Figure 0007733815000001 
Figure 0007733815000002 
Figure 0007733815000003
Abstract
Description
[Technical Field]
[0001] One aspect of the present invention relates to a virtual space providing device. [Background technology]
[0002] Patent Document 1 discloses a system for generating an image of a virtual space including images of multiple users as a system for realizing communication between multiple users via a virtual space. Also known is a technology for capturing images of the user as a subject from all directions using multiple cameras or the like, and generating 3D content (Volumetric Video) that reproduces the subject's appearance, shape, movement, etc. as they are with high accuracy. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2014-56308 Summary of the Invention [Problem to be solved by the invention]
[0004] In a system such as that disclosed in Patent Document 1, in order to promote communication between multiple users via a virtual space, it is conceivable to reflect each user's 3D content in real time on a person's image (avatar) in the virtual space. However, if we try to reflect each user's full-body 3D content in real time on each user's avatar placed in the virtual space, the load on the GPU (Graphic Processing Unit) will increase and the amount of data transmission will increase. As a result, transmission delays and processing lags will occur, causing the avatar's movements in the virtual space to become jerky, which may hinder smooth communication.
[0005] Therefore, an object of one aspect of the present invention is to provide a virtual space providing device that can facilitate smooth communication between users via a virtual space. [Means for solving the problem]
[0006] A virtual space providing device according to one aspect of the present invention is a virtual space providing device that provides each user with a three-dimensional virtual space shared by the users, and includes: an acquisition unit that acquires video data of each user; a generation unit that generates an avatar to be placed in the virtual space corresponding to each user based on the video data of each user; and a provision unit that generates and provides, for each user, a video corresponding to a field of view from each user's virtual viewpoint set in the virtual space. The acquisition unit is configured to acquire first video data that shows a first part of the first user's body from video data obtained by photographing a first user from a plurality of different directions during a first period, while not acquiring second video data that shows a second part of the first user's body that is different from the first part. When the acquisition unit acquires the first video data during the first period but does not acquire the second video data, the generation unit generates the first part of the first avatar corresponding to the first user during the first period based on the first video data acquired during the first period, and generates the second part of the first avatar during the first period based on the second video data acquired during a second period prior to the first period.
[0007] According to one aspect of the present invention, a virtual space providing device selectively acquires only a portion of the first user's video data during a first period, thereby reducing the amount of video data transmitted for the first user. As a result, transmission delays, processing slowdowns, and the like caused by large amounts of data transmission can be suppressed. Furthermore, for the second portion of the video data not acquired during the first period, the second portion is supplemented with second video data acquired during a second period that precedes the first period. This allows the first avatar corresponding to the first user during the second period to be presented in a manner that appears natural to other users. As a result, the virtual space providing device can facilitate communication between users via the virtual space. [Effects of the Invention]
[0008] According to one aspect of the present invention, it is possible to provide a virtual space providing device that can facilitate communication between users via a virtual space. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of a functional configuration of a virtual space providing system according to an embodiment. [Figure 2] FIG. 10 is a diagram showing an example of a virtual space image provided to a user U2. [Figure 3] FIG. 10 is a sequence diagram showing an example of the operation of the virtual space providing system. [Figure 4] 4 is a flowchart showing a first example of the process in step S7 of FIG. 3. [Figure 5] 10 is a flowchart showing a second example of the process in step S7 of FIG. 3. [Figure 6] FIG. 2 is a diagram illustrating an example of a hardware configuration of a server included in the virtual space providing system. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the accompanying drawings. In the description of the drawings, the same or equivalent elements are designated by the same reference numerals, and redundant description will be omitted.
[0011] 1 is a diagram illustrating an example of a virtual space providing system 1 according to an embodiment. The virtual space providing system 1 is a system that provides communication between multiple users located at multiple distant locations via a virtual space.
[0012] As an example, the virtual space providing system 1 is composed of a server 10 (virtual space providing device), user terminals 20A, 20B installed at each base, HMDs (Head Mounted Displays) 30A, 30B worn on the heads of users U1, U2 at each base, and multiple cameras C arranged at each base.
[0013] Although only two locations B1 and B2 are illustrated in FIG. 1, if there are three or more users, there may be three or more locations. Also, multiple users may exist within one location. In that case, a separate user terminal may be installed for each user, or one user terminal may be shared by multiple users.
[0014] At site B1, a user terminal 20A and multiple cameras C are installed, and a user U1 (first user) wearing an HMD 30A is present. The multiple cameras C installed at site B1 are arranged around user U1 so that user U1 can be photographed from multiple different directions. The user terminal 20A acquires video data of the user U1's entire body by acquiring video data captured by each camera C. Note that if the number of cameras C installed at site B1 is insufficient (i.e., if it is not possible to acquire video data of the user U1's entire body (video data of the user U1 viewed from any direction) simply by combining the video data captured by each camera C), the missing video data may be supplemented in the user terminal 20A using AI or the like. The video data of user U1 acquired in this manner by user terminal 20A is transmitted to server 10.
[0015] At site B2, similar to site B1, a user terminal 20B and multiple cameras C are installed, and a user U2 (second user) wearing an HMD 30B is present. The multiple cameras C installed at site B2 are arranged around user U2 so that user U2 can be photographed from multiple different directions. User terminal 20B acquires full-body video data of user U2 by acquiring video data captured by each camera C. Note that if the number of cameras C installed at site B2 is insufficient (i.e., if full-body video data of user U2 (video data of user U2 viewed from any direction) cannot be acquired simply by combining the video data captured by each camera C), the missing video data may be supplemented in user terminal 20B using AI or the like. The video data of user U2 acquired in this manner by user terminal 20B is transmitted to server 10.
[0016] The user terminals 20A and 20B are computer devices configured to be able to communicate with the server 10 and multiple cameras C installed at the same location. The user terminals 20A and 20B are not limited to a specific form. Examples of the user terminals 20A and 20B include desktop PCs, laptop PCs, smartphones, tablet terminals, and wearable terminals.
[0017] In this embodiment, the user terminal 20A is configured to be able to communicate with the HMD 30A. That is, the HMD 30A is configured to be able to communicate with the server 10 via the user terminal 20A. Similarly, the user terminal 20B is configured to be able to communicate with the HMD 30B, and the HMD 30B is configured to be able to communicate with the server 10 via the user terminal 20B. However, the form of communication between the HMDs 30A and 30B and the server 10 is not limited to the above form. For example, the HMDs 30A and 30B may be configured to perform data communication directly with the server 10 without relaying through the user terminals 20A and 20B.
[0018] The HMDs 30A and 30B are devices worn on the heads of the users U1 and U2. For example, the HMDs 30A and 30B include displays (display units) placed in front of the eyes of the users U1 and U2, sensors that detect the attitudes (direction, tilt, etc.) of the HMDs 30A and 30B, and communication devices for transmitting and receiving data to and from the user terminals 20A and 20B. The HMDs 30A and 30B also include a control unit (e.g., a computer device configured with a processor, memory, etc.) that controls the operations of the above-mentioned displays, sensors, communication devices, etc. Examples of the HMDs 30A and 30B include eyeglass-type devices (e.g., smart glasses such as so-called XR glasses), goggle-type devices, hat-type devices, etc.
[0019] By viewing the images (virtual space images described later) displayed on the displays of the HMDs 30A and 30B, the users U1 and U2 can enjoy a VR experience that makes them feel as if they are in a virtual space.
[0020] FIG. 2 is a diagram showing an example of a virtual space image IM, which is an image provided to user U2 from server 10. The virtual space image IM provided to user U2 is an image corresponding to the field of view from the virtual viewpoint of user U2 set in virtual space VS. In this embodiment, the virtual viewpoint of user U2 corresponds to the first-person viewpoint of avatar A2, which is placed in virtual space VS corresponding to user U2. The virtual viewpoint of each user set in virtual space VS may change in response to the movement of the head of each user (i.e., the HMD worn on the head) (e.g., a change in posture detected by a sensor mounted on the HMD). For example, if user U2 makes a motion to turn right in real space, the head of avatar A2 in virtual space VS may also turn right in response to that motion, resulting in a change in user U2's virtual viewpoint and the field of view from the virtual viewpoint.
[0021] In the example of FIG. 2, the virtual space VS is a space simulating a virtual office room, and is arranged with an avatar A1 corresponding to user U1, an avatar A2 corresponding to user U2, and an avatar A3 corresponding to a user U3 other than users U1 and U2. More specifically, the avatars A1 to A3 are arranged to surround a table arranged in the virtual space VS. Note that the virtual space image IM shown in FIG. 2 is an image corresponding to the field of view from the virtual viewpoint of user U2 (the field of view of avatar A2), and therefore does not show avatar A2. Users U1 and U3 are provided with virtual space images corresponding to the first-person viewpoints of avatars A1 and A3.
[0022] The server 10 is a device that realizes communication between multiple users via a virtual space VS by providing each user with a three-dimensional virtual space VS shared by the multiple users. As shown in FIG. 1, the server 10 includes an acquisition unit 11, a generation unit 12, a provision unit 13, and a setting unit 14.
[0023] The acquisition unit 11 acquires video data of each user. In this embodiment, the acquisition unit 11 acquires, from the user terminal 20A at the site B1, video data of the user U1 captured by a plurality of cameras C installed at the site B1 (i.e., video data obtained by capturing images of the user U1 from a plurality of different directions). Similarly, the acquisition unit 11 acquires, from the user terminal 20B at the site B2, video data of the user U2 captured by a plurality of cameras C installed at the site B2 (i.e., video data obtained by capturing images of the user U2 from a plurality of different directions). The acquisition unit 11 similarly acquires video data of other users.
[0024] Here, in order to reduce the amount of data transmitted from each of the user terminals 20A and 20B to the server 10, the acquisition unit 11 is configured to be able to selectively acquire only a portion of the video data of each user. The configuration of the acquisition unit 11 will be described below, focusing on the user U1. That is, the process in which the acquisition unit 11 selectively acquires only a portion of the video data of the user U1 from the user terminal 20A in order to reduce the amount of data transmitted from the user terminal 20A to the server 10 will be described.
[0025] The acquisition unit 11 acquires, during a period T1 (a second period), video data of the user U1's entire body obtained by capturing images of the user U1 from multiple different directions (e.g., video data captured by all cameras C installed at the base B1). The period T1 is, for example, a period (e.g., several seconds) after the completion of the login process of the user U1 (e.g., immediately after the user terminal 20A accesses the server 10, a predetermined authentication process is completed, and the user U1 is able to use communication via the virtual space VS provided by the server 10). That is, as an example, the acquisition unit 11 acquires video data of the user U1's entire body in an initial state immediately after the user U1 logs in. The video data of the user U1 acquired during the period T1 is stored in a location accessible by the generation unit 12 (e.g., a memory 1002 or a storage 1003, which will be described later). The video data of the user U1 acquired during the period T1 is used to complement a part (a second part P2, which will be described later) of the avatar A1 corresponding to an arbitrary period T2 (a first period) after the period T1.
[0026] The acquisition unit 11 is configured to acquire, during the period T2, first image data showing a first portion P1 of the user U1's body from among the video data of the user U1's entire body obtained by photographing the user U1 from multiple different directions (e.g., image data captured by all cameras C installed at the location B1), while not acquiring second image data showing a second portion P2 of the user U1's body that is different from the first portion P1. In other words, during the period T2, the acquisition unit 11 is configured to selectively acquire (receive) only the first image data showing the first portion P1 of the user U1's body from the user terminal 20A, among the video data of the user U1's entire body, and not acquire (receive) the second image data showing other portions (the second portion P2) from the user terminal 20A. With this configuration, transmission of the second image data from the user terminal 20A to the server 10 is omitted during the period T2, thereby reducing the amount of data transmitted from the user terminal 20A to the server 10.
[0027] The generation unit 12 generates an avatar to be placed in the virtual space VS in correspondence with each user, based on the video data of each user acquired by the acquisition unit 11.
[0028] When the acquisition unit 11 acquires whole-body video data of the user U1 during the period T2 (for example, image data captured by all cameras C installed at the base B1), the generation unit 12 can generate 3D content (for example, volumetric video image) of the user U1 based on the whole-body video data and apply the 3D content to the avatar A1 of the user U1. That is, the realistic movements of the whole body of the user U1 during the period T2 can be reflected in the avatar A1 placed in the virtual space VS.
[0029] On the other hand, if the acquisition unit 11 acquires the first video data (i.e., video data showing the first part P1 of user U1) during period T2 but does not acquire the second video data (i.e., video data showing the second part P2 of user U1), the generation unit 12 performs the following processing.
[0030] That is, the generation unit 12 generates a first portion P1 of the avatar A1 (first avatar) for the period T2 based on the first video data acquired during the period T2. For example, the generation unit 12 generates partial 3D content in which the second portion P2 is missing based on the first video data acquired during the period T2. That is, for the first portion P1 in which video data (first video data) capturing the actual movement of the user U1 during the period T2 exists, the generation unit 12 can reflect the actual movement of the user U1 by using the video data.
[0031] Meanwhile, the generation unit 12 generates the second portion P2 of the avatar A1 in the period T2 (i.e., the missing portion of the partial 3D content) based on the second video data acquired during the period T1 (second period) prior to the period T2. The generation unit 12 complements the avatar A1 in the period T2 by, for example, attaching a part configured to repeatedly play the video of the second portion P2 of the avatar A1 acquired during the period T1 to the second portion P2 of the avatar A1 in the period T2, or by attaching an image of the second portion P2 at a point in time included in the period T1. This process prevents the avatar A1 in the period T2 from becoming an avatar missing the second portion P2 for which video data was not acquired during the period T2. Note that, since the shape of the avatar A1 is recognized when the generation unit 12 creates the 3D content, if the first portion P1 of the avatar A1 moves, the second portion P2 can be configured to move in accordance with the movement of the first portion P1.
[0032] The providing unit 13 generates and provides, for each user, an image corresponding to the field of view from each user's virtual viewpoint set in the virtual space VS. As described above, for example, the providing unit 13 generates an image corresponding to the field of view from the virtual viewpoint of user U2 (in this embodiment, the first-person viewpoint of avatar A2 corresponding to user U2) as virtual space image IM (see FIG. 2) for user U2 and transmits it to the user terminal 20B. The virtual space image IM transmitted to the user terminal 20B is transmitted to the HMD 30B of user U2 and displayed on the display provided in the HMD 30B. The same processing as above is also performed for users other than user U2.
[0033] The setting unit 14 sets the first part P1 and the second part P2 described above. The setting unit 14 dynamically sets the first part P1 and the second part P2. That is, the setting unit 14 appropriately updates the first part P1 and the second part P2 in response to changes in the situation. The setting unit 14 sets the first part P1 and the second part P2, for example, as follows.
[0034] (First example) Based on the virtual viewpoint of a user (second user) different from user U1 among the multiple users, the setting unit 14 sets a portion of the avatar A1 that is visible to the second user as a first portion P1, and sets a portion of the avatar A1 that is not visible to the second user as a second portion P2. That is, in the first example, the portion of the avatar A1 of user U1 that is visible to other users (i.e., the portion that can promote non-verbal communication between user U1 and other users by reflecting the user U1's realistic movements) is set to the first portion P1 in order to reflect the movements of user U1 in real time. On the other hand, the portion of the avatar A1 of user U1 that is not visible (invisible) to other users is set to the second portion P2 because it is considered that the portion does not contribute much to promoting the non-verbal communication.
[0035] To simplify the explanation of the first example, the explanation will be given assuming that user U3 does not exist in the example of Fig. 2. That is, the processing of the setting unit 14 in the first example will be explained assuming that user U2 is the only second user who views avatar A1. In this case, as shown in Fig. 2, the setting unit 14 sets a part of avatar A1 that is viewable by user U2 (mainly a part including the right half of user U1's body) as the first part P1, and sets a part of avatar A1 that is not viewable by user U2 (mainly a part including the left half of user U1's body, which is the part of avatar A1 opposite to the side where user U2's virtual viewpoint is located) as the second part P2.
[0036] According to the first example, the first portion P1 and the second portion P2 can be appropriately set based on the criterion of whether or not they are visible to other users (i.e., whether or not it is desirable to reflect the user's real movements in order to promote communication between users). That is, by not acquiring video data (second video data) for the second portion P2 of the avatar A1 of user U1 that is not visible to other user U2, the amount of data transmission can be reduced. On the other hand, by acquiring real-time video data (first video data) for the first portion P1 that is visible to other user U2 and reflecting it in the avatar A1, communication between users U1 and U2 can be facilitated.
[0037] (Second example) The setting unit 14 acquires motion information regarding the body motion of the user U1, and based on the motion information, sets a portion of the user U1's body in which motion of a predetermined magnitude or greater is detected as a first portion P1, and sets a portion of the user U1's body in which motion of a predetermined magnitude or greater is not detected as a second portion P2. For example, the portion of the user U1's body in which motion of a predetermined magnitude or greater (or a portion in which motion of a predetermined magnitude or greater is not detected) may be detected by the user terminal 20A based on video data captured by multiple cameras C installed at the base B1. In this case, the setting unit 14 may acquire the detection results from the user terminal 20A to identify the portion of the user U1's body in which motion of a predetermined magnitude or greater is detected (or a portion in which motion of a predetermined magnitude or greater is not detected). Here, "motion of a predetermined magnitude or greater" refers to motion that exceeds a predetermined arbitrary standard (e.g., standard related to the distance of movement, the speed of movement, etc.). For example, motion of a predetermined magnitude or greater may be movement of a predetermined threshold distance or greater within a predetermined threshold period, or movement of a predetermined threshold distance or greater at a speed of a predetermined threshold speed or greater.
[0038] According to the second example, by acquiring video data (first video data) of a first portion P1 of the user U1's body that is moving, the real movement of the user U1 can be reflected in the avatar A1. On the other hand, for a second portion P2 of the user U1's body that is not moving, real-time video data (second video data for period T2) is not acquired, and the avatar A1 is supplemented with past video data (second video data already acquired during period T1), thereby reducing the amount of data transmission.
[0039] In the second example, if a method were adopted in which part A of user U1's body is set to the second part P2 until movement is detected and then switched to the first part P1 upon detection of movement, the following problem could occur. That is, a time lag occurs between when part A, which was set to the second part P2, moves and when part A is set to the first part P1. As a result, the acquisition unit 11 may not be able to acquire video data for the period X from when part A starts moving until it is set to the first part P1, and the generation unit 12 may not be able to reflect the movement of part A during this period X in avatar A1. As a result, when the movement of part A of user U1 is reflected in avatar A1 after the period X has elapsed (i.e., after the video data of part A of user U1 has been acquired), other users may feel that part A of avatar A1 has warped. That is, due to the loss of video data for the period X corresponding to the time lag, the movement of avatar A1 may appear unnatural to other users.
[0040] Therefore, in the second example, the setting unit 14 may set the entire body of the user U1 as the first part P1 as the initial state. Then, the setting unit 14 may change a part of the first part P1 in which no movement of a predetermined amount or more has been detected for a predetermined period of time (e.g., 10 seconds) to the second part P2. Furthermore, when movement of a predetermined amount or more has been detected in the second part P2, the setting unit 14 may change the second part P2 in which the movement has been detected to the first part P1. According to the above configuration, it is possible to avoid the occurrence of the above-mentioned problems and to allow the avatar A1 to move more naturally in the virtual space VS.
[0041] The setting unit 14 notifies the user terminal 20A of setting information indicating the first part P1 and the second part P2 of the user U1. As a result, when the user terminal 20A executes a process of transmitting the video data of the user U1 to the server 10, it can selectively transmit only the video data of the first part P1 (first video data) to the server 10 by referring to the setting information.
[0042] Next, an example of the operation of the virtual space providing system 1 will be described with reference to Fig. 3. Here, attention is focused on the process of providing a virtual space image IM including an avatar A1 generated based on video data of user U1 to another user (user U2). That is, the server 10 also performs a process in which the relationship between user U1 and user U2 is reversed (i.e., a process of generating a virtual space image including an avatar A2 generated based on video data of user U2 and providing it to user U1). However, since such a process is similar to the process described below, a description thereof will be omitted.
[0043] In step S1, the user terminal 20A transmits whole-body video data of the user U1 during a period T1 (second period) (for example, image data captured by all cameras C installed at the location B1 during the period T1) to the server 10. The period T1 is, for example, a certain period (several seconds) immediately after the completion of the login process of the user U1.
[0044] In step S2, the acquisition unit 11 acquires (receives) whole-body video data of the user U1 for the period T1 from the user terminal 20A.
[0045] In step S3, the generation unit 12 generates an avatar A1 for the period T1 based on the video data for the period T1 acquired by the acquisition unit 11. For example, the generation unit 12 generates 3D content (e.g., volumetric video image) of the user U1 based on the video data of the whole body of the user U1 for the period T1, and applies the 3D content to the avatar A1 of the user U1.
[0046] In steps S4 and S5, the providing unit 13 generates a virtual space image IM (see FIG. 2) according to the field of view from the virtual viewpoint of the user U2 set in the virtual space VS, and transmits the virtual space image IM to the user terminal 20B.
[0047] In step S6, the user terminal 20B receives the virtual space image IM from the server 10 and displays the virtual space image IM on the display of the HMD 30B worn on the head of the user U2. Through the above process, an image of the virtual space VS including the avatar A1 that realistically reflects the whole-body movements of the user U1 during the period T1 is provided to the user U2.
[0048] In step S7, the setting unit 14 sets the first part P1 and the second part P2 of the user U1. When executing the process of the first example described above, the setting unit 14 executes the process (steps S21 to S23) shown in the flowchart of Fig. 4. Here, for simplicity of explanation, it is assumed that the only user who can see the avatar A1 is the user U2.
[0049] In step S21, the setting unit 14 acquires information about the virtual viewpoint of the user U2. For example, the setting unit 14 identifies the field of view of the user U2 from the virtual viewpoint of the user U2 (i.e., the area included in the virtual space image IM as shown in FIG. 2). As described above, if the virtual viewpoint (and line of sight) of the user U2 changes depending on the posture of the HMD 30B, the setting unit 14 may identify the field of view of the user U2 based on information about the posture of the HMD 30B. Alternatively, if the positional relationship between the avatars of each user and the virtual line of sight are fixed in the virtual space VS, the field of view of the user U2 may be identified based on setting information about the positional relationship between the avatars and the virtual viewpoint.
[0050] In step S22, the setting unit 14 sets the portion of the avatar A1 that is visible to the user U2 as the first portion P1.
[0051] In step S23, the setting unit 14 sets the part of the avatar A1 that is not visible to the user U2 as the second part P2.
[0052] On the other hand, when executing the processing of the second example described above, the setting unit 14 executes the processing (steps S31 to S35) shown in the flowchart of FIG.
[0053] In step S31, the setting unit 14 sets the entire body of the user U1 as the first part P1 as an initial state.
[0054] In step S32, the setting unit 14 determines whether or not there is a portion in the first portion P1 where a movement equal to or greater than a predetermined value has not been detected for a predetermined period of time.
[0055] If it is determined in step S32 that there is a portion in the first portion P1 where no movement of a predetermined amount or more has been detected for a predetermined period of time (step S32: YES), the setting unit 14 sets that portion as the second portion P2 (step S33). On the other hand, if it is not determined in step S32 that there is a portion in the first portion P1 where no movement of a predetermined amount or more has been detected for a predetermined period of time (step S32: NO), the processing of step S33 is skipped.
[0056] In step S34, the setting unit 14 determines whether or not there is a part in the second part P2 where a movement equal to or greater than a predetermined amount has been detected.
[0057] If it is determined in step S34 that there is a portion of the second portion P2 where a movement equal to or greater than a predetermined value has been detected (step S33: YES), the setting unit 14 sets that portion as the first portion P1 (step S35). On the other hand, if it is not determined in step S34 that there is a portion of the second portion P2 where a movement equal to or greater than a predetermined value has been detected (step S32: NO), the processing of step S35 is skipped.
[0058] The setting information indicating the first part P1 and the second part P2 set in step S7 is notified from the server 10 to the user terminal 20A. After this setting information is notified, the processes of steps S8 to S14 are executed. Note that the process of step S7 and the process of notifying the setting information may be executed periodically. That is, the first part P1 and the second part P2 may change dynamically depending on changes in the situation.
[0059] In step S8, the user terminal 20A transmits to the server 10 the video data (first video data) of the first portion P1 of the user U1 in the period T2 (first period) that comes after the period T1 (second period).
[0060] In step S9, the acquisition unit 11 acquires (receives) the first video data of the first portion P1 of the user U1 in the period T2 from the user terminal 20A.
[0061] In step S10, the generation unit 12 generates a first portion P1 of the avatar A1 for the period T2 based on the first video data acquired during the period T2. That is, the generation unit 12 generates the first portion P1 of the avatar A1 so as to reflect the actual movement of the user U1 during the period T2. For example, the generation unit 12 generates partial 3D content in which the second portion P2 is missing.
[0062] In step S11, the generation unit 12 generates the second portion P2 of the avatar A1 during the period T2 (i.e., the missing portion of the partial 3D content) based on the second video data already acquired during the period T1 prior to the period T2 (in the example of FIG. 3, the data already acquired in step S2). That is, the generation unit 12 complements the second portion P2 of the avatar A1 based on the past video data. As a result, although the second portion P2 does not reflect the actual movements of the user U1 during the period T2, it is possible to generate an avatar A1 with a more natural shape (a shape that gives less discomfort to other users) in which the second portion P2 is not missing.
[0063] The processes of steps S12 and S13 are the same as those of steps S4 and S5. That is, the providing unit 13 generates a virtual space image IM (see FIG. 2) according to the field of view from the virtual viewpoint of the user U2 set in the virtual space VS, and transmits the virtual space image IM to the user terminal 20B.
[0064] The processing of step S14 is the same as step S6. That is, upon receiving the virtual space image IM from the server 10, the user terminal 20B displays the virtual space image IM on the display of the HMD 30B worn on the head of the user U2. Through the above processing, the movements of the first part P1 of the user U1 during the period T2 are realistically reflected for the user U2, while for the second part P2, an image of the virtual space VS including the avatar A1 complemented based on the data of the second part P2 of the user U1 in the past (period T1) is provided.
[0065] According to the server 10 (virtual space provision system 1), by selectively acquiring only a portion of the first video data of the user U1 during the period T2, the amount of video data transmitted for the user U1 can be reduced. As a result, the occurrence of transmission delays, processing slowdowns, and the like caused by an increase in the amount of data transmitted can be suppressed. Furthermore, for the second portion P2 for which video data was not acquired during the period T2, the second portion P2 is supplemented with video data acquired during the period T1, which is earlier than the period T2 (second video data showing the second portion P2). This allows the avatar A1 corresponding to the user U1 during the period T1 to be presented in a manner that is less unnatural to another user U2. As described above, the server 10 (virtual space provision system 1) can facilitate communication between users via the virtual space VS.
[0066] Note that when a portion invisible to the other user U2 is set as the second portion P2 as in the first example of the setting unit 14, it seems unnecessary to complement the second portion P2 of the avatar A1 based on past video data. In other words, if the user U2 cannot view the area corresponding to the second portion P2 in the first place, it seems that there is no problem in leaving the second portion P2 of the avatar A1 missing. However, for example, there is a possibility that the virtual viewpoint of the other user U2 set in the virtual space VS suddenly changes (e.g., the first-person viewpoint of the avatar A2 is switched to a position that allows a bird's-eye view of the virtual space VS). Also, there is a possibility that the orientation of the avatar A1 suddenly changes in conjunction with the user U1 performing an action to change the orientation of his or her body (the first portion P1 of the avatar A1 moves suddenly). In such a case, the second portion P2 of the avatar A1, which was previously invisible to the user U2, may suddenly become visible to the user U2. In this case, if the second portion P2 of the avatar A1 is missing, the missing portion will be visible to the user U2, causing the user U2 to feel uncomfortable, which may result in a problem of impairing the quality of the VR experience of the user U2. Therefore, even when the setting unit 14 executes the processing of the first example, by generating (complementing) the second portion P2 based on past video data, it is possible to avoid the above-mentioned problems and maintain the quality of the VR experience of the user U2.
[0067] The aspects of the virtual space providing device of the present disclosure are not limited to the above-described embodiment. For example, in the first example, when the avatar A1 is visible to a plurality of users U2 and U3, the setting unit 14 sets a portion of the avatar A1 that is visible to at least one of the users U2 and U3 as the first portion P1, and sets a portion of the avatar A1 that is not visible to either of the users U2 and U3 as the second portion P2.
[0068] In the above embodiment, the virtual space providing device is configured only by the server 10, but some functions of the server 10 may be executed by other devices (for example, user terminals at each base). In this case, the virtual space providing device is configured by a system including the server 10 and the user terminals.
[0069] Furthermore, the virtual space providing system 1 does not require an HMD to be worn on the head of each user. For example, instead of an HMD, a regular display device may be placed in front of the user at each location. In this case, although the sense of immersion in the virtual space VS is reduced compared to when an HMD is used, the user can enjoy communication with other users via the virtual space VS by viewing the virtual space image IM displayed on the display device.
[0070] Furthermore, the block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of at least one of hardware and software. Furthermore, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are connected directly or indirectly (for example, using wires, wirelessly, etc.) and these multiple devices. The functional block may also be realized by combining the single device or the multiple devices with software.
[0071] Functions include, but are not limited to, judging, determining, calculating, computing, processing, deriving, investigating, searching, verifying, receiving, transmitting, outputting, accessing, resolving, selecting, choosing, establishing, comparing, assuming, expecting, regarding, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assigning.
[0072] For example, the server 10 according to an embodiment of the present disclosure may function as a computer that performs the virtual space providing method of the present disclosure. Fig. 6 is a diagram illustrating an example of the hardware configuration of the server 10 according to an embodiment of the present disclosure. The server 10 may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc.
[0073] In the following description, the term "device" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the server 10 may be configured to include one or more of the devices shown in Fig. 6, or may be configured to exclude some of the devices.
[0074] Each function of the server 10 is realized by loading predetermined software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations, control communication via the communication device 1004, and control at least one of reading and writing data in the memory 1002 and storage 1003.
[0075] The processor 1001 controls the entire computer by running, for example, an operating system, and may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc.
[0076] Furthermore, the processor 1001 reads programs (program codes), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002, and executes various processes in accordance with these. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, each functional unit of the server 10 (e.g., the acquisition unit 11, etc.) may be implemented by a control program stored in the memory 1002 and running on the processor 1001, and similar implementations may be made for other functional blocks. While the above-described various processes have been described as being executed by one processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may be transmitted from a network via a telecommunications line.
[0077] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be called a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for implementing a virtual space providing method according to one embodiment of the present disclosure.
[0078] Storage 1003 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray disc), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including at least one of memory 1002 and storage 1003.
[0079] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also called, for example, a network device, a network controller, a network card, or a communication module.
[0080] The input device 1005 is an input device (for example, a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (for example, a display, a speaker, an LED lamp, etc.) that outputs to the outside. The input device 1005 and the output device 1006 may be integrated into one device (for example, a touch panel).
[0081] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.
[0082] The server 10 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.
[0083] Although the present embodiment has been described in detail above, it is clear to those skilled in the art that the present embodiment is not limited to the embodiment described in this specification. The present embodiment can be implemented in modified and altered forms without departing from the spirit and scope of the present invention as defined by the claims. Therefore, the description in this specification is intended to be illustrative and does not have any limiting meaning on the present embodiment.
[0084] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.
[0085] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.
[0086] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).
[0087] Each aspect / embodiment described in this disclosure may be used alone, in combination, or switched depending on the implementation. Furthermore, notification of predetermined information (e.g., notification that "X is true") is not limited to being done explicitly, but may be done implicitly (e.g., by not notifying the predetermined information).
[0088] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0089] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.
[0090] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0091] Furthermore, the information, parameters, etc. described in this disclosure may be expressed using absolute values, may be expressed using relative values from a predetermined value, or may be expressed using other corresponding information.
[0092] The names used for the parameters described above are not intended to be limiting in any way. Furthermore, the mathematical formulas and the like that use these parameters may differ from those explicitly disclosed in this disclosure. The various information elements may be identified by any suitable names, and the various names assigned to these various information elements are not intended to be limiting in any way.
[0093] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0094] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.
[0095] When used in this disclosure, the terms "include," "including," and variations thereof are intended to be inclusive, similar to the term "comprising." Furthermore, when used in this disclosure, the term "or" is not intended to be an exclusive or.
[0096] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.
[0097] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different." [Explanation of symbols]
[0098] 1...virtual space providing system, 10...server (virtual space providing device), 11...acquisition unit, 12...generation unit, 13...providing unit, 14...setting unit, 20A, 20B...user terminal, 30A, 30B...HMD, A1...avatar (first avatar), A3...avatar, IM...virtual space image, P1...first part, P2...second part, VS...virtual space.
Claims
1. A virtual space providing device that provides a three-dimensional virtual space shared by a plurality of users to each of the users, an acquisition unit that acquires video data of each of the users; a generation unit that generates an avatar to be placed in the virtual space corresponding to each of the users based on the video data of each of the users; a providing unit that generates and provides, for each of the users, an image corresponding to a field of view from a virtual viewpoint of each of the users that is set in the virtual space; the acquisition unit is configured to acquire first video data in which a first part of a body of the first user is captured, from video data obtained by photographing a first user from a plurality of different directions during a first period, while not acquiring second video data in which a second part of the body of the first user that is different from the first part is captured; When the acquisition unit acquires the first video data but does not acquire the second video data during the first period, the generation unit generating the first portion of a first avatar corresponding to the first user in the first period based on the first video data acquired in the first period; generating the second portion of the first avatar in the first period based on the second video data already acquired in a second period prior to the first period; Virtual space providing device.
2. a setting unit that sets a portion of the first avatar that is visible to the second user as the first portion, and sets a portion of the first avatar that is not visible to the second user as the second portion, based on the virtual viewpoint of a second user who is different from the first user among the plurality of users; The virtual space providing device according to claim 1 .
3. a setting unit that acquires movement information regarding a movement of the body of the first user, and sets a part of the body of the first user in which a movement of a predetermined amount or more is detected as the first part based on the movement information, and sets a part of the body of the first user in which a movement of the predetermined amount or more is not detected as the second part, The virtual space providing device according to claim 1 .
4. The setting unit setting the first user's entire body as the first part as an initial state; A portion of the first portion in which movement equal to or greater than the predetermined value has not been detected for a predetermined period of time is changed to the second portion; When a movement equal to or greater than the predetermined value is detected in the second portion, the second portion in which the movement is detected is changed to the first portion. The virtual space providing device according to claim 3 .
Citation Information
Patent Citations
Video communication system, and video communication method
JP2014056308A
Video communication method, video communication device, and video communication program
JP2020065229A
System and method for augmented reality multi-view telepresence
US20190253667A1