Systems and methods for virtual reality immersive calling
The system addresses inconsistent user positions and lighting in virtual reality communication by processing image streams to align user postures and lighting, ensuring a fully immersive virtual meeting experience.
Patent Information
- Application Number
- JP2024539734
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-30
- Filing Date
- 2022-12-29
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2042-12-29
AI Technical Summary
Inconsistent user positions and lighting conditions during virtual reality communication using headsets or head-mounted displays (HMDs) affect the immersive experience, particularly in scenarios like pandemics where real-time 3D face interactions are crucial.
A system comprising capture devices, networks, and virtual reality devices that capture and process image streams to render a virtual environment with consistent user positions and lighting, adjusting users' postures and lighting conditions for an immersive experience.
The system ensures a fully immersive virtual meeting experience by aligning user positions and lighting conditions, providing a shared virtual environment with realistic interactions.
Smart Images

Figure 0007770577000009 
Figure 0007770577000010 
Figure 0007770577000011
Abstract
Description
[Technical Field]
[0001] The present invention relates to virtual reality, and more particularly to methods and systems for immersive virtual reality communication. [Background technology]
[0002] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of priority from U.S. Provisional Patent Applications Nos. 63 / 295,501 and 63 / 295,505, both filed December 30, 2021, the entireties of which are incorporated herein by reference.
[0003] Given the great advances in virtual or mixed reality, it is becoming practical to use headsets or head-mounted displays (HMDs) to participate in virtual meetings or social gatherings and see each other's 3D faces in real time. In some scenarios, such as pandemics and other disease outbreaks, the need for such gatherings becomes more important as people are unable to meet in person.
[0004] However, the images of various users used in a virtual environment are often taken with different devices at different positions and angles. These inconsistent user positions / postures and lighting conditions have a significant impact on participants' ability to have a fully immersive virtual meeting experience. Summary of the Invention
[0005] According to one embodiment, a system is provided for immersive virtual reality communication, the system including a first capture device configured to capture an image stream of a first user, a first network configured to transmit the captured image stream of the first user, a second network configured to receive data based at least in part on the captured image stream of the first user, a first virtual reality device used by a second user, a second capture device configured to capture the image stream of the second user, and a second virtual reality device used by the first user, wherein the first virtual reality device is configured to render a virtual environment and generate a rendition of the first user based at least in part on the data based at least in part on the image stream of the first user generated by the first capture device, and the second virtual reality device is configured to render a virtual environment and generate a rendition of the second user based at least in part on data based at least in part on the captured image stream of the second user generated by the second capture device.
[0006] In some embodiments, the virtual environment is substantially common between the first virtual reality device and the second virtual reality device, and a viewpoint of the first virtual reality device is different from a viewpoint of the second virtual reality device. In other embodiments, the virtual environment may provide a common feel, but may be selectively configured based on viewpoints of individual users. In a further embodiment, the system further includes instructing the first user and the second user to move and rotate to optimize the position of the first user and the position of the second user relative to the first capture device and the second capture device, respectively, based on a desired rendering environment before renditions of the first user and the second user are generated via user interfaces at the second virtual reality device and the second virtual reality device, respectively.
[0007] According to yet another embodiment, the first network includes at least one graphics processing unit, and the data based at least in part on the captured image stream of the first user is generated entirely in the graphics processing unit before being transmitted to the second network. [Brief explanation of the drawings]
[0008] These and other objects, features, and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure when considered in conjunction with the accompanying drawings and the appended claims. [Figure 1] FIG. 1 is a diagram illustrating a virtual reality capture and display system. [Figure 2] FIG. 2 is a diagram illustrating an embodiment of a system in which two users are in two respective user environments according to the first embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of a virtual reality environment rendered to a user. [Figure 4] FIG. 4 is a diagram illustrating an example of a second virtual perspective 400 of the second user 270 of FIG. [Figure 5A] , [Figure 5B] 5A and 5B illustrate an example of an immersive call in a virtual environment from the perspective of a user's starting position. [Figure 6] FIG. 6 is a flow chart illustrating a call initiation flow that places the first and second users of the system in the proper positions to implement the desired call characteristics. [Figure 7] FIG. 7 is a diagram showing an example of directing a user to a desired location. [Figure 8] FIG. 8 is a diagram showing an example of detecting a person via a capture device and estimating a skeleton as three-dimensional points. [Figure 9A] , [Figure 9B] , [Figure 9C] , [Figure 9D] , [Figure 9E] , [Figure 9F] 9A-9E show exemplary user interfaces for user-recommended actions of moving right, moving left, moving backward, moving forward, turning left, and turning right, respectively. [Figure 10] FIG. 10 is a diagram illustrating the overall flow of an immersive virtual call according to one embodiment of the present invention. [Figure 11] FIG. 11 is a diagram illustrating an example of a system for a virtual reality immersive communication system. [Figure 12A] , [Figure 12B] , [Figure 12C] , [Figure 12D] 12A-D show a user workflow for adjusting to a sitting or standing position in an immersive calling system. [Figure 13A] , [Figure 13B] , [Figure 13C] , [Figure 13D] , [Figure 13E] 13A-E show various embodiments for boundary setting in an immersive communication system. [Figure 14] , [Figure 15] , [Figure 16] , [Figure 17] , [Figure 18] 14, 15, 16, 17 and 18 show various scenarios for user interaction in an immersive call system. [Figure 19] FIG. 19 is a diagram showing an example of converting an image displayed by a game engine in a GPU. [Figure 20] FIG. 20 shows a wireless version of FIG. [Figure 21] FIG. 21 shows a version of FIG. 19 that uses a stereo camera for capture. [Figure 22]FIG. 22 shows an exemplary workflow for region-based object relighting using the Lab color space. [Figure 23] FIG. 23 illustrates a region-based method for object or environment relighting using the Lab color space, using the human face as an example. [Figure 24] FIG. 24 shows a user wearing a VR headset. [Figure 25A] , [Figure 25B] 25A and 25B show an example of 468 facial feature points extracted from both the input image and the target image. [Figure 26] FIG. 26 shows a region-based method for relighting an object or environment using the covariance matrix of the RGB channels.
[0009] Throughout the figures, the same reference numerals and characters, unless otherwise stated, are used to denote like features, elements, components, or portions of the illustrated embodiments. Moreover, while the present disclosure will be described in detail with reference to the figures, it is done so in connection with illustrative exemplary embodiments. It is intended that changes and modifications can be made to the exemplary embodiments described without departing from the true scope and spirit of the subject disclosure as defined by the appended claims. DETAILED DESCRIPTION OF THE INVENTION
[0010] Exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that the following exemplary embodiments are merely examples for implementing the present disclosure and can be appropriately modified or altered depending on the individual configuration of the device to which the present disclosure is applied and various conditions. Therefore, the present disclosure is not limited to the following exemplary embodiments. According to the figures and embodiments described below, the described embodiments can be applied / implemented in situations other than those described below as examples. Furthermore, when multiple embodiments are described, unless expressly specified otherwise, the respective embodiments can be combined with each other. This includes the ability to substitute various steps and functions between embodiments as deemed appropriate by those skilled in the art. FIG. 1 illustrates a virtual reality capture and display system 100. The virtual reality capture system includes a capture device 110. The capture device may be, for example, a camera with a sensor and optics designed to capture 2D RGB images or video. Some embodiments use specialized optics, such as a binocular camera or a light field camera, to capture multiple images from different viewpoints. Some embodiments include one or more such cameras. In some embodiments, the capture device may include a range sensor that effectively captures RGBD (red, green, blue, depth) images directly or through software / firmware fusion of multiple sensors, such as an RGB sensor and a range sensor (e.g., a lidar system or a point cloud-based depth sensor). The capture device may be connected to local or remote (e.g., cloud-based) systems 150 and 140 (hereinafter referred to as server 140), respectively, via a network 160. Capture device 110 is configured to communicate with server 140 via network connection 160, and the capture device transmits a series of images (e.g., a video stream) to server 140 for further processing. Also shown in FIG. 1 is a user of system 120. In an exemplary embodiment, the user is wearing a virtual reality (VR) device 130 configured to transmit stereo video to the left and right eyes of user 120.As an example, the VR device may be a headset worn by a user. Other examples may include a stereoscopic display panel or any display device that enables implementation of embodiments described in this disclosure. The VR device is configured to receive incoming data from server 140 via second network 170. In some embodiments, network 170 may be the same physical network as network 160, but data transmitted from capture device 110 to server 140 may differ from data transmitted between server 140 and VR device 130. Some embodiments of the system do not include VR device 130, as described below. The system may also include microphone 180 and speaker / headphone device 190. In some embodiments, the microphone and speaker device are part of VR device 130.
[0011] 2 illustrates an embodiment of a system 200 with two users 220 and 270 in two separate user environments 205 and 255. In this exemplary embodiment, each of the users 220 and 270 is equipped with a separate capture device 210 and 260, a separate VR device 230 and 280, and is connected to a server 250 via a separate network 240 and 290. In some examples, only one user may have a capture device 210 or 260, while the other user may have only a VR device. In this case, one user environment may be considered a transmitter and the other user environment may be considered a receiver with respect to video capture. However, in embodiments where the roles of the transmitter and receiver are different, audio content may be transmitted and received by only or both the transmitter and receiver, or the roles may be reversed.
[0012] FIG. 3 illustrates a virtual reality environment 300 rendered to a user. The environment includes a computer graphics model 320 of a virtual world with a computer graphics projection of a captured user 310. For example, user 220 of FIG. 2 can view virtual world 320 and rendition 310 of second user 270 of FIG. 2 through a separate VR device 230. In this example, capture device 260 captures images of user 270, processes them on server 250, and renders them in virtual reality environment 300. In the example of FIG. 3, user rendition 310 of user 270 of FIG. 2 shows the user without a separate VR device 280. Some embodiments show the user with a VR device 280. In other embodiments, user 270 does not use a wearable VR device 280. Furthermore, in some embodiments, the captured image of user 270 captures a wearable VR device, but processing of the user image removes the wearable VR device and replaces it with a portrait of the user's face.
[0013] Additionally, adding the user rendition 310 to the virtual reality environment 300 along with the VR content 320 may include a lighting adjustment step to adjust the lighting of the captured and depicted user 310 to better match the VR content 320.
[0014] In this disclosure, first user 220 of Figure 2 is shown VR rendition 300 of Figure 3 via a separate VR device 230. Thus, first user 220 sees user 270 and virtual environment content 320. Similarly, in some embodiments, second user 270 of Figure 2 is in the same VR environment 320 but from a different perspective, such as the perspective of a virtual character rendition of 310.
[0015] Figure 4 shows a second virtual perspective 400 of the second user 270 of Figure 2. The second virtual perspective 400 is shown on the virtual device 230 of the first user 220 of Figure 2. The second virtual perspective may be based on the same virtual content 320 of Figure 3, but includes virtual content 420 from the perspective of a virtual rendition of the character 310 of Figure 3 that represents the point of view of the user 220 of Figure 2. The second virtual perspective may also include a virtual rendition of the second user 270 of Figure 2.
[0016] FIG. 6 illustrates a call initiation flow that positions a first and second user of a system to implement desired call characteristics, such as the two examples shown in FIGS. 5A and 5B. The flow begins in block B610 with a first user initiating an immersive call to a second user. The call can be initiated through an application on the user's VR device or through other intermediate devices, such as the user's local computer, mobile phone, or voice assistant (e.g., Alexa, Google Assistant, Siri, etc.). The call initiation executes instructions that notify a server, such as 140 in FIG. 1 or 250 in FIG. 2, that the user intends to make an immersive call with the second user. The first user can be selected, via an app, from a list of contacts, for example, that are known to have immersive calling capabilities. The server responds to the call initiation in block B620 by notifying the second user that the first user intends to initiate an immersive call with the second user. By way of just a few examples, an application on the second user's local device, such as the user's VR device, mobile phone, computer, or voice assistant, provides a final notification to the second user giving the second user an opportunity to accept the call. If the call is not accepted in block B630, either by the second user's selection or via a timeout period while waiting for an answer, flow proceeds to block B640, where the call is rejected. In the call, the first user may be notified that the second user did not actively or passively accept the call. Other embodiments include when the second user is detected to be in do not disturb mode or engaged in another active call. In these cases, the call may not be accepted. If the call is accepted in block B630, flow proceeds to block B650 for the first user and to block B670 for the second user. At blocks B650 and B670, each user is notified that they should wear their respective VR devices. At this point, the system begins video streams from the first and second users' respective image capture devices.The video stream is processed via a server to detect the presence of the first and second users and determine their positions in the captured image. Blocks B660 and B680 then provide cues via the VR device application for the first and second users, respectively, to move to appropriate positions for an effective immersive conversation.
[0017] Thus, the collective effect of the system is to present a virtual world comprising 300 of FIG. 3 and 400 of FIG. 4 that presents the illusion of a meeting of a first user and a second user in a shared virtual environment.
[0018] 5A and 5B show two examples of immersive conversations in a virtual environment from the perspective of user starting positions. For example, in some instances shown in FIG. 5A, user renditions 510 and 520 are positioned side-by-side. Side-by-side placement may be preferable if both users intend to watch VR content together. This may be the case, for example, when watching a live event or video or other content. In other examples, as shown in FIG. 5B, a first user's rendition 560 and a second user's rendition 570 (representing users 200 and 270, respectively, in FIG. 2) are positioned in the virtual environment so that they are face-to-face. For example, if the intention of the immersive experience is for two users to meet, they may wish to enter the environment face-to-face.
[0019] FIG. 7 illustrates an exemplary embodiment for directing a user to a desired location. The flow can be used for both a first and a second user. The flow begins in block B710, where an image capture device providing video frames to a server is analyzed to determine whether a person is present in the captured image. One such embodiment performs face detection to determine whether a face is present in the image. Another embodiment uses a full-body detector. Such a detector can detect the presence of a person and estimate a "human skeleton" that can provide some estimate of the detected person's pose.
[0020] Block B720 determines whether a person is detected. Some embodiments may include a detector capable of detecting multiple people, although for purposes of an immersive call, only one person is of interest. In the case of multiple person detection, some embodiments alert the user that there are multiple detections and ask the user to orient other people outside the camera's field of view. In other embodiments, the centermost detected person is used, and in still other embodiments, the largest detected person may be selected. It should be understood that other detection techniques may be used. If block B720 determines that a person is not detected, flow proceeds to block B725, and in some embodiments, the user is presented with streaming video from the camera in the VR device headset alongside captured video, if available, taken from the VR device headset. In this way, the user can see both their camera from their perspective and the scene being captured from the capture device's perspective. These images may be presented side-by-side or picture-in-picture, for example. Flow then proceeds back to block B710, and detection is repeated. If block B720 determines that a person is detected, flow proceeds to block B730.
[0021] In block B730, the VR device's boundaries are obtained (if available) for the current user position. Some VR devices provide guardian boundaries to prevent the user from colliding with other real-world objects while wearing the headset and immersed in the virtual world, and can detect when the user moves near or outside the virtual boundary. VR boundaries are described in more detail, for example, with respect to Figures 13A-13E. Flow then proceeds to block B740.
[0022] Block B740 determines the user's pose relative to the capture device. For example, in one embodiment shown in FIG. 8, a person 830 is detected via the capture device 810 and a skeleton 840 is estimated as a 3D point. The user's pose relative to the capture device may be the pose of the user's shoulders relative to the capture device. For example, if left and right shoulder points 860 and 870 are estimated in 3D, a unit vector n 890 may be determined to emanate from the midpoint of the two shoulder points, be perpendicular to the line 880 connecting the two shoulder points, and be parallel to the horizontal x-axis and the depth z (optical) axis 820 of the capture device. In this embodiment, the dot product of vector n 890 with the negative capture device axis -z axis generates the cosine of the user's pose relative to the capture device. If n and z are both unit vectors, a dot product close to 1 indicates the user is positioned with their shoulders facing the camera, which is ideal for capture for a face-to-face scenario such as that shown in FIG. 5B. However, if the dot product is close to zero, it indicates the user is facing to the side, which is ideal for the scenario shown in FIG. 5A. Additionally, in a side-by-side scenario, one user may be captured from the right side and should be positioned to the left of the other user, who should be captured from the left side. To determine whether the user is oriented so that their left or right side is captured, the depths of the two shoulder points 860 and 870 may be compared to determine which shoulder is closer in depth to the capture device. In some embodiments, other joints, such as the hips or eyeballs, are used as reference joints in a similar manner.
[0023] Additionally, the detected skeleton in FIG. 8 may be used to determine the size of the detected person based on estimated joint lengths. Through calibration, joint lengths can be used to estimate a person's standing size, even if they are not fully upright. This allows the detection system to determine the physical height of a user's bounding box if the user's height is known approximately a priori. Other reference lengths can also be used to estimate a user's height; for example, headset size is known for a given device and varies little across devices. Thus, a user's height can be estimated based on reference lengths, such as the size of a headset, when they appear together in a captured frame. Some embodiments ask the user for their height when they create a contact profile so that they can be appropriately scaled when depicted in a virtual environment.
[0024] In some embodiments, the virtual environment is a real indoor / outdoor environment captured, for example, via 3D scanning and photogrammetry. Thus, a virtual camera corresponds to the real 3D physical world, whose dimensions and sizes are known, and through which that world is rendered to the user, allowing the system to place renditions of people at various locations within the environment, independent of the virtual camera's position within the environment. Therefore, to provide a realistic interactive experience, a program is required to accurately project the view of the person captured by the real camera onto the desired location and pose within the environment. This can be done by creating a person-centered coordinate frame based on skeletal joints, and the system derives a reprojection matrix.
[0025] Some embodiments display a rendition of a user on a 2D projection screen (flat or curved) rendered in a 3D virtual environment (possibly stereoscopically or via a light-field display device). Note that in these cases, if the viewing angle differs significantly from the capture angle, the projected figure no longer appears realistic; in the extreme case where the projection screen is parallel to the optical axis of the virtual camera, the user simply sees a line representing the projection screen. However, due to the flexibility of the visual system, for a moderate range of difference between the capture angle in the physical world and the virtual viewing angle in the virtual world, a second user can see the projected figure as a 3D figure, positioned nearby. This means that both communicating users can perform a limited range of motion without destroying the other's 3D perception. This range can be quantified, and this information can be used to guide the design of various embodiments for positioning users relative to their respective capture devices.
[0026] In some embodiments, the user's rendition is presented as a 3D mesh instead of a planar projection. Such embodiments may allow for greater flexibility in the user's range of movement, further impacting positioning purposes.
[0027] Returning to FIG. 7, once block B740 determines the user's posture relative to the capture device, flow continues to block B750.
[0028] Block B750 determines the size and position of the user within the capture device frame. Some embodiments prefer to capture the user's entire body and determine whether the entire body is visible. Additionally, in some embodiments, an estimated bounding box of the user may be determined, such that the center, height, and width of the box in the capture frame are determined. Flow then proceeds to block B760.
[0029] In block B760, an optimal position is determined. First, the estimated user's pose relative to the capture device is compared to the desired pose given the desired scenario. Second, the user's bounding box is compared to the user's ideal bounding box. For example, some embodiments determine that the estimated user bounding box should not extend beyond the capture frame to capture the entire body, and that there is sufficient margin above and below the top and bottom of the box to allow the user to move without risking moving outside the capture device area. Third, to optimize movement margin, the user's position should be determined (e.g., the center of the bounding box) and compared to the center of the capture area. Some embodiments also check the VR boundary to ensure that the user's current placement relative to the VR boundary provides sufficient margin for movement.
[0030] Some embodiments include a position score S based at least in part on one or more of an orientation pose score p, a position score x, a size score s, and a bounds score b.
[0031] The pose score may be based on the dot product of vector n 890 and vector z 820 in Figure 8. Additionally, the detected z positions 860 and 870 of the left and right shoulders, respectively, are used to transform the pose angle θ. JPEG0007770577000001.jpg25146
[0032] Therefore, one pose score can be expressed as: JPEG0007770577000002.jpg1560
[0033] Here, the above norm must take into account the periodicity of θ. For example, one embodiment uses ||θ-θ desired || is defined as follows: JPEG0007770577000003.jpg10137
[0034] The position score can measure the position of the detected person within the capture device frame. An example embodiment of the position score is based at least in part on the captured person bounding box center c and the capture frame width W and height H: JPEG0007770577000004.jpg2384
[0035] The boundary score b provides a score for the user's position within the VR device boundary. In this case, given a user position (u,v) on the ground plane, the position (0,0) provides the position on the ground plane that can be used to construct a circle of maximum radius circumscribing the defined boundary. In this embodiment, the boundary score may be given as: JPEG0007770577000005.jpg1548
[0036] A total score for assessing the user pose and position can then be given as objective J: JPEG0007770577000006.jpg16150where, λ p , λ x , λ s and λ b is a weighting factor that gives the relative weight to each score, and f is a score shaping function with parameter Γ that describes the monotonic shaping function of the scores. As an example, JPEG0007770577000007.jpg1529Here, Γ b is a positive number.
[0037] Flow then moves to block B770, where it is determined whether the user's position and pose are acceptable. If not, flow continues to block B780, where visual cues are provided to the user to assist in moving to a better position. Exemplary UIs for user-recommended actions of move right, move left, move backward, move forward, turn left, and turn right, respectively, are shown in Figures 9A, 9B, 9C, 9D, 9E, and 9F. A combination of these may also be capable of indicating the flow for the user.
[0038] 7, once the user has been provided with a visual cue, flow returns to block B710 and the process repeats. If block B770 ultimately determines that the pose and position are acceptable, flow moves to block B790 and the process ends. In some embodiments, the process continues for the duration of the immersive call, and if block B770 determines that the position is acceptable, flow skips block B780, which provides the positioning cues, and returns to block B710.
[0039] The overall flow of an immersive call embodiment is shown in Figure 10. The flow begins in block B1010, where a call is initiated by a first user based on a selected contact as the second user and a selected VR scenario. Next, the flow proceeds to block B1020, where the second user accepts the call, or the call is not accepted, in which case the flow ends. If the call is accepted, the flow continues to block B1030, where the first and second users are prompted to put on their VR devices (headsets). Next, the flow proceeds to block B1040, where the users are individually oriented to their appropriate positions and postures based on their positions relative to their respective capture devices and the selected scenario. Once the users are in an acceptable position, the flow continues to B1050, where the VR scenario begins. During the VR scenario, the user will have the option to end the call at any time. In block B1060, the call is ended, and the flow ends.
[0040] 11 illustrates an exemplary embodiment of a system for a virtual reality immersive communication system. System 11 includes two user environment systems 1100 and 1110, which are specially configured computing devices, two respective virtual reality devices 1104 and 1114, and two respective image capture devices 1105 and 1115. In this embodiment, the two user environment systems 1100 and 1110 communicate over one or more networks 1120, which may include wired networks, wireless networks, local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), and personal area networks (PANs). In some embodiments, the devices also communicate over other wired or wireless channels.
[0041] The two user environment systems 1100 and 1110 include one or more respective processors 1101 and 1111, one or more respective I / O components 1102 and 1112, and respective storage 1103 and 1113. The hardware components of the two user environment systems 1100 and 1110 also communicate via one or more buses or other electrical connections. Examples of buses include a Universal Serial Bus (USB), an IEEE 1394 bus, a PCI bus, an Accelerated Graphics Port (AGP) bus, a Serial AT Attachment (SATA) bus, and a Small Computer System Interface (SCSI) bus.
[0042] The one or more processors 1101 and 1111 include one or more central processing units (CPUs), which may include one or more microprocessors (e.g., single-core microprocessors, multi-core microprocessors), one or more graphics processing units (GPUs), one or more tensor processing units (TPUs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more digital signal processors (DSPs), or other electronic circuits (e.g., other integrated circuits). The I / O components 1102 and 1112 include communication components (e.g., graphics cards, network interface controllers) that communicate with the respective virtual reality devices 1104 and 1114, the respective capture devices 1105 and 1115, the network 1120, and other input or output devices (not shown), which may include a keyboard, mouse, printing device, touchscreen, light pen, optical storage device, scanner, microphone, drives, and game controllers (e.g., joystick, gamepad).
[0043] The storages 1103 and 1113 include one or more computer-readable storage media. As used herein, a computer-readable storage medium includes, for example, magnetic disks (e.g., floppy disks, hard disks), optical disks (e.g., CDs, DVDs, Blu-rays), magneto-optical disks, magnetic tapes, and manufactured products such as semiconductor memory (e.g., non-volatile memory cards, flash memory, solid-state drives, SRAM, DRAM, EPROM, EEPROM). The storages 1103 and 1113, which may include both ROM and RAM, can store computer-readable data or computer-executable instructions. The two user environment systems 1100 and 1110 also include communication modules 1103A and 1113A, capture modules 1103B and 1113B, rendering modules 1103C and 1113C, positioning modules 1103D and 1113D, and user performance modules 1103E and 1113E. A module comprises logic, computer-readable data, or computer-executable instructions. In the embodiment shown in FIG. 11 , the module is implemented in software (e.g., Assembly, C, C++, C#, Java, BASIC, Perl, Visual Basic, Python, Swift). However, in some embodiments, the module is implemented in hardware (e.g., customized circuitry) or, alternatively, a combination of software and hardware. When a module is implemented at least partially in software, the software may be stored in storages 1103 and 1113. Also, in some embodiments, the two user environment systems 1100 and 1110 include additional or fewer modules, modules are combined into fewer modules, or modules are divided into more modules. One environment system may be similar to the other environment system or may differ in terms of the inclusion or organization of modules.
[0044] Respective capture modules 1103B and 1113B include programmed operations to perform captures such as those shown in 110 of Figure 1, 210 and 260 of Figure 2, and 810 of Figure 8, and used in block B710 of Figure 7 and block B1040 of Figure 10. Respective rendering modules 1103C and 1113C include programmed operations to perform functions described in, for example, blocks B660 and B680 of Figure 6, block B780 of Figure 7, block B1050 of Figure 10, and examples of Figures 9A-9F. Respective positioning modules 1103D and 1113D include programmed operations to perform processes described in Figures 5A and 5B, B660 and B680 of Figure 6, Figures 7, 8, and 9. Each user rendering module 1103E and 1113E includes operations programmed to perform user renderings such as those shown in Figures 3, 4, 5A, and 5B.
[0045] In another embodiment, user environment systems 1100 and 1110 are embedded in VR devices 1104 and 1114, respectively. In some embodiments, modules are stored and executed on an intermediate system, such as a cloud server.
[0046] 12A-D illustrate a user workflow for arranging a seated or standing meeting between user A and user B, as described in blocks B660 and B680 of FIG. 6. In FIG. 12A, user A, who is seated, talks to user B, and the system prompts user B, who is standing, to put on a headset and sit in a designated area. As a result, FIG. 12B illustrates a virtual meeting taking place in a seated position. In another scenario described in FIG. 12C, user A, who is standing, talks to user B. The system prompts user B, who is seated, to put on a headset and sit in a designated area. As a result, FIG. 12D illustrates a virtual meeting taking place in a standing position.
[0047] Next, Figures 13, 14, 15, 16, 17 and 18 illustrate various scenarios for user interaction in the immersive call system described herein.
[0048] 13A shows an example configuration of room-scale boundaries in a skybox setting in a VR environment. In this example, the chairs of users A and B face the same direction; FIG. 13B shows an example configuration of room-scale boundaries in a park setting in a virtual environment. In this example, the chairs of users A and B face each other, and the cameras are positioned opposite each other.
[0049] 13C shows another example of setting VR boundaries, where stationary boundaries for user A and user B and corresponding room-scale boundaries are shown.
[0050] Figure 13D shows another example configuration of a room-scale boundary in a beach setting in a VR environment. In this example, the chairs of users A and B face the same direction; Figure 13E shows an example configuration of a room-scale boundary in a train setting in a VR environment. In this example, the chairs of users A and B face each other, and the cameras are positioned opposite each other.
[0051] 14 and 15 show various embodiments of a VR setup where the user's side profile is captured by a capture device.
[0052] Figure 16 shows an example of performing an immersive VR activity (e.g., table tennis) where a first user faces another user and a camera is positioned directly in front to capture the user's front view. This allows the user's gaze to be fixed on the other user in the VR environment. Figure 18 shows an example of two users standing in a sports venue, with each user's profile captured by their capture device.
[0053] To capture a proper image of the user, the user is prompted to move to an appropriate position and posture. Figure 17 shows an example in which the user puts on the HMD and the HMD provides instructions to the user to move their physical chair to a specified position and in a direction opposite to that specified so that a proper image of the user can be captured.
[0054] Below, we describe an embodiment for capturing an image from a camera, applying arbitrary code on a GPU to transform the image, and sending the transformed image to a game engine for display without leaving the GPU. An example is described in more detail in connection with FIG. 19.
[0055] This capture method has the following advantages: a single copy function from CPU memory to GPU memory; all operations are performed on the GPU; the GPU's high parallel processing capabilities allow images to be processed much faster than using the CPU; textures are shared with the game engine without leaving the GPU, making data transmission to the game engine more efficient; and reducing the time between capturing and displaying an image in the game engine application.
[0056] In the example shown in Figure 19, when a camera is connected to an application, the camera transfers frames of a video stream captured by the camera to the application to facilitate displaying the frames via a game engine. In this example, the video data is transferred to the application via an audio / video interface, such as an HDMI-USB capture card, capable of capturing uncompressed video frames at high resolution with low latency. In another embodiment, as shown in Figure 20, the camera wirelessly transmits the video stream to a computer, where it is decoded.
[0057] Next, the system acquires frames in the camera provided in the native format. In this step, the system acquires frames in the native format provided by the camera. In this embodiment, for the purpose of explanation only, the native format is the YUV format. The use of the YUV format is not considered limiting, and any native format that allows the implementation of this embodiment is applicable.
[0058] The data is then loaded into the GPU, and the YUV-encoded frames are loaded into GPU memory, allowing highly parallel operations to be performed on the YUV-encoded frames. Once the images are loaded into the GPU, they are converted from YUV format to RGB to allow for additional downstream processing. A mapping function is then applied to remove image distortions caused by the camera lens. Deep learning techniques are then employed to separate the subject from the background to remove the subject's background. To send the images to the game engine, GPU texture sharing is used to write textures to memory that the game engine reads. This process prevents data from being copied from the CPU to the GPU. The game engine receives the textures from the GPU and uses them to display them to users on various devices. Any game engine capable of implementing this embodiment is applicable.
[0059] In another embodiment, a stereoscopic camera is used and lens correction is performed on each half of the image, as shown in Figure 21. The stereoscopic camera displays the image captured from the left lens only to the left eye and the image captured from the right lens only to the right eye, thereby providing the user with a 3D effect of the image. This can be achieved by using a VR headset. Any VR headset that allows the implementation of this embodiment is applicable.
[0060] Relighting of objects or environments can be very important in augmented VR. Generally, images of the virtual environment and images of different users are captured at different times and in different locations. These differences in location and time make it impossible to maintain perfectly identical lighting conditions between the user and the environment.
[0061] Under different lighting conditions, the appearance of images captured from a subject changes. Humans can use this difference in appearance to extract the lighting conditions of the environment. When different subjects are captured under different lighting conditions and then composited into VR without any processing, users will notice some inconsistencies in the extracted lighting conditions from the different subjects, which will cause an unnatural perception of the VR environment.
[0062] In addition to lighting conditions, the cameras used to capture images for different users, as well as the virtual environments, are often different. Each camera has its own hardware-dependent nonlinear color correction function. Different cameras will have different color correction functions. This difference in color correction can lead to different perceived lighting appearances for different subjects, even in the same lighting environment.
[0063] Given all of these lighting and camera differences and variations, it is important in augmented VR to relight the RAW capture images of different subjects so that the lighting information provided by the subjects matches each other.
[0064] 22 shows a workflow diagram for implementing a region-based object relighting technique using the Lab color space, according to an exemplary embodiment. The Lab color space is provided as an example, and any color space that enables the implementation of this embodiment is applicable.
[0065] Given an input image 2201 and a target image 2202, first, feature extraction algorithms 2203 and 2204 are applied to identify feature points in the target image and the input image. Next, in step 2205, a shared reference region is determined based on the feature extraction. After that, the shared region is transformed from RGB color space to Lab color space (e.g., CIE Lab color space) in steps 2206, 2207, and 2208, respectively, for both the input image and the target image. The Lab information obtained from the shared region will be used to determine a transformation matrix in steps 2209 and 2210. Then, this transformation matrix is applied to the whole or a specific region of the input image to adjust the Lab components in step 2211, and outputs the final relighting of the input image after being converted to RGB color space in step 2212.
[0066] To explain each step in detail, an example is provided in connection with FIG.
[0067] Figure 23 shows the workflow of the process of this embodiment for relighting an input image based on illumination and color information from a target image. The input image is shown in A1, and the target image is shown in B1. The objective of this invention is to make the illumination of the face in the input image closer to the illumination of the target image. First, a face is extracted using a face detection application. Any face detection application that allows the implementation of this embodiment can be applied.
[0068] In this example, the entire face from the two images was not used as a reference for relighting. In a VR environment, the user typically wears a head-mounted display (HMD), as shown in FIG. 24 , so the entire face was not used. As shown in FIG. 24 , when the user wears the HMD, the HMD typically blocks the entire upper half of the user's face, making only the lower half of the user's face visible to the camera. Another reason the entire face was not used is that the content of the two faces may differ even if the user is not wearing an HMD. For example, the faces in FIGS. 23(A1) and 23(A2) have open mouths, while the faces in FIGS. 23(B1) and 23(B2) have closed mouths. The difference in the mouth region between these images may result in inaccurate adjustments to the input image if the entire face region should be used. Also, not using the entire face provides flexibility with respect to relighting the subject. The region-based approach allows for different control to be provided for different regions of the subject.
[0069] As shown in A2 and B2 of Figure 23, a common region was selected from the lower right region of the face, shown as a rectangle in A2 (2310) and B2 (2320), and replotted as A3 and B3. The selected region served as a reference region used to relight the input image. However, any region of the image can be selected that allows for the implementation of this embodiment.
[0070] While the above description describes manual selection of specific regions for the input and target images, in another exemplary embodiment, the selection can be determined automatically based on feature points detected from the face. An example is shown in FIGS. 25A and 25B. Application of a facial feature identifier application results in the identification of 468 facial feature points for both images. Any facial feature identifier application that enables implementation of the present embodiment can be applied. These feature points serve as guidelines for selecting shared regions for relighting. Any facial region can then be selected. For example, the entire lower face region can be selected as the shared region, as indicated by the boundary lines A and B.
[0071] After the shared region is obtained, relighting of the face in the input image is performed. The first step in this process is to convert the existing RGB color space. There are many color spaces available for representing colors, but the RGB color space is the most typical color space. However, the RGB color space is device-dependent, and different devices produce different colors. Therefore, it is not ideal to serve as a framework for color and lighting adjustment, and conversion to a device-independent color space will provide better results.
[0072] As mentioned above, this embodiment uses the CIELAB or Lab color space. It is device-independent and is calculated from the XYZ color space by normalizing to the white point. "The CIELAB color space represents any color using three values: L*, a*, and b*. L* indicates the perceived lightness, and a* and b* represent the four colors specific to human vision" (https: / / en.wikipedia.org / wiki / CIELAB_color_space).
[0073] The Lab components from the RGB color space can be obtained, for example, by open-source computer vision color conversion. After the Lab components of the shared reference region in both the input image and the target image, their mean and standard deviation are calculated. Some embodiments use other measures of centrality and variation other than the mean and standard deviation. For example, the median and median absolute deviation can be used to robustly estimate these measures. Of course, other means are possible, and this description is not limited to these. Next, the Lab components of all or some specific selected regions of the input image are adjusted by executing the following formula: JPEG0007770577000008.jpg21161Here, x is one of the three components L^*, a^*, or b^* in CIELAB space.
[0074] In another exemplary embodiment that is more data-driven, the covariance matrix of the RGB channels is used. Whitening the covariance matrix allows for separation of the RGB channels, similar to what is done using a Lab color space such as the CIELab color space. Detailed steps are shown in Figure 26.
[0075] In Figure 26, a feature extraction algorithm (2630 and 2640) is used to identify key feature points in both the input image and the target image. A shared reference region is then determined based on the locations of the key points in both images. Figure 26 differs from Figure 22 in that the RGB channels are not transformed to the lab color space. Instead, the covariance matrices of the shared region in both the input image and the target image are calculated directly from the RGB channels. Single value decomposition (SVD) is then applied to obtain a transformation matrix from these two covariance matrices. The transformation matrix is applied to the entire input image to obtain a corresponding relighting matrix that is ultimately used and applied to correct the image displayed to the user.
[0076] At least some of the above-described devices, systems, and methods can be implemented, at least in part, by providing one or more computer-readable media containing computer-executable instructions for performing the operations described above to one or more computing devices configured to read and execute the computer-executable instructions. The system or device performs the operations of the above-described embodiments when executing the computer-executable instructions. Also, an operating system on one or more systems or devices may implement at least some of the operations of the above-described embodiments.
[0077] Furthermore, some embodiments use one or more functional units to implement the devices, systems, and methods described above. The functional units may be implemented solely in hardware (e.g., customized circuitry) or in a combination of software and hardware (e.g., a microprocessor running software).
[0078] Furthermore, some embodiments of the devices, systems, and methods combine features from two or more of the embodiments described herein. Also, as used herein, the conjunction "or" generally refers to an inclusive "or," but may refer to an exclusive "or" if expressly indicated, or if the context indicates that the "or" must be an exclusive "or."
[0079] While the present disclosure has been described with reference to exemplary embodiments, it is to be understood that the invention is not limited to the disclosed exemplary embodiments.
Claims
1. 1. A system for immersive virtual reality communication, comprising: a first capture device for capturing an image stream of a first user; a second capture device for capturing an image stream of a second user; a first virtual reality device used by the first user; a second virtual reality device used by the second user; Including, the first virtual reality device displays a virtual environment and a representation of the second user based at least in part on an image stream of the second user captured by the second capture device; the second virtual reality device displays the virtual environment and a representation of the first user based at least in part on the image stream of the first user captured by the first capture device; the virtual environment is modified based on a scenario selected for the immersive virtual reality communication; the first virtual reality device indicates a state of the first user relative to the first capture device in response to the selected scenario; the second virtual reality device indicates a state of the second user relative to the second capture device in response to the selected scenario. A system characterized by:
2. A viewpoint from which the virtual environment is rendered on the first virtual reality device is different from a viewpoint from which the virtual environment is rendered on the second virtual reality device.
2. The system of claim 1.
3. The indication of the state of the first user is given before the drawn image of the second user is displayed on the first virtual reality device; the indication of the state of the second user occurs before the representation of the first user is displayed on the second virtual reality device.
2. The system of claim 1.
4. The instruction of the state of the first user includes an instruction of movement and rotation of the first user, The instruction of the state of the second user includes an instruction of movement and rotation of the second user.
2. The system of claim 1.
5. The instruction of the state of the first user and the instruction of the state of the second user are determined based on at least one of a user pose, a user position, a user scale, or a virtual reality device boundary.
5. The system of claim 4.
6. A first network for transmitting the captured image stream of the first user; a second network for receiving data based at least in part on the captured image stream of the first user; and the first network includes a graphics processing unit; the data is generated entirely in the graphics processing unit before being transmitted to the second network.
2. The system of claim 1.
7. The first capture device is a camera that captures multiple images from different viewpoints.
2. The system of claim 1.
8. The captured image stream of the second user shows the second user wearing the second virtual reality device, and the depicted image of the second user shows the second user with the second virtual reality device removed.
2. The system of claim 1.
9. The first virtual reality device evaluating a state of the first user relative to the first capture device according to the selected scenario; indicating a state of the first user relative to the first capture device until an evaluation of the state of the first user relative to the first capture device satisfies a predetermined criterion; 2. The system of claim 1.
10. The state of the first user relative to the first capture device is evaluated based on at least one of the following: a pose of the first user; a position of the first user in the captured image stream of the first user; a size of the first user in the captured image stream of the first user; and a position of the first user within a boundary established for the first virtual reality device.
10. The system of claim 9.
11. 1. A method for immersive virtual reality communication, comprising: a first capture device capturing an image stream of a first user; a second capture device capturing an image stream of a second user; displaying a virtual environment and a representation of the second user based at least in part on an image stream of the second user captured by the second capture device; displaying the virtual environment and a representation of the first user based at least in part on an image stream of the first user captured by the first capture device; Including, the virtual environment is modified based on a scenario selected for the immersive virtual reality communication; The method comprises: indicating a state of the first user relative to the first capture device in response to the selected scenario; indicating a state of the second user relative to the second capture device in response to the selected scenario; further comprising: A method characterized by:
12. 1. A virtual reality capture and display system for immersive virtual reality communication, comprising: A capture device, a virtual reality device used by a first user; Including, The capture device is capture means for capturing an image stream of the first user; transmitting means for transmitting the image stream of the first user to a network; and the virtual reality device receiving means for receiving data from the network based at least in part on the image stream of a second user; a display means for displaying a virtual environment and a drawn image of the second user based on the data; and the virtual environment is modified based on a scenario selected for the immersive virtual reality communication; the virtual reality device further comprises an indicating means for indicating a state of the first user to be captured as an image stream of the first user in response to the selected scenario.
1. A virtual reality capture and display system comprising:
13. A program for causing a first virtual reality device and a second virtual reality device that perform immersive virtual reality communication to perform the operation of the system described in claim 1.
14. A method for controlling a capture and display system for immersive virtual reality communication, the method comprising: a capture device; and a virtual reality device used by a first user, the method comprising: The capture device, capturing an image stream of the first user; a transmitting step of transmitting the first user's image stream to a network; Execute The virtual reality device, receiving data from the network based at least in part on the image stream of a second user; a display step of displaying a virtual environment and a representation of the second user based on the data; Execute the virtual environment is modified based on a scenario selected for the immersive virtual reality communication; and causing the virtual reality device to further perform an indicating step of indicating a state of the first user to be captured as an image stream of the first user in response to the selected scenario. A control method comprising:
15. A program for causing the capture device to execute each step executed by the capture device of the control method described in claim 14.
16. A program for causing the virtual reality device to execute each step of the control method described in claim 14.
Citation Information
Patent Citations
Selective hand occlusion on virtual projection onto a physical surface using skeletal tracking
JP2014514652A
Calibration of virtual reality system
JP2017142813A
Content generation in an immersive environment
US10373342B1