A device for creating a shared virtual AR conversation space over a network for multiple participants
The system addresses the challenge of synchronizing virtual overlays in AR by generating a common scene description for AR devices, reducing overhead and ensuring consistent participant placement across different physical spaces.
Patent Information
- Application Number
- JP2023565567
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-12-08
- Filing Date
- 2022-12-13
- Publication Date
- 2025-09-25
- Estimated Expiration
- 2042-12-13
AI Technical Summary
Existing augmented reality (AR) streaming devices fail to synchronize the virtual overlays of participants across different physical spaces, leading to inconsistent orientations and placements, and result in high network and server computational overhead.
A system that includes a memory and processors to acquire video data from multiple AR devices, determine user orientations, generate a common scene description, and control AR devices to display participants overlaid on objects in a shared virtual space, reducing network and server overhead.
Enables synchronized virtual overlays of participants in different physical spaces with reduced network and server computational overhead, providing a consistent shared AR experience across diverse environments.
Smart Images

Figure 0007744436000001 
Figure 0007744436000002 
Figure 0007744436000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 307,552, filed February 7, 2022, and U.S. Application No. 18 / 063,403, filed December 8, 2022, the contents of which are expressly incorporated by reference in their entireties.
[0002] The present disclosure is directed, according to an example embodiment, to providing a virtual conversation session using an augmented reality (AR) device, in which each participant sees all other participants in their local space, but the participants' placement in the local space is the same as the others, i.e., they are sitting / standing / etc. in the same setting, as if they were all in a common location and facing the same or similar directions. [Background technology]
[0003] Even if AR streaming devices can provide images of other participants to the conference, they cannot synchronize the virtual overlays of other participants with the various objects and depths of each user in different physical spaces, or enable each AR user to see other AR users in their own respective environments and with similar orientations shared by each AR participant. Summary of the Invention [Means for solving the problem]
[0004] To address one or more different technical problems, the present disclosure provides technical solutions that reduce network overhead and server computational overhead, while providing options for applying various operations to the resolved elements, so that the use of these options may improve their practicality and some of the technical signaling functions.
[0005] Methods and apparatuses are included that include a memory configured to store computer program code and one or more processors configured to access the computer program code and act as instructed by the computer program code. The computer program code includes: an acquisition code configured to cause at least one processor to acquire video data from a first AR device and a second AR device, where the first AR device is worn by a first user in a first physical space and the second AR device is worn by a second user in a second physical space separate from the first physical space; a determination code configured to cause the at least one processor to determine, based on the video data, a first orientation of the first user with respect to an object in the first physical space and a second orientation of the second user with respect to an object in the second physical space; a generation code configured to cause the at least one processor to generate a common scene description for both the first AR device and the second AR device based on the determining of the first orientation and the second orientation; and a control code configured to cause the at least one processor to control the first AR device to display, to the first user, a representation of the second user overlaid on an object in a virtual scene including the first physical space in which the first user is located, based on the common scene description. As described herein, the physical space may be a room or may be another space other than a room, such as an outdoor space, and the object may be one or more pieces of furniture, etc., according to an exemplary embodiment.
[0006] According to an exemplary embodiment, the step of generating a common scene description is performed by the first AR device.
[0007] According to an exemplary embodiment, the step of generating the common scene description is performed by a network device separate from each of the first AR device and the second AR device.
[0008] According to an exemplary embodiment, the step of controlling the first AR device to display the second user virtually overlaid on an object in the first physical space is further based on the step of checking whether the second user is to be virtually overlaid on a portion of at least one other object in the first physical space.
[0009] According to an example embodiment, the computer program code includes further acquisition code configured to cause the at least one processor to acquire third video data from a third AR device, the third AR device being worn by a third user in a third physical space; further determination code configured to cause the at least one processor to determine a third orientation of the third user relative to any object in the third physical space based on the third video data; and further generation code configured to cause the at least one processor to generate a common scene description further based on the determining the third orientation.
[0010] According to an exemplary embodiment, the computer program code further includes further control code configured to cause the at least one processor to control the first AR device to display the third user to the first user such that the third user is virtually overlaid on a second object in the first physical space based on the common scene description.
[0011] According to an exemplary embodiment, the orientation of the second user virtually overlaid on the object and the third user overlaid on the second object is determined based on a determination that the first user should be virtually overlaid on the other object in a corresponding orientation in each of the second physical space and the third physical space, relative to the relative viewpoints of each of the second user and the third user.
[0012] According to an exemplary embodiment, at least one of the first physical space, the second physical space, and the third physical space is a physical space within a residence, and others of the first physical space, the second physical space, and the third physical space are located in one of an office and a public space.
[0013] According to an exemplary embodiment, at least one of the first object and the second object is an office chair, and the third object is one of a loveseat, a chaise lounge, and a coffee table.
[0014] According to an example embodiment, a common scene description is provided to each of a first AR device, a second AR device, and a third AR device.
[0015] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0016] [Figure 1] FIG. 1 is a simplified schematic diagram according to an embodiment. [Figure 2] FIG. 1 is a simplified schematic diagram according to an embodiment. [Figure 3] FIG. 2 is a simplified block diagram of a decoder according to an embodiment. [Figure 4] FIG. 2 is a simplified block diagram of an encoder according to an embodiment. [Figure 5] FIG. 1 is a simplified block diagram according to an embodiment. [Figure 6] FIG. 1 is a simplified block diagram according to an embodiment. [Figure 7] FIG. 1 is a simplified block diagram according to an embodiment. [Figure 8] FIG. 1 is a simplified block diagram according to an embodiment. [Figure 9] FIG. 1 is a simplified block diagram according to an embodiment. [Figure 10]FIG. 1 is a simplified diagram according to an embodiment. [Figure 11] FIG. 1 is a simplified diagram according to an embodiment. [Figure 12] FIG. 1 is a simplified diagram according to an embodiment. [Figure 13] 1 is a simplified flowchart according to an embodiment. [Figure 14] 1 is a schematic diagram according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0017] The proposed features discussed below may be used separately or combined in any order. Furthermore, the embodiments may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program recorded on a non-transitory computer-readable medium.
[0018] 1 illustrates a simplified block diagram of a communication system 100 according to one embodiment of the present disclosure. The communication system 100 may include at least two terminals 102, 103 interconnected via a network 105. For unidirectional data transmission, a first terminal 103 may encode video data at a local location for transmission to the other terminal 102 via the network 105. The second terminal 102 may receive the other terminal's encoded video data from the network 105, decode the encoded data, and display the recovered video data. Unidirectional data transmission may be common in media serving applications, etc.
[0019] 1 illustrates a second pair of terminals 101 and 104 provided to support bidirectional transmission of encoded video, such as might occur during a video conference. For the bidirectional transmission of data, each terminal 101 and 104 may encode captured video data at a local location for transmission to the other terminal over network 105. Each terminal 101 and 104 may also receive encoded video data transmitted by the other terminal, decode the encoded data, and display the recovered video data on a local display device.
[0020] In FIG. 1 , terminals 101, 102, 103, and 104 may be illustrated as a server, a personal computer, and a smartphone, although the principles of the present disclosure are not so limited. Embodiments of the present disclosure find application with laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. Network 105 represents any number of networks that convey coded video data among terminals 101, 102, 103, and 104, including, for example, wired and / or wireless communication networks. Communication network 105 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of network 105 may not be important to the operation of the present disclosure, unless otherwise described herein below.
[0021] 2 illustrates the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application of the disclosed subject matter, which may be equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, and storage of compressed video on digital media including CDs, DVDs, memory sticks, and the like.
[0022] The streaming system may include a capture subsystem 203, which may include a video source 201, such as a digital camera, that creates an uncompressed video sample stream 213. The sample stream 213 may be enhanced as a high data volume compared to an encoded video bitstream and may be processed by an encoder 202 coupled to the camera 201. The encoder 202 may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream 204 may be enhanced as a lower data volume compared to the sample stream and may be recorded on a streaming server 205 for future use. One or more streaming clients 212 and 207 may access the streaming server 205 to retrieve copies 208 and 206 of the encoded video bitstream 204. The client 212 may include a video decoder 211 that decodes the incoming copy 208 of the encoded video bitstream and creates an outgoing video sample stream 210 that can be rendered on a display 209 or other rendering device (not shown). In some streaming systems, the video bitstreams 204, 206, and 208 may be encoded according to particular video coding / compression standards, examples of which are mentioned above and further described herein.
[0023] FIG. 3 may be a functional block diagram of a video decoder 300 according to one embodiment of the present invention.
[0024] Receiver 302 may receive one or more codec video sequences to be decoded by decoder 300. In the same or another embodiment, receiver 302 may receive one coded video sequence at a time, with the decoding of each coded video sequence being independent of other coded video sequences. The coded video sequences may be received from channel 301, which may be a hardware / software link to a storage device that records the coded video data. Receiver 302 may receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, that may be forwarded to a respective using entity (not shown). Receiver 302 may separate the coded video sequences from other data. To combat network jitter, buffer memory 303 may be coupled between receiver 302 and entropy decoder / parser 304 (hereinafter "parser"). If receiver 302 is receiving data from a storage / forwarding device with sufficient bandwidth and controllability or from an isosynchronous network, buffer 303 may not be necessary or may be small. For use in a best effort packet network such as the Internet, buffer 303 may be required and may be relatively large, and may advantageously be adaptively sized.
[0025] The video decoder 300 may include a parser 304 for reconstructing symbols 313 from the entropy-coded video sequence. These symbol categories include information used to manage the operation of the decoder 300 and, potentially, information for controlling a rendering device, such as a display 312, that is not an integral part of the decoder but may be coupled to it. The rendering device control information may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information Parameter Set Fragment (not shown). The parser 304 may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 304 may extract from the coded video sequence a set of subgroup parameters for at least one of a subgroup of pixels in the video decoder based on at least one parameter corresponding to that group. Subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The entropy decoder / parser may also extract from the coded video sequence information such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0026] The parser 304 may perform entropy decoding / parsing operations on the video sequence received from the buffer 303 to create symbols 313. The parser 304 may receive the encoded data and selectively decode particular symbols 313. Additionally, the parser 304 may determine whether a particular symbol 313 should be provided to the motion compensated prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.
[0027] The reconstruction of symbols 313 may involve several different units, depending on the type of coded video picture or portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block, etc.), as well as other factors. Which units are involved and how may be governed by subgroup control information parsed from the coded video sequence by parser 304. The flow of such subgroup control information between parser 304 and the following units is not shown for clarity.
[0028] In addition to the functional blocks already mentioned, decoder 300 may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate:
[0029] The first unit is the scalar / inverse transform unit 305. The scalar / inverse transform unit 305 receives quantized transform coefficients and control information from the parser 304 as symbols 313, including the transform to use, block size, quantization factor, quantization scaling matrix, etc. It can output blocks containing sample values that can be input to the aggregator 310.
[0030] In some cases, the output samples of the scaler / inverse transform unit 305 may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 uses surrounding already reconstructed information fetched from the current (partially reconstructed) picture 309 to generate blocks of the same size and shape as the block being reconstructed. The aggregator 310 may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305.
[0031] In other cases, the output samples of the scalar / inverse transform unit 305 may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion-compensated prediction unit 306 may access the reference picture memory 308 to fetch samples used for prediction. After motion-compensating the fetched samples according to symbols 313 related to the block, these samples may be added by an aggregator 310 to the output of the scalar / inverse transform unit to generate output sample information (in this case, referred to as residual samples or residual signals). The addresses within the reference picture memory from which the motion compensation unit fetches prediction samples may be controlled by a motion vector and may be made available to the motion compensation unit in the form of symbols 313, which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values fetched from the reference picture memory when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.
[0032] The output samples of aggregator 310 may be subjected to various loop filtering techniques in loop filter unit 311. Video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video bitstream and made available to loop filter unit 311 as symbols 313 from parser 304, but may also be responsive to meta-information obtained during decoding of previous (in decoding order) parts of the coded picture or coded video sequence, or to previously reconstructed and loop-filtered sample values.
[0033] The output of the loop filter unit 311 can be a sample stream that can be output to the rendering device 312 as well as recorded in the reference picture memory 557 for use in future inter-picture prediction.
[0034] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture has been identified as a reference picture (e.g., by parser 304), the current reference picture 309 can become part of reference picture buffer 308, and new current picture memory can be reallocated before starting reconstruction of the following coded picture.
[0035] The video decoder 300 may perform decoding operations according to a predetermined video compression technology, which may be documented in a standard such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technology or standard being used, in the sense of adhering to the syntax of the video compression technology or standard as specified in the video compression technology document or standard, specifically the profile document therein. Compliance may also require that the complexity of the coded video sequence be within a range defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained by a hypothetical reference decoder (HRD) specification and HRD buffer management metadata signaled in the coded video sequence.
[0036] In one embodiment, the receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder 300 to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, a temporal layer, a spatial layer, or a signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.
[0037] FIG. 4 may be a functional block diagram of a video encoder 400 according to one embodiment of the present disclosure.
[0038] The encoder 400 may receive video samples from a video source 401 (not part of the encoder) that may capture video images to be encoded by the encoder 400 .
[0039] The video source 401 may provide the source video sequence to be coded by the encoder (303) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCB, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media distribution system, the video source 401 may be a storage device that records previously prepared video. In a video conferencing system, the video source 401 may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that, when viewed sequentially, create motion. The pictures themselves are organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following discussion focuses on samples.
[0040] According to one embodiment, the encoder 400 may encode and compress pictures of a source video sequence into a coded video sequence 410 in real time or under any other time constraint required by the application. Achieving an appropriate coding rate is one function of the controller 402. The controller controls and is operatively coupled to other functional units, as described below. For clarity, coupling is not depicted. Parameters set by the controller may include rate control-related parameters (e.g., picture skip, quantizer, lambda value for rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art will readily identify other functions of the controller 402 as they may pertain to optimizing the video encoder 400 for a particular system design.
[0041] Some video encoders operate in what those skilled in the art will readily recognize as a "coding loop." As an overly simplified explanation, the coding loop may consist of an encoding portion, an encoder 402 (hereafter "source coder") (responsible for creating symbols based on the input picture to be coded and reference pictures), and a (local) decoder 406 embedded in the encoder 400 that reconstructs the symbols to create sample data that a (remote) decoder will also create (since any compression between the symbols and the coded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream is input to a reference picture memory 405. Because decoding of the symbol stream yields bit-exact results regardless of the decoder's location (local or remote), the reference picture buffer contents are also bit-exact between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift when synchronization cannot be maintained, for example due to channel errors) is well known to those skilled in the art.
[0042] The operation of the "local" decoder 406 may be the same as that of the "remote" decoder 300, which has already been described in detail above in connection with Figure 3. However, briefly referring also to Figure 4, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder 408 and parser 304 may be lossless, the entropy decoding portion of the decoder 300, including the channel 301, receiver 302, buffer 303 and parser 304, may not be fully implemented in the local decoder 406.
[0043] An observation that can be made at this point is that any decoder technology other than analysis / entropy decoding that is present in a decoder must also necessarily be present in the corresponding encoder, in substantially the same functional form. The description of the encoder technology can be omitted, since it is the inverse of the decoder technology that has been described generically. Only in certain areas is more detailed description required, and is provided below.
[0044] As part of its operation, source coder 403 may perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence, designated as “reference frames.” In this method, coding engine 407 codes the differences between pixel blocks of the input frame and pixel blocks of reference frames that may be selected as predictive references for the input frame.
[0045] The local video decoder 406 may decode the encoded video data of frames that may be designated as reference frames based on the symbols created by the source coder 403. The operation of the coding engine 407 may advantageously be a lossy process. When the encoded video data is decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may typically be a copy of the source video sequence, with some errors. The local video decoder 406 may replicate the decoding process that may be performed by the video decoder on the reference frames and record the reconstructed reference frames in the reference picture cache 405. In this way, the encoder 400 may locally record copies of reconstructed reference frames that have common content as the reconstructed reference frames that will be retrieved by a far-end video decoder (without transmission errors).
[0046] The predictor 404 may perform the prediction search for the coding engine 407. That is, for a new frame to be coded, the predictor 404 may search the reference picture memory 405 for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable prediction references for the new picture. The predictor 404 may operate on a pixel block-by-pixel block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor 404, the input picture may have prediction references drawn from multiple reference pictures recorded in the reference picture memory 405.
[0047] Controller 402 may manage the coding operations of video coder 403, including, for example, setting parameters and subgroup parameters used to encode the video data.
[0048] The output of all the aforementioned functional units may undergo entropy coding in entropy coder 408. The entropy coder converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.
[0049] The transmitter 409 may buffer the encoded video sequence created by the entropy coder 408 and prepare it for transmission over a communication channel 411, which may be a hardware / software link to a storage device that will record the encoded video data. The transmitter 409 may merge the encoded video data from the video coder 403 with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0050] The controller 402 may manage the operation of the encoder 400. During coding, the controller 405 may assign a particular coding picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures may often be assigned as one of the following frame types:
[0051] An intra-picture (I-picture) may be a picture that can be coded and decoded without using any other frame in a sequence as a source of prediction. Some video codecs allow various types of intra-pictures, including, for example, independent decoder refresh pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.
[0052] A predicted picture (P picture) may be a picture that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict sample values for each block.
[0053] Bidirectionally predicted pictures (B-pictures) may be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0054] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be nonpredictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be nonpredictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Pixel blocks of a B-picture may be nonpredictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0055] Video coder 400 may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, video coder 400 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to a syntax specified by the video coding technique or standard being used.
[0056] In one embodiment, the transmitter 409 may transmit additional data along with the coded video. The source coder 403 may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplemental Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.
[0057] Figure 5 is an example end-to-end architecture 500 for a standalone AR (STAR) device according to an example embodiment, showing a 5G STAR user equipment (UE) receiver 600, a network / cloud 501, and a 5G UE (transmitter) 700. Figure 6 is a further detailed example 600 of one or more configurations for the STAR UE receiver 600 according to an example embodiment, and Figure 7 is a further detailed example 700 of one or more configurations for the 5G UE transmitter 700 according to an example embodiment. 3GPP TR 26.998 ("3GPP" is a registered trademark) specifies support for glasses-based augmented reality / mixed reality (AR / MR) devices in 5G networks. Also, in accordance with example embodiments herein, at least two classes of devices are contemplated: 1) devices that are fully capable of decoding and playing complex AR / MR content (standalone AR, i.e., STAR); and 2) devices that have smaller computational resources and / or smaller physical size (and therefore battery) and can only run such applications when the majority of the computation is performed on the 5G edge server, network, or cloud rather than on the device (edge-dependent AR, i.e., EDGAR).
[0058] Also, according to exemplary embodiments, as described below, a shared conversation use case can be experienced in which all participants in a shared AR conversation experience have AR devices, each participant sees the other participants in the AR scene, the participants are overlays in the local physical scene, and the placement of participants in the scene is consistent in all receiving devices, e.g., people in each local space have the same positions / seating arrangement relative to each other, and such virtual spaces create the feeling of being in the same space, but the physical space is different for each participant because it is the actual physical space or space in which each person is physically located.
[0059] For example, according to the exemplary embodiment illustrated with reference to FIGS. 5-7 , an immersive media processing function on the network / cloud 501 receives uplink streams from various devices and constructs a scene description that defines the placement of individual participants within a single virtual conference physical space. The scene description and encoded media streams are distributed to each receiving participant. The receiving participant's 5G STAR UE 600 receives, decodes, processes, and renders the 3D video and audio streams using the received scene description and information received from its AR runtime to create an AR scene of the virtual conference physical space with all other participants. While a participant's virtual physical space is based on their own physical space, the seating / positioning of all other participants in the physical space matches the virtual physical spaces of all other participants in the session.
[0060]
[0041] According to an exemplary embodiment, see also Figure 8, which illustrates an example EDGAR device architecture 800; a device such as a 5G EDGAR UE 900 itself cannot perform significant processing. Therefore, scene analysis and media analysis on received content is performed in the cloud / edge 801, and then a simplified AR scene with a small number of media components is delivered to the device for processing and rendering. Figure 9 illustrates a more detailed example of a 5G EDGAR UE 900 according to an exemplary embodiment.
[0061] However, even with functionality such as that associated with the exemplary embodiments of Figures 5-9, there may be one or more technical challenges associated with constructing a common virtual space scene description for immersive media functionality. Also, as described below, such embodiments are technically improved in the context of immersive media processing functionality to generate a scene description that is provided to all participants so that all participants experience the same relative placement of the participants within the local AR scene.
[0062] 10 illustrates an example 1000 in which users A10, B11, and T12 are participating in an AR conference physical space. As shown, user A10 is in an office 1001, seated in the conference physical space with a varying number of chairs, with user A10 sitting in a chair. User B11 is in his / her living physical space 1002, seated on a loveseat, along with one or more loveseats and other objects in his / her living physical space, such as chairs and tables. User T12 is on a bench in an airport lobby 1003, across from a coffee table among one or more other coffee tables.
[0063] As described in the example flowchart 1300 of Figure 13, Figure 12 illustrates an example 1200 in which users A10, B11, and T12 share a virtual meeting in their own respective areas, each represented with a relative orientation that matches each other's viewpoint, objects, and pose by each AR device receiving information about each participant's position and overall layout such that the immersive media functionality generates a general scene description that describes the relative positions. That general scene description is shown in the example of Figure 1100 as description 1101 and corresponds to a common scene graph (CSG) shown as CSG 1102, which shows a root node 13 and nodes: user A10's node 10n (which may be named Alice, at least for graph purposes), user B11's node 11n (which may be named Bob, at least for graph purposes), and user T12's node 12n (which may be named Tom, at least for graph purposes).
[0064] The CSG 1102 and description 1101 can be sent to all participating devices. The AR runtime engine of each device customizes a scene graph based on the actual layout and positions of the chairs, tables, sofas, and benches. It decodes and renders each participant's media objects and overlays them on the AR scene according to the customized scene graph.
[0065] FIG. 13 shows an example flowchart 1300 of an immersive media processing function (IMPF) that may be performed by each device of a user, such as user A10, user B11, and user T12, and / or any other participant.
[0066] In S131, one or more photos of what a person sees in their space are displayed: Alice's meeting space and table, Bob's living space, and Tom's airport lobby. Depending on the user's respective capture device, one or more of the following may be provided: one or more photos or videos, a depth map of the physical spaces, and the relative sizes of the physical spaces. Alternatively, in S132, such information may be derived in addition to the IMPF, which derives the number of participants for the CSG. Also, in S133, for each received physical space, such as the office 1001, the living space 1002, and the airport lobby 1003, the process uses object detection to search for locations in each physical space, identifies one or more possible correct positions for overlaying participants, and continues to find the respective positions of each participant. In this example, two positions for the two participants are added to each of the office 1001, the living space 1002, and the airport lobby 1003, thereby generating a scene description that includes the participants' locations within the physical spaces.
[0067] In S134, a comparison of all derived physical space scene descriptions is performed, and a common scene description (CSG) is derived from it, in which the positions of all participants in all corresponding physical spaces are the same and which is semantically valid for the physical spaces. For example, referring to example 1200 in FIG. 12 , in office 1001, user A10's AR shows user A10 virtual user B11v1 corresponding to user B11 and virtual user T12v1 corresponding to user T12, and virtual user B11v1 and virtual user T12v1 are shown to user A10 as sitting in an object in office 1001, i.e., an office chair, just like user A10. Also, see living physical space 1202 in example 1200, in which the AR for user B11 corresponds to user T12, but virtual user T12v2 is sitting on a couch in living physical space 1202, and virtual user A10v1 corresponding to user A10 is also sitting in an object in living physical space 1202, not in an office chair in office 1201. See also airport lobby 1203, where user T12's AR shows virtual user A10v2, who corresponds to user A10 but is seated at a table in airport lobby 1203, and virtual user B11v2, who is seated at a table opposite virtual user A10v2. Also, in each of these office 1201, living physical space 1202, and airport lobby 1203, the updated scene description of each physical space matches the other physical spaces in terms of location / seating arrangement. For example, user A10 is shown with its virtual representation oriented counterclockwise relative to user T11 or clockwise relative to user T12, or its virtual representation per physical space. Also, in S135, if a satisfactory orientation and arrangement was reached in S134, further information is obtained from one or more of S131, S132, and S133 while checking the derived common scene description (CSG) for all devices; otherwise, the process provides a satisfactory CSG for all participating devices.
[0068] Each step of the example flowchart 1300 of FIG. 13 may be performed by each participating device, such as the 5G STAR US 600 of example 500 of FIG. 5, or may be performed in part or in large part by a cloud edge, such as the cloud / edge 801 of example 800 of FIG. 8, or by one or more of the devices illustrated with respect to either of FIGS. 5 and 13.
[0069] The techniques described above may be implemented using computer-readable instructions, as computer software physically stored on one or more computer-readable media, or by one or more tangibly configured hardware processors. For example, Figure 14 illustrates a computer system 1400 suitable for implementing certain embodiments of the disclosed subject matter.
[0070] Computer software can be coded using any suitable machine code or computer language that can be subjected to mechanisms such as assembly, compilation, linking, etc. to create code containing instructions that can be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., directly or via interpretation, microcode execution, etc.
[0071] The instructions may be executed on various types of computers or computer components including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.
[0072] 14 for computer system 1400 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. Neither the arrangement of components should be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of computer system 1400.
[0073] The computer system 1400 may include certain human interface input devices that can respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).
[0074] The input human interface devices may include one or more of a keyboard 1401, a mouse 1402, a trackpad 1403, a touchscreen 1410, a joystick 1405, a microphone 1406, a scanner 1408, and a camera 1407 (only one of each is shown).
[0075] The computer system 1400 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., touchscreen 1410 or haptic feedback via joystick 1405, although there may also be haptic feedback devices that do not function as input devices), audio output devices (such as speakers 1409, headphones (not shown)), visual output devices (such as screens 1410, including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capabilities, each with or without haptic feedback capabilities, some of which may be capable of outputting two-dimensional visual output or output in more than three dimensions via means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0076] The computer system 1400 may also include human-accessible storage devices and their associated media, such as CD / DVD 1411 or CD / DVD ROM / RW 1420 with similar media, thumb drives 1422, removable hard drives or solid state drives 1423, legacy magnetic media such as tape and floppy disks (not shown), optical media including dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), and the like.
[0077] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.
[0078] The computer system 1400 may also include an interface 1499 to one or more communication networks 1498. The network 1498 may be, for example, wireless, wired, optical, etc. Furthermore, the network 1498 may be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks 1498 include local area networks such as Ethernet, cellular networks including WLAN, GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicular and industrial networks including CANBus, etc. Particular networks 1498 typically require external network interface adapters coupled to particular general-purpose data ports or peripheral buses (1450 and 1451) (e.g., USB ports on computer system 1400), while others are typically integrated into the core of computer system 1400 by coupling to a system bus, as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks 1498, computer system 1400 can communicate with other entities. Such communications may be unidirectional receive only (e.g., broadcast TV), unidirectional transmit only (e.g., CANbus to certain CANbus devices), or bidirectional, e.g., to other computer systems using local-area or wide-area digital networks. Specific protocols and protocol stacks may be used with each of these networks and network interfaces, as described above.
[0079] The aforementioned human interface devices, human-accessible storage, and network interfaces may be coupled to the core 1440 of the computer system 1400 .
[0080] The core 1440 may include one or more central processing units (CPUs) 1441, graphics processing units (GPUs) 1442, graphics adapters 1417, dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) 1443, hardware accelerators for specific tasks 1444, etc. These devices may be connected via a system bus 1448, along with read-only memory (ROM) 1445, random access memory 1446, and internal mass storage devices 1447 such as internal hard drives, SSDs, etc. that are not user accessible. In some computer systems, the system bus 1448 may also be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be coupled directly to the core's system bus 1448 or via a peripheral bus 1451. Architectures for peripheral buses include PCI, USB, etc.
[0081] The CPU 1441, GPU 1442, FPGA 1443, and accelerator 1444 may combine to execute specific instructions that may constitute the aforementioned computer code, which may be stored in ROM 1445 or RAM 1446. Persistent data may be stored, for example, in internal mass storage device 1447, while transient data may also be stored in RAM 1446. Cache memory, which may be closely associated with one or more of the CPU 1441, GPU 1442, mass storage device 1447, ROM 1445, RAM 1446, etc., may be used to enable fast storage and retrieval to any of the memory devices.
[0082] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0083] By way of example, and not limitation, the architecture, and in particular the computer system 1400 having the core 1440, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage devices, as introduced above, as well as media associated with specific storage devices of the core 1440 that are non-transitory in nature, such as the core's internal mass storage device 1447 or the ROM 1445. Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core 1440. The computer-readable medium can include one or more memory devices or chips, depending on particular needs. The software can cause the core 1440, and in particular the processors therein (including a CPU, GPU, FPGA, etc.), to perform particular processes or portions of particular processes described herein, including defining data structures stored in RAM 1446 and modifying such data structures in accordance with the software-defined processes. Additionally, or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 1444) that can operate in place of or together with software to perform particular processes or portions of particular processes described herein. References to software can encompass logic, and vice versa, as appropriate. References to computer-readable media can encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.
[0084] While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope. [Explanation of symbols]
[0085] 100 Communication Systems 101 terminals 102 Second Terminal 103 First Terminal 105 Network 201 Video Sources 201 Camera 202 Encoder 203 Capture Subsystem 204 Video Bitstream 205 Streaming Server 208 copies 209 Display 210 Stream 211 Video Decoder 212 Streaming Client 213 Stream 300 decoder 301 Channel 302 Receiver 303 Buffer Memory 303 Encoder 304 Parser 305 Scaler / Descaler Unit 306 Motion Compensation Prediction Unit 307 Intra Prediction Unit 308 Reference Picture Memory 308 Reference Picture Buffer 309 Current Reference Picture 310 Aggregator 311 Loop Filter 311 units 312 Rendering Device 312 Display 313 Symbol 400 Encoder 401 Source 402 Encoder 402 Controller 403 Source Coder 403 Video Coder 404 Predictor 405 Controller 405 Reference Picture Cache 405 Reference Picture Memory 406 decoder 407 Coding Engine 408 Entropy Coder 409 Transmitter 410 coded video sequence 411 Communication Channel 500 examples 501 Cloud 557 Reference Picture Memory 600 receiver 700 examples 800 examples 801 Edge 1000 examples 1001 Office 1002 Living physical space 1003 Lobby 1101 Description 1102 CSG 1200 examples 1201 Office 1202 Living physical space 1203 Lobby 1300 Flowchart 1400 Computer Systems 1401 keyboard 1402 Mouse 1403 Trackpad 1405 Joystick 1406 Mike 1407 Camera 1408 Scanner 1409 Device Speaker 1410 Touchscreen 1417 Graphics Adapter 1422 thumb drive 1423 Solid State Drive 1440 cores 1441 CPU 1442 GPU 1443 FPGA 1444 Accelerator 1445 ROM 1446 RAM 1447 Mass Storage 1448 System Bus 1450 Peripheral Bus 1451 Peripheral bus 1498 Communication Network 1499 Interface
Claims
1. 1. A method for augmented reality (AR) video streaming, comprising: acquiring video data from a first AR device, a second AR device, and a third AR device, wherein the first AR device is worn by a first user in a first physical space, the second AR device is worn by a second user in a second physical space separate from the first physical space, and the third AR device is worn by a third user in a third physical space; determining a first orientation of the first user relative to any object in the first physical space based on the video data; determining a second orientation of a second user relative to any object in the second physical space based on the video data; determining a third orientation of the third user relative to any object in the third physical space based on the video data; generating a common scene description for all of the first AR device, the second AR device, and the third AR device based on determining the first orientation, the second orientation, and the third orientation; controlling the first AR device to display to the first user a representation of the second user overlaid on objects in a virtual scene including the first physical space in which the first user is located based on the common scene description; controlling the first AR device to display the third user to the first user such that the third user is virtually overlaid on a second object in the first physical space based on the common scene description; Including, the relative placement of the first, second, and third users in each of the virtual scenes including the first, second, and third physical spaces is the same; method.
2. generating the common scene description is performed by the first AR device; The method of claim 1.
3. generating the common scene description is performed by a network device separate from each of the first AR device and the second AR device; The method of claim 1.
4. and wherein controlling the first AR device to display the second user virtually overlaid on the object in the first physical space is further based on checking whether the second user is virtually overlaid on a portion of at least one other object in the first physical space. The method of claim 1.
5. orientations of the second user virtually overlaid on the object and the third user overlaid on the second object are determined based on a determination that the first user should be virtually overlaid on another object in a corresponding orientation in each of the second physical space and the third physical space with respect to a relative viewpoint of each of the second user and the third user; The method of claim 1.
6. at least one of the first physical space, the second physical space, and the third physical space is a physical space within a residence; the other of the first physical space, the second physical space, and the third physical space is in one of an office and a public space; The method of claim 1.
7. at least one of the first object and the second object is an office chair; The third object is one of a loveseat, a chaise lounge, and a coffee table. The method of claim 1.
8. the common scene description is provided to each of the first AR device, the second AR device, and the third AR device; The method of claim 1.
9. 1. An apparatus for augmented reality (AR) video streaming, comprising: at least one memory configured to store computer program code; accessing said computer program code; at least one processor configured to operate as instructed by the code, said computer program code comprising: acquisition code configured to cause the at least one processor to acquire video data from a first AR device, a second AR device, and a third AR device, wherein the first AR device is worn by a first user in a first physical space, the second AR device is worn by a second user in a second physical space separate from the first physical space, and the third AR device is worn by a third user in a third physical space; and determination code configured to cause the at least one processor to determine, based on the video data, a first orientation of the first user relative to any object in the first physical space, a second orientation of a second user relative to any object in the second physical space, and a third orientation of the third user relative to any object in the third physical space; generation code configured to cause the at least one processor to generate a common scene description for all of the first AR device, the second AR device, and the third AR device based on determining the first orientation, the second orientation, and the third orientation; control code configured to cause the at least one processor to control the first AR device to display to the first user a representation of the second user overlaid on an object in a virtual scene including the first physical space in which the first user is located based on the common scene description; and further control code configured to cause the at least one processor to control the first AR device to display the third user to the first user such that the third user is virtually overlaid on a second object in the first physical space based on the common scene description; Including, the relative placement of the first, second, and third users in each of the virtual scenes including the first, second, and third physical spaces is the same; Device.
10. generating the common scene description is performed by the first AR device; 10. The apparatus of claim 9.
11. generating the common scene description is performed by a network device separate from each of the first AR device and the second AR device; 10. The apparatus of claim 9.
12. controlling the first AR device to display the second user virtually overlaid on the object in the first physical space is further based on checking whether the second user is virtually overlaid on a portion of at least one other object in the first physical space.
10. The apparatus of claim 9.
13. orientations of the second user virtually overlaid on the object and the third user overlaid on the second object are determined based on a determination that the first user should be virtually overlaid on another object in a corresponding orientation in each of the second physical space and the third physical space with respect to a relative viewpoint of each of the second user and the third user; 10. The apparatus of claim 9.
14. at least one of the first physical space, the second physical space, and the third physical space is a physical space within a residence; the other of the first physical space, the second physical space, and the third physical space is in one of an office and a public space; 10. The apparatus of claim 9.
15. at least one of the first object and the second object is an office chair; The third object is one of a loveseat, a chaise lounge, and a coffee table.
10. The apparatus of claim 9.
16. A non-transitory computer-readable medium having recorded thereon a program for causing a computer to execute a process, the process comprising: acquiring video data from a first AR device, a second AR device, and a third AR device, wherein the first AR device is worn by a first user in a first physical space, the second AR device is worn by a second user in a second physical space separate from the first physical space, and the third AR device is worn by a third user in a third physical space; determining a first orientation of the first user relative to any object in the first physical space based on the video data; determining a second orientation of a second user relative to any object in the second physical space based on the video data; and determining a third orientation of the third user relative to any object in the third physical space based on the video data; and generating a common scene description for all of the first AR device, the second AR device, and the third AR device based on determining the first orientation, the second orientation, and the third orientation; controlling the first AR device to display to the first user a representation of the second user virtually overlaid on objects in a virtual scene including the first physical space in which the first user is located based on the common scene description; controlling the first AR device to display the third user to the first user such that the third user is virtually overlaid on a second object in the first physical space based on the common scene description; Including, the relative placement of the first, second, and third users in each of the virtual scenes including the first, second, and third physical spaces is the same; Non-transitory computer-readable medium.
Citation Information
Patent Citations
Systems and methods for presentation of augmented reality supplemental content in combination with presentation of media content
US20190130655A1
Augmented reality computing environments - collaborative workspaces
US20190313059A1
Automated understanding of three dimensional (3D) scenes for augmented reality applications
US20190392630A1