Apparatus for generating a shared virtual conversation space with AR and non-AR devices using edge processing
The system allows non-AR devices to participate in AR video conferences by generating a virtual scene from non-AR data, addressing network and computation overhead, and ensuring consistent participant positioning, enhancing immersive experiences.
Patent Information
- Application Number
- JP2023548251
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-12-08
- Filing Date
- 2022-12-13
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2042-12-13
Smart Images

Figure 0007778804000001 
Figure 0007778804000002 
Figure 0007778804000003
Abstract
Description
[Technical Field]
[0001] The present disclosure is directed to providing a virtual conversation session using an augmented reality (AR) device, where each participant sees all other participants in their local space, but in the same arrangement as the other participants in that local space. That is, according to an exemplary embodiment, people are sitting / standing / etc. in the same configuration, as if they were all ordinary and of the same or similar orientation.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application US63 / 307,534, filed February 7, 2022, and U.S. Provisional Patent Application US18 / 077,672, filed December 8, 2022, the contents of which are expressly incorporated by reference herein in their entireties. [Background technology]
[0003] Even if an AR streaming device provides images of other participants in the conference, a non-AR device may not be able to participate in an AR video conference, even if the non-AR device has 360 video or 2D video capabilities. Summary of the Invention
[0004] To address one or more different technical problems, this disclosure provides technical solutions to reduce network overhead and server computation overhead, while providing options to apply various operations to the resolved elements, so that when using these operations, their practicality and some of the technical signaling characteristics may be improved.
[0005] Methods and apparatuses are included that include a memory configured to store computer program code, and one or more processors configured to access the computer program code and to act as directed by the computer program code. The computer program code includes acquisition code configured to cause at least one processor to acquire video data from a non-AR device and an AR device, respectively, where the non-AR device is used by a first user in a first room and the AR device is worn by a second user in a second room separate from the first room; acquisition code configured to cause at least one processor to acquire an AR scene description from the non-AR device that does not render an AR scene; generation code configured to cause the at least one processor to generate a virtual scene by analyzing and rendering the scene description acquired from the non-AR device; acquisition code configured to cause the at least one processor to determine an orientation of the non-AR device relative to a position where the second user is displayed in the AR scene in the first room based on the AR scene description acquired from the non-AR device; and streaming code configured to cause the at least one processor to stream the rendered virtual scene to the non-AR device based on the orientation determination. According to an exemplary embodiment, a non-AR device may be a device that is not configured to render AR scenes, such as a laptop, a smart TV, a smartphone, etc., according to an exemplary embodiment, and an AR device may be a device that is configured to render AR scenes and may include glass-type AR / mixed reality devices, etc.
[0006] According to an example embodiment, the position in the AR scene where the second user should be displayed is determined based on the view selection of the first user via the non-AR device.
[0007] According to an exemplary embodiment, streaming the scene information to the non-AR device includes streaming at least one of 360 video and 2D video in response to a selection by the first user via the non-AR device.
[0008] According to an exemplary embodiment, the scene information is generated on a cloud device separate from the non-AR device.
[0009] According to an exemplary embodiment, the cloud device performs AR rendering based on the video data and provides scene information to the non-AR device.
[0010] According to an exemplary embodiment, the scene information includes a second user virtually overlaid on a position in the first room.
[0011] According to an exemplary embodiment, the location at which the second user is virtually overlaid in the first room is a location in the first room that at least one of the non-AR device and the cloud device has determined to be a dedicated location in the first room for overlaying the second user during streaming of scene information.
[0012] According to an exemplary embodiment, the loud device further provides updated scene information to the non-AR device based on a view switch of the non-AR device via the first user moving the non-AR device in the first room.
[0013] According to an exemplary embodiment, audio from the first room and audio from the second room are mixed and provided to the non-AR device along with the scene information.
[0014] According to an exemplary embodiment, a second user of the AR device views a scene in the AR environment, while a first user of the non-AR device views the scene in the non-AR environment according to the scene description. [Brief explanation of the drawings]
[0015] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Figure 1] FIG. 1 is a simplified schematic diagram according to an embodiment. [Figure 2] FIG. 2 is a simplified schematic diagram according to an embodiment. [Figure 3] FIG. 3 is a simplified block diagram of a decoder according to an embodiment. [Figure 4] FIG. 4 is a simplified block diagram of an encoder according to an embodiment. [Figure 5] FIG. 5 is a simplified block diagram according to an embodiment. [Figure 6] FIG. 6 is a simplified block diagram according to an embodiment. [Figure 7] FIG. 7 is a simplified block diagram according to an embodiment. [Figure 8] FIG. 8 is a simplified block diagram according to an embodiment. [Figure 9] FIG. 9 is a simplified block diagram according to an embodiment. [Figure 10] FIG. 10 is a simplified diagram according to an embodiment. [Figure 11] FIG. 11 is a simplified block diagram according to an embodiment. [Figure 12] FIG. 12 is a simplified block diagram according to an embodiment. [Figure 13] FIG. 13 is a simplified block and timing diagram according to an embodiment. [Figure 14] FIG. 14 is a schematic diagram according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0016] The suggested functions described below may be used individually or in any order and combination. Furthermore, the embodiments may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.
[0017] 1 shows a simplified block diagram of a communication system 100 according to one embodiment of the present disclosure. The communication system 100 may include at least two terminals 102 and 103 interconnected via a network 105. For unidirectional transmission of data, a first terminal 103 may code video data at a local location for transmission to the other terminal 102 via the network 105. A second terminal 102 may receive the other terminal's coded video data from the network 105, decode the coded data, and display the recovered video data. Unidirectional data transmission may be common in media service applications, etc.
[0018] 1 shows a second pair of terminals 101 and 104 provided to support two-way transmission of coded video, such as may occur during a video conference. For the two-way transmission of data, each terminal 101 and 104 can code video data captured at a local location for transmission to the other terminal over network 105. Each terminal 101 and 104 can also receive coded video data transmitted by the other terminal, decode the coded data, and display the recovered video data on a local display device.
[0019] In FIG. 1 , terminals 101, 102, 103, and 104 may be depicted as servers, personal computers, and smartphones, but the principles of the present disclosure are not so limited. Embodiments of the present disclosure apply to laptop computers, tablet computers, media players, and / or dedicated videoconferencing devices. Network 105 represents any number of networks that convey coded video data between terminals 101, 102, 103, and 104, including, for example, wired and / or wireless communication networks. Communication network 105 can exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this description, the architecture and topology of network 105 are not important to the operation of the present disclosure, unless otherwise described herein below.
[0020] 2 shows the arrangement of a video encoder and decoder in a streaming environment as one example application of the disclosed technical subject matter, which may be equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0021] The streaming system may include a capture subsystem 203, which may include, for example, a video source 201, e.g., a digital camera, that generates an uncompressed video sample stream 213. The sample stream 213 may be emphasized as a high-data-volume stream when compared to an encoded video bitstream and may be processed by an encoder 202 coupled to the camera 201. The encoder 202 may include hardware, software, or a combination thereof for enabling or implementing aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream 204 may be emphasized as a low-data-volume stream when compared to the sample stream and may be stored on a streaming server 205 for future use. One or more streaming clients 212 and 207 may access the streaming server 205 to retrieve copies 208 and 206 of the encoded video bitstream 204. Client 212 may include a video decoder 211, which decodes an incoming copy of encoded video bitstream 208 and generates an outgoing video sample stream 210 that can be rendered on a display 209 or other rendering device (not shown). In some streaming systems, video bitstreams 204, 206, and 208 may be encoded according to a predetermined video coding / compression standard. Examples of these standards are described above and further herein.
[0022] figure 3 is a functional block diagram of a video decoder 300 according to one embodiment of the present invention.
[0023] Receiver 302 can receive one or more codec video sequences to be decoded by decoder 300, in the same or another embodiment, one coded video sequence at a time, where the decoding of each coded video sequence is independent of the other coded video sequences. The coded video sequences may be received from channel 301, which may be a hardware / software link to a storage device that stores the coded video data. Receiver 302 can receive the encoded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not shown). Receiver 302 can separate the coded video sequences from the other data. To combat network jitter, buffer memory 303 is coupled between receiver 302 and entropy decoder / parser 304 (hereinafter referred to as the "parser"). If receiver 302 is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from an isosychronous network, buffer 303 may not be needed or may be small. For use over a best-effort packet network such as the Internet, buffer 303 may be needed, which may be relatively large and advantageously adaptively sized.
[0024] The video decoder 300 may include a parser 304 for reconstructing symbols 313 from the entropy-coded video sequence. These symbol categories include information used to manage the operation of the decoder 300 and potential information for controlling a rendering device, such as a display 312, that is not an integral part of the decoder but can be coupled to it. The control information for the rendering device may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information parameter set fragment (not shown). The parser 304 may parse / entropy-decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 304 may extract, from the coded video sequence, a set of subgroup parameters for at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The entropy decoder / parser may also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.
[0025] Parser 304 may perform entropy decoding / parsing operations on the video sequence received from buffer 303 to generate symbols 313. Parser 304 may receive the encoded data and selectively decode particular symbols 313. Additionally, parser 304 may determine whether a particular symbol 313 should be provided to motion compensated prediction unit 306, scaler / inverse transform unit 305, intra prediction unit 307, or loop filter 311.
[0026] The reconstruction of symbols 313 may involve several different units, depending on the type of coded video picture or portion thereof (inter and intra picture, inter and intra block, etc.) and other factors. Which units participate and how can be controlled by subgroup control information parsed from the coded video sequence by parser 304. The flow of such subgroup control information between parser 304 and the following units is not shown for clarity.
[0027] In addition to the functional blocks already described, decoder 300 can be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed technical subject matter, the following conceptual subdivision into functional units is appropriate:
[0028] The first unit is a scalar / inverse transform unit 305. The scalar / inverse transform unit 305 receives quantized transform coefficients as well as control information from the parser 304 as symbols 313, including which transform to use, block size, quantization factors, quantization scaling matrices, etc. It can output blocks containing sample values, which can be input to an aggregator 310.
[0029] In some cases, the output samples of the scaler / inverse transform unit 305 may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 uses surrounding already reconstructed information fetched from the current (partially reconstructed) picture 309 to generate blocks of the same size and shape as the block under reconstruction. The aggregator 310 optionally adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit 307 to the output sample information as provided by the scaler / inverse transform unit 305.
[0030] In other cases, the output samples of the scalar / inverse transform unit 305 may relate to an inter-coded and potentially motion-compensated block. In such cases, the motion-compensated prediction unit 306 may access the reference picture memory 308 to fetch samples used for prediction. After motion-compensating the fetched samples according to the block-related symbols 313, these samples may be added by the aggregator 310 to the output of the scalar / inverse transform unit to generate output sample information (in this case, referred to as residual samples or residual signals). The addresses in the reference picture memory from which the motion compensation unit fetches prediction samples may be controlled by a motion vector. The motion vector is available to the motion compensation unit in the form of a symbol 313, which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values fetched from the reference picture memory when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.
[0031] The output samples of aggregator 310 may be subjected to various loop filtering techniques in loop filter unit 311. Video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video bitstream and made available to loop filter unit 311 as symbols 313 from parser 304, but may also be responsive to meta-information obtained during decoding of a coded picture or previous portion of a coded video sequence (in decoding order), as well as to previously reconstructed and loop-filtered sample values.
[0032] The output of the loop filter unit 311 may be a sample stream that can be output to a display 312, which may be a rendering device, or may be stored in a reference picture memory 557 for use in future intra-picture prediction.
[0033] Once a given coded picture is fully reconstructed, it can be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by parser 304), the current reference picture 309 can become part of reference picture buffer 308, and a new current picture memory follows. Ko The frames can be reallocated before starting the reconstruction of the coded picture.
[0034] Video decoder 300 can perform decoding operations in accordance with a predetermined video compression technique, which may be documented in a standard, such as ITU-TRec. H.265. The coded video sequence can comply with the syntax specified by the video compression technique or standard being used, in the sense that it conforms to the syntax of the video compression technique or standard as specified in the video compression technique document or standard, and particularly in the profile document therein. Compliance also requires that the complexity of the coded video sequence be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level can, in some cases, be further constrained through a Hypothetical Reference Decoder (HRD) specification signaled in the coded video sequence and metadata for HRD buffer management.
[0035] In one embodiment, the receiver 302 can receive additional (redundant) data along with the coded video. The additional data may be included as part of the coded video sequence. The additional data can be used by the video decoder 300 to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0036] FIG. 4 may be a functional block diagram of a video encoder 400 according to one embodiment of the present disclosure.
[0037] The encoder 400 may receive video samples from a video source 401 (not part of the encoder) that may capture video pictures to be coded by the encoder 400 .
[0038] The video source 401 can provide a source video sequence to be coded by the encoder (303) in the form of a digital video sample stream, which can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media services system, the video source 401 can be a storage device storing pre-prepared video. In a video conferencing system, the video source 401 can be a camera capturing local picture information as a video sequence. The video data can be provided as multiple individual pictures that convey motion when viewed sequentially. The picture itself can be organized as a spatial array of pixels, where each pixel can contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art can readily understand the relationship between pixels and samples. The following description focuses on samples.
[0039] According to one embodiment, the encoder 400 can code and compress pictures of a source video sequence into a coded video sequence 410 in real time or under any other time constraint required by the application. Enforcing the appropriate coding rate is one function of the controller 402. The controller controls and is operatively coupled to other functional units, as described below. The coupling is not shown for clarity. Parameters set by the controller may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can readily identify other functions of the controller 402 as may be relevant for a video encoder 400 optimized for a given system design.
[0040] Some video encoders operate in what those skilled in the art would readily recognize as a "coding loop." As an overly simplistic explanation, the coding loop can consist of an encoding portion of the encoder (e.g., a source coder 403) (responsible for creating symbols based on an input picture to be coded and reference pictures), and a (local) decoder 406 embedded in the encoder 400 that reconstructs the symbols to create sample data that a (remote) decoder would also create (since in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols and the coded video bitstream is lossless). The reconstructed sample stream is input to a reference picture memory 405. Because decoding of the symbol stream yields bit-exact results regardless of the location of the decoder (local or remote), the contents of the reference picture buffer are also bit-exact between the local and remote encoders. In other words, the prediction portion of the encoder "sees" exactly the same sample values as the reference picture samples that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronicity (and the resulting drift when synchronicity cannot be maintained, e.g., due to channel errors) is well known to those skilled in the art.
[0041] The operation of the "local" decoder 406 is identical to the operation of the "remote" decoder 300 already described in detail above in connection with Figure 3. However, and referring also briefly to Figure 4, the entropy decoding portion of the decoder 300, including the channel 301, receiver 302, buffer 303, and parser 304, may not be fully implemented in the local decoder 406 because symbols are available and the encoding / decoding of the symbols into a coded video sequence by the entropy coder 408 and parser 304 may be lossless.
[0042] An observation that can be made at this point is that any decoder technique other than analysis / entropy decoding that is present in the decoder must also be present, in substantially identical functional form, in the corresponding encoder. The description of the encoder techniques can be simplified because they are the inverse of the decoder techniques that have been comprehensively described. Only in certain areas is a more detailed description required and is provided below.
[0043] As part of its operation, the source coder 403 may perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence designated as "reference frames." In this manner, the coding engine 407 codes differences between pixel blocks of the input frame and pixel blocks of reference frames that may be selected as predictive references for the input frame.
[0044] The local video decoder 406 can decode coded video data of frames that may be designated as reference frames based on symbols generated by the source coder 403. The operation of the coding engine 407 can advantageously be a lossy process. When the coded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence can be a replica of the source video sequence, typically with some errors. The local video decoder 406 can replicate the decoding process that may be performed by the video decoder on the reference frames and cause the reconstructed reference frames to be stored in the reference picture memory 405, which may be, for example, a cache. In this way, the encoder 400 can locally store copies of reconstructed reference frames that have content in common with reconstructed reference frames acquired by a far-end video decoder (without transmission errors).
[0045] The predictor 404 can perform a prediction search for the coding engine 407. That is, for a new frame to be coded, the predictor 404 can search the reference picture memory 405 for sample data (as candidate reference pixel blocks) or predetermined metadata, such as reference picture motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new picture. The predictor 404 can operate on a block-by-pixel basis of samples to find appropriate prediction references. In some cases, as determined by the search results obtained by the predictor 404, the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory 405.
[0046] The controller 402 may manage the coding operations of the video coder 403, including, for example, setting the parameters and subgroup parameters used to encode the video data.
[0047] The output of all the above-mentioned functional units may be entropy coded in entropy coder 408. The entropy coder converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.
[0048] Transmitter 409 can buffer the coded video sequence created by entropy coder 408 and prepare it for transmission over communication channel 411, which can be a hardware / software link to a storage device that will store the encoded video data. Transmitter 409 can merge the coded video data from video coder 403 with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0049] A controller 402 can manage the operation of the encoder 400. During coding, the controller 405 can assign each coded picture a predetermined coded picture type, which can affect the coding technique that can be applied to the respective picture. For example, pictures may often be assigned as one of the following frame types:
[0050] An intra picture (I-picture) may be one that can be coded and decoded without using other frames in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh Pictures. Those skilled in the art are aware of these variations of I-pictures and their respective applications and functions.
[0051] A predictive picture (P picture) may be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values of each block.
[0052] A Bi-directionally Predictive Picture (B picture) may be coded and decoded using intra- or inter-prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, a multiple-predictive picture may use more than one reference picture and associated metadata for the reconstruction of a single block.
[0053] A source picture may generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and coded on a block-by-block basis. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be nonpredictively coded, or they may be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be nonpredictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be nonpredictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0054] Video coder 400 may perform coding operations in accordance with a given video coding technique or standard, such as ITU-T Rec. H.265. In doing so, video coder 400 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in an input video sequence. The coded video data may therefore conform to a syntax specified by the video coding technique or standard being used.
[0055] In one embodiment, the transmitter 409 can transmit additional data along with the encoded video. The source coder 403 may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.
[0056] Figure 5 is an example end-to-end architecture 500 for a standalone AR (STAR) device, showing a 5G STAR user equipment (UE) receiver 600, a network / cloud 501, and a 5G UE (sender) 700, according to an example embodiment. Figure 6 is a more detailed example 600 of one or more configurations for the STAR UE receiver 600, according to an example embodiment, and Figure 7 is a more detailed example 700 of one or more configurations for a 5G UE transmitter 700, according to an example embodiment. 3GPP TR26.998 defines support for glass-type augmented reality / mixed reality (AR / MR) devices in 5G networks, and according to example embodiments herein, at least two device classes are considered: 1) devices fully capable of decoding and playing complex AR / MR content (Standalone AR or STAR), and 2) devices with smaller computational resources and / or smaller physical size (and therefore battery) and capable of running such applications only if the majority of the computation is performed on the 5G edge server, network, or cloud rather than on the device (Edge-bound AR or EDGAR).
[0057] Thus, according to an example embodiment, as described below, a shared conversation use case can be experienced in which all participants in the shared AR conversation experience have AR devices, and each participant sees the other participants in an AR scene, where participants are overlays in a local physical scene, their placement in that scene is consistent across all receiving devices, e.g., people in each local space have the same positions / seating arrangement relative to each other, and such a virtual space creates the sense of being in the same space, but the room varies from participant to participant because the room is the actual room or space in which each person is physically located.
[0058] For example, according to the exemplary embodiment shown with respect to FIGS. 5-7 , an immersive media processing function on the network / cloud 501 receives uplink streams from various devices and creates a scene description that defines the arrangement of individual participants in a single virtual conference room. The scene description and encoded media streams are distributed to each receiving participant. The receiving participant's 5G STAR UE 600 receives, decodes, and processes the 3D video and audio streams and renders them using the received scene description and information received from its AR runtime to create an AR scene of the virtual conference room with all other participants. The virtual room for a participant is based on their own physical space, but the seating / position arrangement of all other participants in the room is consistent with the virtual rooms of all other participants in the session.
[0059]
[0043] Also see Figure 8, which illustrates an example EDGAR device architecture 800, according to an exemplary embodiment. Here, the device itself, such as the 5G EDGAR UE 900, is not capable of heavy processing. Therefore, scene analysis and media analysis of the received content is performed in the cloud / edge 801, and then a simplified AR scene with a small number of media components is delivered to the device for processing and rendering. Figure 9 illustrates a more detailed example of the 5G EDGAR UE 900, according to an exemplary embodiment.
[0060] However, even with capabilities such as those associated with the exemplary embodiments of Figures 5-9, immersive media functionality may present one or more technical challenges with constructing a common virtual space scene description, if any, and as described below, such embodiments may be technically improved in the context of immersive media processing functionality to generate a scene description that would be provided to all participants so that they all experience the same relative placement of participants within the local AR scene.
[0061] FIG. 10 shows a case where a user A 10, a user B 11, and a user T 12 participate in an AR conference room, and one or more of the users AR 10 shows an example embodiment 1000 without a device. As shown, user A 10 is in an office 1001, sitting in a conference room with a varying number of chairs, and user A 10 is in charge of that chair. User B 11 is in their living room 1002, sitting on a loveseat; his living room also has one or more two-person couches, as well as other furniture such as chairs and tables. User T 12 is in an airport lounge 1003, sitting on a bench with a bench across a coffee table among one or more other coffee tables.
[0062] Then, when looking at the AR environment, in the office 1001, the AR of user A 10 shows to user A 10 a virtual user B 11v1 corresponding to user B 11 and a virtual user T 12v1 corresponding to user T 12, and as a result, virtual user B 11v1 and virtual user T 12v1 are shown to user A 10 as if they are sitting in furniture, such as office chairs, in the office 1001, just like user A 10. 1000 In the living room 1002 In this example, the AR of user B 11 corresponds to user T 12, but in the living room 1002 Virtual user T 12v2 sitting on the couch and in the office1001 Not an office chair, but a living room chair 1002 The figure shows a virtual user A 10v1 corresponding to a user A 10 sitting on furniture in an airport lounge. 1003 Also, there, the AR of user T 12 corresponds to user A 10, but in the airport lounge 1003 It also shows virtual user A 10v2 sitting at the table in front of virtual user A 10v2, and virtual user B 11v2 sitting at the table opposite virtual user A 10v2. 1001 ,living room 1002 , and airport lounges 1003 In each of the above, the updated scene description for each room is consistent with the other rooms in terms of location / seating arrangement. For example, user A 10 is shown counterclockwise relative to user 11 or his virtual representation, which is shown clockwise relative to user T 12 or his virtual representation per room.
[0063] However, AR technology has been limited to any attempt to incorporate the creation and use of virtual spaces for devices that do not support AR but are capable of analyzing VR or 2D video, and embodiments herein provide improved technical procedures for creating virtual scenes that are consistent with AR scenes when such devices participate in a shared AR conversation service.
[0064] 11 illustrates an example end-to-end architecture 1100 involving a non-AR device 1101 and a cloud / edge 1102, according to an example embodiment. FIG. 12 illustrates a more detailed block diagram of the example non-AR device 1101.
[0065] As shown in Figures 11 and 12, non-AR UE 1101 is a device capable of rendering 360 or 2D video, but does not have any AR capabilities. However, edge functionality on cloud / edge 1102 is capable of receiving scenes, rendering scenes, and AR rendering of immersive visual and audio objects in a virtual room selected from a library. The entire video is then encoded and delivered to device 1101 for decoding and rendering.
[0066] Thus, there may be multi-view capabilities, such that the AR processing on the edge / cloud 1102 can generate multiple videos of the same virtual room, from different angles and using different viewports, and the device 1101 can receive one or more of these videos and switch between them as needed, or send commands to the edge / cloud processing to stream only the desired viewport / angle.
[0067] There may also be a background change capability, where a user on device 1101 can select a desired room background from a provided library, for example, one of different conference room or even living room layouts, and cloud / edge 1102 then uses the selected background and creates the virtual room accordingly.
[0068] 13 shows an example timing diagram 1300 for an example call flow for an immersive AR conversation to a receiving non-AR UE 1101. For purposes of explanation, only one sender is shown in this diagram without showing the detailed call flow.
[0069] Shown are an AR application module 21, a media player module 22, and a media access function module 23, which may be considered modules of the receiving non-AR UE 1101. Also shown is a cloud / edge partitioned rendering module 24. Also shown are a media delivery module 25 and a scene graph composition module 26 of each of the network clouds 1102. Also shown is a 5G sender UE module 700.
[0070] S1-S6 can be considered a session establishment phase. In S1, the AR application module 21 can request the media access function module 23 to start a session, and in S2, the media access function module 23 can request the cloud / edge split rendering module 24 to start a session.
[0071] The cloud / edge partitioned rendering module 24 may perform session negotiation with the scene graph synthesis module 26 in S3, which may then negotiate with the 5G sender UE 700. If successful, then in S5 the cloud / edge partitioned rendering module may send an acknowledgement to the media access function module 23, and the media access function module 23 may send an acknowledgement to the AR application module 21.
[0072] S7 can then be considered a media pipeline configuration stage, where the media access function module 23 and the cloud / edge partitioned rendering module 24 each configure their respective pipelines. Then, after that pipeline configuration, a session can be initiated by signals from the AR application module to the media player module 22 at s8, from the media player module 22 to the media access function module 23 at s9, and from the media access function module 23 to the cloud / edge partitioned rendering module 24 at S10.
[0073] There is then a pose loop stage from S11 to S13, where in S11 pose data is provided from the media player module 22 to the AR application module 21, and in S12 the AR application module provides the pose data 12 to the media access function module 23, which can then provide the pose data to the cloud / edge split rendering module 24.
[0074] S14 through S16 can be considered a shared experience stream stage, in which the 5G sender UE 700 can provide a media stream to the media delivery module 25 at S14 and provide AR data to the scene graph composition module 26 at S15. The scene graph composition module 25 can then compose one or more scenes based on the received AR data and provide the scene and scene updates to the cloud / edge partitioned rendering module 24 at S16, which in turn can provide the media stream to the cloud / edge partitioned rendering module at S17. This can include, according to an example embodiment, acquiring an AR scene descriptor from a non-AR device that does not render the AR scene, and generating a virtual scene by the cloud device by parsing and rendering the scene description acquired from the non-AR device.
[0075] S18 to S19 can be considered a media uplink stage, where the media player module 22 captures and processes media data from the local user and, at S18, can provide the media data to the media access function module 23. The media access module 23 can then encode the media and, at S19, provide the media stream to the cloud / edge split rendering module 24.
[0076] Between S19 and S20, the cloud / edge split rendering module 24 can perform scene analysis and complete AR rendering. S20 and S21 can then be considered to constitute a media stream loop stage. At S20, the cloud / edge split rendering module 24 can provide a media stream to the media access function module 23, which can then decode the media and provide media rendering to the media player 22 at S21.
[0077] Such functionality in accordance with an example embodiment allows a non-AR UE 1101 that does not have a see-through display, and therefore cannot create AR scenes, to nevertheless take advantage of a display that is capable of rendering VR or 2D video. Thus, its immersive media processing capabilities only generate a common scene description and are not shared with other participants. To Each participant's position relative to and Scene The scene itself needs to be calibrated with pose information on each device before being rendered as an AR scene, as described above. An AR rendering process on the edge or in the cloud can then analyze the AR scene and create a simplified VR-2D scene.
[0078] According to an example embodiment, this disclosure uses a similar split rendering process of an EDGAR device for non-AR devices, such as VR or 2-d video equipment, where the edge / cloud AR rendering process does not generate any AR scenes. Instead, it generates a virtual scene by parsing and rendering a scene description received from an immersive media processing capability against a given background (such as a conference room), and then renders each participant in the location described by the scene description in the conference room.
[0079] The resulting video can also be 360 video or 2D video depending on the capabilities of the receiving non-AR device, and the resulting video is generated taking into account location information received from the non-AR device in accordance with an exemplary embodiment.
[0080] Also, each other participant with a non-AR device is added as a 2D video overlay on the 360 / 2D video of the conference room, as shown in FIG. 10, and the conference room is also a furniture area on which virtual pictures are overlaid, as shown in FIG. region It is also possible to have a dedicated area used for these overlays, such as:
[0081] Additionally, audio signals from all participants may be mixed, if desired, to produce a single channel of audio carrying the voices within the room. Video may be encoded as a single 360 or 2-D video and distributed to the devices, and, optionally, multiple video (multi-view) sources may be generated, each capturing the same virtual conference room from a different view and providing those views to the devices, according to an exemplary embodiment.
[0082] Additionally, the non-AR UE device 1101 、3 60 videos and / or one or more multiview videos of your choice With audio Reception and rendered on the device display The user can then switch between different views or change the viewport of the 360 video by moving or rotating the viewing device and thus navigate within the virtual room while watching the video.
[0083] The techniques described above can be implemented using computer-readable instructions and as computer software physically stored on one or more computer-readable media, or by one or more specifically configured hardware processors. For example, Figure 14 illustrates a computer system 1400 suitable for implementing certain embodiments of the disclosed subject matter.
[0084] Computer software can be coded using any suitable machine code or computer language and can be subject to mechanisms such as assembly, compilation, linking, etc. to generate code including instructions that can be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., directly or via interpretation, microcode execution, etc.
[0085] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0086] 14 for computer system 1400 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of computer system 1400.
[0087] The computer system 1400 may include certain human interface input devices that can respond to input by one or more human users through, for example, tactile input (such as keystrokes, swipes, data glove movements), audio input (such as speech, clapping), visual input (such as gestures), and olfactory input (not shown). The human interface devices may also be used to capture certain media that do not necessarily involve direct conscious human input, such as audio (such as voice, music, and ambient sounds), images (such as scanned images and photographic images acquired from still-image cameras), and video (such as two-dimensional video and three-dimensional video, including stereoscopic video).
[0088] The input human interface devices may include one or more of a keyboard 1401, a mouse 1402, a trackpad 1403, a touchscreen 1410, a joystick 1405, a microphone 1406, a scanner 1408, and a camera 1407 (only one of each is shown).
[0089] The computer system 1400 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen 1410 or joystick 1405, although some haptic feedback devices may not function as input devices), audio output devices (such as speakers 1409, headphones (not shown)), visual output devices (such as screens 1410, including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capability and each with or without haptic feedback capability—some of which may provide two-dimensional visual output or three-dimensional or higher-dimensional output via means such as stereo output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0090] Computer system 1400 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 1420 using CD / DVD 1411 or similar media, thumb-drive 1422, removable hard drive or solid-state drive 1423, conventional magnetic media such as tape and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.
[0091] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not encompass transmission media, carrier waves, or other transitory signals.
[0092] The computer system 1400 may also include an interface 1499 to one or more communication networks 1498. The network 1498 may be, for example, wireless, wired, optical, etc. The network 1498 may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks 1498 include local area networks such as Ethernet; cellular networks including WLAN, GSM, 3G, 4G, 5G, LTE, etc.; TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CANbus; etc. Certain networks 1498 generally require an external network interface adapter (e.g., a USB port on computer system 1400) connected to certain general-purpose data ports or peripheral buses (1450 and 1451); others are generally integrated into the core of computer system 1400 by connecting to a system bus, as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks 1498, computer system 1400 can communicate with other entities. Such communication may be unidirectional, receive only (e.g., broadcast TV), unidirectional transmit only (e.g., a CAN bus to a certain CANbus device), or bidirectional. For example, to other computer systems using local or wide area digital networks. Predetermined protocols and protocol stacks may be used in each of these networks and network interfaces, as described above.
[0093] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be connected to the core 1440 of the computer system 1400 .
[0094] The core 1440 may include one or more central processing units (CPUs) 1441, graphics processing units (GPUs) 1442, graphics adapters 1417, dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) 1443, hardware accelerators 1444 for certain tasks, etc. These devices, along with read-only memory (ROM) 1445, random access memory 1446, and internal mass storage devices 1447 such as non-user-accessible internal hard drives, SSDs, etc., may be connected via a system bus 1448. In some computer systems, the system bus 1448 may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices are connected directly to the core's system bus 1448 or via a peripheral bus 1451. Architectures for peripheral buses include PCI, USB, etc.
[0095] The CPU 1441, GPU 1442, FPGA 1443, and accelerator 1444 can execute predetermined instructions, which, in combination, may constitute the above-mentioned computer code, which may be stored in ROM 1445 or RAM 1446. Transient data may also be stored in RAM 1446, while permanent data may be stored, for example, in an internal mass storage device 1447. Fast storage and retrieval from any memory device may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU 1441, GPU 1442, mass storage device 1447, ROM 1445, RAM 1446, etc.
[0096] The computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those having skill in the computer software arts.
[0097] As one example, and not by way of limitation, an architecture corresponding to computer system 1400, and specifically core 1440, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible, computer-readable media. Such computer-readable media can be user-accessible mass storage devices, as described above, as well as media associated with a given storage device associated with core 1440 that is non-transitory in nature, such as core-internal mass storage device 1447 or ROM 1445. Software implementing various embodiments of the present disclosure can be stored on such devices and executed by core 1440. Computer-readable media can include one or more memory devices or chips, depending on particular needs. The software can cause core 1440, and specifically the processors therein (including a CPU, GPU, FPGA, etc.), to perform particular processes or portions of particular processes described herein, including defining data structures stored in RAM 1446 and modifying those data structures in accordance with the software-defined processes. Additionally, or alternatively, a computer system may provide functionality as a result of logic embedded in hardware or otherwise embedded in circuitry (e.g., accelerator 1444) that may operate in place of, or in conjunction with, software executing particular processes or portions of particular processes described herein. References to software may encompass logic, where appropriate, and vice versa. References to computer-readable media may encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware and software.
[0098] While this disclosure has described several exemplary embodiments, there are alterations, substitutions, and various alternatives and equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art can devise numerous systems and methods that, while not explicitly shown or described herein, embody the principles of the present disclosure and, therefore, are within its spirit and scope.
Claims
1. A method for augmented reality (AR) video streaming performed by a separate cloud / edge device, comprising: acquiring video data from a non-AR device and an AR device, respectively; the non-AR device is being used by a first user in a first room, and the non-AR device does not have the capability to render an AR scene; the AR device is worn by a second user in a second room different from the first room; Steps and acquiring a scene description from the non-AR device and the AR device, the scene description describing the scene and each participant's relative position with respect to other participants, and provided to all participants such that all participants experience the same relative placement of the participants within the scene; generating a virtual scene by analyzing and rendering the scene description obtained from the non-AR device and the AR device, wherein the virtual scene is generated by adjusting the scene with pose information from the non-AR device; determining an orientation of the non-AR device relative to a position where the second user should be displayed within an AR scene in the first room based on the scene description; streaming the generated virtual scene to the non-AR device based on the orientation determination; A method comprising:
2. the location where the second user should be displayed within the AR scene is determined based on a view selection of the first user via the non-AR device. The method of claim 1.
3. Streaming the generated virtual scene to the non-AR device includes: Streaming at least one of 360 video and 2D video in response to a selection of the first user via the non-AR device; The method of claim 1 , comprising:
4. The cloud / edge device performs AR rendering based on the video data and provides the generated virtual scene to the non-AR device. The method of claim 1.
5. the generated virtual scene includes the second user virtually overlaid at the position in the first room. The method of claim 4.
6. The location at which the second user is virtually overlaid in the first room is a location in the first room that is determined by at least one of the non-AR device and the cloud / edge device to be a dedicated location in the first room for overlaying the second user during streaming of the generated virtual scene. The method of claim 5.
7. The cloud / edge device further providing updated scene information to the non-AR device based on a change in a view of the non-AR device caused by the first user moving the non-AR device in the first room; The method of claim 4.
8. The step of streaming the generated virtual scene to the non-AR device, comprising: mixing audio from the first room and audio from the second room and providing the audio to the non-AR device along with the generated virtual scene; The method of claim 1 , comprising:
9. the AR device presents a scene in an AR environment to the second user; Meanwhile, the non-AR device presents the scene to the first user in a non-AR environment according to the generated virtual scene. The method of claim 1.
10. An apparatus for augmented reality (AR) video streaming configured as a cloud / edge device, comprising: at least one memory configured to store computer program code; at least one processor configured to access the computer program code and to act as directed by the computer program code; The computer program code causes the at least one processor to: acquiring video data from a non-AR device and an AR device, respectively, wherein the non-AR device is used by a first user in a first room, the non-AR device does not have a function for rendering an AR scene, and the AR device is worn by a second user in a second room separate from the first room; acquiring a scene description from the non-AR device and the AR device, the scene description describing the scene and each participant's relative position with respect to other participants, and providing the scene description to all participants such that all participants experience the same relative placement of the participants within the scene; generating a virtual scene by analyzing and rendering the scene description obtained from the non-AR device and the AR device, wherein the virtual scene is generated by adjusting the scene with pose information from the non-AR device; determining an orientation of the non-AR device relative to a position where the second user should be displayed within an AR scene in the first room based on the scene description; Streaming the generated virtual scene to the non-AR device based on the orientation determination; and A device that performs the following.
11. the location where the second user should be displayed within the AR scene is determined based on a view selection of the first user via the non-AR device.
11. The apparatus of claim 10.
12. Streaming the generated virtual scene to the non-AR device includes: Streaming at least one of 360 video and 2D video in response to a selection of the first user via the non-AR device.
11. The apparatus of claim 10.
13. The apparatus performs AR rendering based on the video data and provides the generated virtual scene to the non-AR device.
11. The apparatus of claim 10.
14. the generated virtual scene includes the second user virtually overlaid at a position in the first room.
14. The apparatus of claim 13.
15. The location at which the second user is virtually overlaid in the first room is a location in the first room that is determined by at least one of the non-AR device and the apparatus to be a dedicated location in the first room for overlaying the second user during streaming of the generated virtual scene.
15. The apparatus of claim 14.
16. The device further comprises: providing updated scene information to the non-AR device based on a change in a view of the non-AR device caused by the first user moving the non-AR device in the first room; 14. The apparatus of claim 13.
17. The method of claim 16, wherein the step of streaming the generated virtual scene to the non-AR device comprises: mixing audio from the first room and audio from the second room and providing the audio to the non-AR device along with the generated virtual scene; The apparatus of claim 10, comprising:
18. A computer program comprising a plurality of instructions executable on a computer configured as a cloud / edge device, the instructions, when executed, causing the computer to perform a process; The process comprises: acquiring video data from a non-AR device and an AR device, respectively; the non-AR device is being used by a first user in a first room, and the non-AR device does not have the capability to render an AR scene; the AR device is worn by a second user in a second room different from the first room; Steps and acquiring a scene description from the non-AR device and the AR device, the scene description describing the scene and each participant's relative position with respect to other participants, and provided to all participants such that all participants experience the same relative placement of the participants within the scene; generating a virtual scene by analyzing and rendering the scene description obtained from the non-AR device and the AR device, wherein the virtual scene is generated by adjusting the scene with pose information from the non-AR device; determining an orientation of the non-AR device relative to a position where the second user should be displayed within an AR scene in the first room based on the scene description; streaming the generated virtual scene to the non-AR device based on the orientation determination; Including, Computer program.
Citation Information
Patent Citations
Placement of Virtual Content in Environments with a Plurality of Physical Participants
US20210084259A1
Systems and methods for generating multi-user augmented reality content
US20220028170A1
System for the rendering of shared digital interfaces relative to each user's point of view
WO2012135554A1