Rendering of augmented reality content and method for implementing AR content rendering
By outsourcing lighting estimation and re-illumination processing to external devices, the problem of limited resources of AR devices is solved, efficient and realistic lighting rendering effect is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202380030566.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-04-07
- Filing Date
- 2023-02-22
- Publication Date
- 2025-07-22
AI Technical Summary
Existing AR devices have limited resources in lighting estimation, resulting in unrealistic rendering of virtual objects in the real world, affecting the user experience.
Light estimation and re-illumination processing are entrusted to external processing equipment. By transmitting encoded video picture data and lighting information, external equipment performs lighting estimation and then transmits it to AR equipment for re-illumination effects of virtual objects.
It realizes efficient and realistic lighting rendering effects on AR devices with limited resources, improving users' immersive experience.
Smart Images

Figure CN120359757A_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to augmented reality (AR) and to devices and methods for rendering or implementing the rendering of AR video frames. In particular, but not exclusively, this application relates to lighting effects applied in AR video frames and to the processing involved in implementing such lighting effects. Background Art
[0002] This section is intended to introduce the reader to aspects of the art that may be related to aspects of at least one embodiment of this application described and / or claimed below. This discussion is considered to be helpful in providing background information to facilitate a better understanding of the various aspects of this application. Accordingly, it should be understood that these statements are to be read from this perspective and not as an admission of related art.
[0003] Augmented reality (hereinafter referred to as "AR") can be defined as an experience in which volumetric information representing (one or more) virtual objects ("the augmentation" part) is superimposed on a video frame representing the real-world environment perceived by the user (the "reality" part). The visual AR information resulting from the combination of the video frame and the virtual information can be displayed by means of various well-known AR devices such as AR glasses, head-mounted headsets or mobile devices.
[0004] The game Pokemon GO TM may be regarded as one of the first large-scale deployed consumer AR services (in 2016). Since then, a large number of AR services have been widely applied in various fields such as tourism, education, healthcare, navigation systems, construction, retail, etc.
[0005] To support various AR functions or services, AR devices may include multiple sensors (cameras, trackers, gyroscopes, accelerometers...) and processing modules (codecs, vision engines, renderers...). For example, according to the 3GPP standardization committee dedicated to the standardization and deployment of AR services over 5G, AR devices can be classified into four categories: "5G Standalone AR UE", "5G EDGe-dependent AR UE", "5G Wireless Tethered AR UE" and "5G Wired Tethered AR UE" (XR over 5G presentation to VR-IF: https: / / www.vr-if.org / wp-content / uploads / VRIF-April-2021-Workshop-3GPP-SA4-presentation.pdf). Each AR device architecture can be tailored to specific use cases or scenarios identified by 3GPP.
[0006] Although distribution technologies for carrying 3D-based models have been studied and standardized for decades (H.264 / MVC, H.265 / MV-HEVC, 3D-HEVC, Google Draco, etc.), due to different integration conditions, augmented reality, virtual reality, and mixed reality (denoted as AR, VR, and MR respectively) are concepts that have only become mainstream in services and products (AR glasses, headsets, etc.) on the market today. 5G is also regarded as the main carrier for deploying XR (XR = AR / VR / MR)-based services in consumers' daily lives. The maturity of 3D capture technologies strongly supports this emergence, especially those based on point cloud capture and multi-view plus depth.
[0007] The concept of AR technology is to insert visual (augmented) information into the captured real-world environment. In particular, virtual objects can be inserted into the video frame of a real-world scene to obtain an AR video frame. For this purpose, lighting is a key aspect for providing a realistic experience to the user. In fact, superimposing virtual objects on an environment with inappropriate lighting and shadows may disrupt the immersion: the objects may look like they are floating in the air or are not actually part of the scene.
[0008] Figure 1 An AR video frame 10 obtained, for example, by inserting a virtual object 14 into the video frame 12 of a real-world scene is shown. The virtual object 14 may create an unwanted impression of floating in its real-world surrounding environment, apparently due to insufficient lighting.
[0009] Therefore, for providing an appropriate user experience, correct re-lighting of the inserted virtual objects is crucial. Non-re-lighted objects lacking coherent shadows may disrupt the immersion and fun of commercial AR experiences.
[0010] Over the past few years, known lighting estimation methods have been developed to determine the lighting information of a perceived or captured scene (where virtual objects are inserted), so that specific effects can be added to the virtual objects, thus integrating gracefully and smoothly into the real-world environment. With the emergence of artificial intelligence, learning-based methods have also emerged.
[0011] However, the problem is that AR devices do not always support light estimation, or if they do, they support it in a limited and unsatisfactory way, which may lead to unrealistic rendering effects. In particular, for AR devices (especially those with limited power resources, such as AR devices implemented as mobile devices, tablets or headsets), achieving a smooth experience through real-time processing is a challenge. Even for pre-trained deep learning-based solutions, light estimation algorithms may require powerful resources. Embedding processing capabilities in a handheld device may be limited, and the battery may quickly run out (loss of autonomy) due to computationally intensive tasks (such as decoding or 3D scene processing or rendering).
[0012] Therefore, the implemented compromise algorithms may produce unrealistic effects that harm the user experience rather than enhancing the immersive scene.
[0013] Therefore, realistic AR content is needed, and in particular, proper re-lighting of virtual objects in the AR video frame is needed to ensure a good and immersive user experience. In particular, for all AR devices, optimal rendering quality of AR content is desired. Summary of the Invention
[0014] The following section presents a simplified summary of at least one embodiment to provide a basic understanding of some aspects of the present application. This summary is not an exhaustive overview of the embodiments. It is not intended to identify key or important elements of the embodiments. The following summary only presents some aspects of at least one embodiment in a simplified form as a prelude to the more detailed description provided elsewhere in the document.
[0015] According to a first aspect of the present application, there is provided a method for performing AR content rendering implemented by an AR device (or AR apparatus), the method comprising:
[0016] - Transmitting encoded source video frame data of a video frame representing a real-world scene to a processing device external to the AR device;
[0017] - Receiving light information from the processing device;
[0018] - Generating an AR video frame by:
[0019] ○ Obtaining an aggregated video frame by aggregating the video frame and volume information representing at least one virtual object;
[0020] And
[0021] ○ Combining a re-lighting effect associated with at least one virtual object into the aggregated video frame based on the received light information; and
[0022] - Render the AR video frame.
[0023] In an embodiment, the illumination information and the environmental map data are obtained by decoding a bitstream of the encoded video frame data received from a processing device. A relighting effect incorporated into the aggregated video frame is determined by applying the illumination information and the environmental map data to at least one virtual object inserted into the aggregated video frame.
[0024] In an embodiment, the incorporating the relighting effect includes at least one of the following:
[0025] - superimposing a reconstructed illumination determined based on the illumination information on at least one virtual object inserted into the aggregated video frame; and
[0026] - superimposing a shadow effect representing at least one reconstructed shadow of the at least one virtual object on the aggregated video frame, the shadow effect being determined based on the illumination information.
[0027] In an embodiment, the volume information is obtained by any one of the following:
[0028] - receiving the volume information from a content server independently of the illumination information received from the processing device; and
[0029] - receiving the volume information and the illumination information as part of a bitstream of the encoded video frame data.
[0030] In an embodiment, the illumination information includes at least one of the following characteristics of light in a source video frame:
[0031] - illumination direction;
[0032] - illumination intensity;
[0033] - illumination position;
[0034] - color temperature;
[0035] - spectrum;
[0036] - light level; and
[0037] - environmental spherical harmonics.
[0038] According to a second aspect of the present application, there is provided a method for implementing AR content rendering implemented by a processing device (or processing apparatus), the method including:
[0039] - receiving encoded video frame data representing a video frame of a real-world scene from an AR device external to the processing device;
[0040] - Obtaining a video frame of a real-world scene by decoding received encoded video frame data;
[0041] - Determining lighting information by performing lighting estimation based on the video frame; and
[0042] - Transmitting the lighting information to an AR device for performing AR rendering of the video frame.
[0043] In an embodiment, the method further includes:
[0044] - Obtaining spatial mapping data included in the encoded video frame data by decoding the encoded video frame data, the spatial mapping data defining a 3D coordinate system;
[0045] wherein the lighting information is determined based on the spatial mapping data such that the processing device and the AR device share the 3D coordinate system.
[0046] In an embodiment, the method further includes:
[0047] - Obtaining environmental map data based on a source video frame, the environmental map data defining a 3D view of an estimated lighting as an image; and
[0048] - Encoding the environmental map data in a bitstream transmitted to the AR device as an encoded image.
[0049] In an embodiment, the method further includes:
[0050] - Obtaining lighting information that defines lighting applied by a lighting device to the real-world scene to capture the source video frame;
[0051] wherein the lighting information is determined based on the lighting information.
[0052] In an embodiment, the method further includes:
[0053] - Inserting the lighting information as metadata in a bitstream of the encoded video frame data transmitted to the AR device.
[0054] In an embodiment, the lighting information includes at least one of the following characteristics of light in the source video frame:
[0055] - Lighting direction;
[0056] - Lighting intensity;
[0057] - Lighting position;
[0058] - Color temperature;
[0059] - Spectrum;
[0060] - Lighting level; and
[0061] - Environmental spherical harmonics.
[0062] According to a third aspect of the present application, there is provided a bitstream of encoded processed data generated by one of the methods according to the second aspect of the present application.
[0063] In an embodiment, the illumination information is formatted in (or included in) a supplementary enhancement information (SEI) message.
[0064] According to a fourth aspect of the present application, there is provided an AR device (or AR equipment) for AR content rendering. The AR device includes means for performing any of the methods according to the first aspect of the present application.
[0065] According to a fifth aspect of the present application, there is provided a processing device (or processing equipment) for implementing AR content rendering. The processing device includes means for performing any of the methods according to the second aspect of the present application.
[0066] According to a sixth aspect of the present application, there is provided a computer program product including instructions that, when the program is executed by one or more processors, cause the one or more processors to perform any of the methods according to the first aspect of the present application.
[0067] According to a seventh aspect of the present application, there is provided a non-transitory storage medium (or storage medium) that carries instructions for program code for performing any of the methods according to the first aspect of the present application.
[0068] According to an eighth aspect of the present application, there is provided a computer program product including instructions that, when the program is executed by one or more processors, cause the one or more processors to perform any of the methods according to the second aspect of the present application.
[0069] According to a ninth aspect of the present application, there is provided a non-transitory storage medium (or storage medium) that carries instructions for program code for performing any of the methods according to the second aspect of the present application.
[0070] This application allows for the elegant and seamless integration of virtual objects into the video frames of real-world scenarios, thereby providing users with a realistic and immersive AR experience, regardless of the resources available at the AR device level and / or the resources for communication between the AR device and the outside. Thus, even with an AR device having limited processing resources (such as a mobile device (e.g., AR on Google)), optimal rendering quality of AR content can be achieved. In particular, for realism, realistic lighting can be efficiently reconstructed in the AR video frames. For example, accurate lighting estimation can be achieved in real time without latency or with limited latency.
[0071] The specific nature of at least one embodiment among the embodiments, as well as other objectives, advantages, features, and uses of the at least one embodiment among the embodiments, will become more apparent from the following description of examples in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Now, examples will be referred to in a manner of reference to the accompanying drawings showing embodiments of the present application, and in which:
[0073] Figure 1 An AR video frame generated according to the related art is shown;
[0074] Figure 2 A schematic block diagram showing the steps of a method 100 for encoding a video frame according to the related art is shown;
[0075] Figure 3 A schematic block diagram showing the steps of a method 200 for decoding a video frame according to the related art is shown;
[0076] Figure 4 A first method for rendering AR content is illustrated;
[0077] Figure 5 A method for rendering AR content according to an embodiment of the present application and a method for implementing AR content rendering are illustrated;
[0078] Figure 6 A method for rendering AR content according to an embodiment of the present application and a method for implementing AR content rendering are shown;
[0079] Figure 7 A real-world scenario and lighting estimation according to an embodiment of the present application are illustrated;
[0080] Figure 8 Each component of the lighting according to an embodiment of the present application is illustrated;
[0081] Figure 9 An AR video frame obtained according to an embodiment of the present application is shown;
[0082] Figure 10 A schematic block diagram showing an example of a system in which various aspects and embodiments are implemented; and
[0083] Figure 11 A schematic block diagram showing an example of a system in which various aspects and embodiments are implemented.
[0084] Like or identical elements are denoted by the same reference numerals. Detailed Description
[0085] Embodiments will be described more fully hereinafter with reference to the accompanying drawings, in which examples of at least one embodiment are depicted. However, the embodiments may be implemented in many alternative forms and should not be construed as limited to the examples set forth herein. Thus, it should be understood that the embodiments are not intended to be limited to the particular forms disclosed. On the contrary, this application is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of this application.
[0086] At least one of these aspects generally relates to video picture encoding and decoding, other aspects generally relate to transmitting the provided or encoded bitstream, and one of the other aspects relates to receiving / accessing the decoded bitstream.
[0087] At least one of these embodiments is described for encoding / decoding video pictures, but extends to encoding / decoding video pictures (picture sequences) because each video picture can be encoded / decoded sequentially as described hereinafter.
[0088] Furthermore, the at least one embodiment is not limited to MPEG standards such as AVC (ISO / IEC 14496-10, Advanced Video Coding for generic audio-visual services, ITU-T Recommendation H.264, https: / / www.itu.int / rec / T-REC-H.264-202108-P / en), EVC (ISO / IEC 23094-1 Basic Video Coding), HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en), VVC (ISO / IEC 23090-3 Versatile Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-I / en), but can be applicable to other standards and recommendations such as AV1 (AOMedia Video 1, http: / / aomedia.org / av1 / specification / ) for example. The at least one embodiment can be applicable to existing or future-developed standards and recommendations and extensions of any such standards and recommendations. Unless otherwise indicated or technically precluded, the aspects described in this application can be used alone or in combination.
[0089] A pixel corresponds to the smallest display unit on the screen and can consist of one or more light sources (1 for a monochrome screen or 3 or more for a color screen).
[0090] A video picture (also referred to as a frame or picture frame) includes at least one component (also referred to as a picture component or channel) determined by a specific picture / video format, and the at least one component specifies all information related to pixel values and all information that can be used by a display unit and / or any other device to display and / or decode video picture data related to the said video picture.
[0091] A video picture includes at least one component typically represented in the form of an array of samples.
[0092] A monochrome video picture includes a single component, while a color video picture may include three components.
[0093] For example, when the picture / video format is the well-known (Y, Cb, Cr) format, a color video picture may include a luminance (or brightness) component and two chrominance components, or when the picture / video format is the well-known (R, G, B) format, a color video picture may include three color components (one for red, one for green, and one for blue).
[0094] Each component of the video picture may include a number of samples related to the number of pixels of the screen on which the video picture is intended to be displayed. For example, the number of samples included in the component may be the same as the number of pixels of the screen on which the video picture is intended to be displayed, or a multiple (or fraction) thereof.
[0095] The number of samples included in the component may also be a multiple (or fraction) of the number of samples included in another component of the same video picture.
[0096] For example, in the case where the video format includes a luminance component and two chrominance components (such as the (Y, Cb, Cr) format), depending on the color format considered, the chrominance components may contain half the number of samples in width and / or height relative to the luminance component.
[0097] A sample is the smallest visual information unit that makes up a component of a video picture. The sample value can be, for example, a luminance or chrominance value, or it can be a color value in the (R, G, B) format.
[0098] A pixel value is the value of a screen pixel. For a monochrome video picture, the pixel value can be represented by one sample, and for a color video image, the pixel value can be represented by multiple co-located samples. Co-located samples associated with a pixel refer to the samples corresponding to the position of the pixel in the screen.
[0099] A video picture is generally regarded as a set of pixel values, with each pixel represented by at least one sample.
[0100] A block of a video picture is a set of samples of a component of the video picture. When the picture / video format is the well-known (Y, Cb, Cr) format, a block of at least one luminance sample or a block of at least one chrominance sample can be considered, or when the picture / video format is the well-known (R, G, B) format, a block of at least one color sample can be considered.
[0101] This at least one embodiment is not limited to a specific picture / video format.
[0102] Figure 2 and Figure 3Provides an overview of video encoding / decoding methods used in current video standard compression systems (such as VVC). As further indicated below, for encoding / decoding purposes, these video encoding / decoding techniques or any suitable variants may be used in this application. However, this application is not limited to these embodiments.
[0103] Figure 2 Shows a schematic block diagram of the steps of method 100 for encoding video picture VP according to the related art.
[0104] In step 110, the video picture VP is partitioned into sample blocks, and the partition information data is signaled into the bitstream. Each block includes samples of one component of the video picture VP. Thus, these blocks include samples defining each component of the video picture VP.
[0105] For example, in HEVC, the picture is divided into coding tree units (CTUs). Each CTU can be further subdivided using quadtree partitioning, where each leaf of the quadtree is represented as a coding unit (CU). Then, the partition information data can include data defining the CTU and the quadtree subdivision of each CTU.
[0106] Then, each block (CU) of samples (simply referred to as a block) is encoded within the encoding loop using an intra or inter prediction coding mode. Hereinafter, "in-loop" is defined to be also assigned to steps, functions, etc. implemented within the loop (i.e., the encoding loop at the encoding stage or the decoding loop at the decoding stage).
[0107] Intra prediction (step 120) consists of predicting the current block by means of a prediction block based on already encoded, decoded, and reconstructed samples located around the current block within the picture (usually at the top and left of the current block). Intra prediction is performed in the spatial domain.
[0108] In the inter - frame prediction mode, motion estimation (step 130) and motion compensation (135) are performed. Motion estimation searches for candidate reference blocks that are good predictors of the current block in one or more reference video pictures used for predictive coding of the current video picture. For example, a good predictor of the current block is a predictor similar to the current block. The output of the motion estimation step 130 is one or more motion vectors and one or more reference picture indices associated with the current block. Next, motion compensation (step 135) obtains the predicted block by means of the (one or more) motion vectors and one or more reference picture indices determined by the motion estimation step 130. Basically, the block belonging to the selected reference picture and pointed to by the motion vector can be used as the predicted block of the current block. Additionally, since the motion vectors are represented as fractions of integer pixel positions (referred to as sub - pixel precision motion vector representation), motion compensation typically involves spatial interpolation of some reconstructed samples of the reference picture to calculate the predicted block samples.
[0109] The prediction information data is signaled into the bitstream. The prediction information may include the prediction mode, the prediction information coding mode, the intra - frame prediction mode, or the (one or more) motion vectors and (one or more) reference image indices, and any other information for obtaining the same predicted block at the decoding side.
[0110] Method 100 optimizes the rate - distortion trade - off by considering, for example, the encoding of the prediction residual block calculated by subtracting the candidate predicted block from the current block, and the signaling of the prediction information data required to determine the candidate predicted block at the decoding side, thereby selecting one of the intra - frame mode or the inter - frame coding mode.
[0111] Generally, the best prediction mode is given as the prediction mode of the best coding mode p* for the current block, given by the following equation:
[0112]
[0113] where P is the set of all candidate coding modes for the current block, p represents a candidate coding mode in the set, and RD cost (p) is the rate - distortion cost of the candidate coding mode p, typically expressed as:
[0114] RD cost(p) = D(p)+λ.R(p).
[0115] D(p) is the distortion between the current block and the reconstructed block obtained after encoding / decoding the current block with the candidate coding mode p, R(p) is the rate cost associated with encoding the current block with the coding mode p, λ is a rate constraint representing the encoding of the current block, and is a Lagrange parameter typically calculated according to the quantization parameter used for encoding the current block.
[0116] The current block is typically encoded according to a prediction residual block PR. More precisely, for example, the prediction residual block PR is calculated by subtracting the best prediction block from the current block. Then, the prediction residual block PR is transformed (step 140) by using, for example, a DCT (Discrete Cosine Transform) or DST (Discrete Sine Transform) type transform or any other suitable transform, and the obtained transformed coefficient block is quantized (step 150).
[0117] In a variant, method 100 may also skip the transform step 140 and apply quantization directly to the prediction residual block PR (step 150) according to a so-called transform-skip coding mode.
[0118] The quantized transformed coefficient block (or quantized prediction residual block) is entropy encoded into a bitstream (step 160).
[0119] Next, as part of an encoding loop, the quantized transformed coefficient block (or quantized residual block) is dequantized (step 170) and inverse transformed (180) (or not inverse transformed), resulting in a decoded prediction residual block. Then, the decoded prediction residual block and the predicted block are combined, typically summed, which provides a reconstructed block.
[0120] Other information data may also be entropy encoded in step 160 to encode the current block of the video picture VP.
[0121] A loop filter (step 190) may be applied to the reconstructed picture (including the reconstructed block) to reduce compression artifacts. The loop filter may be applied after all picture blocks have been reconstructed. For example, they include a deblocking filter, sample adaptive offset (SAO), or an adaptive loop filter.
[0122] The reconstructed block or the filtered reconstructed block forms a reference picture, which may be stored in a decoded picture buffer (DPB) such that it can be used as a reference image for encoding the next current block of the video picture VP or the next video picture to be encoded.
[0123] Figure 3 A schematic block diagram showing the steps of a method 200 for decoding a video picture VP according to the related art is shown.
[0124] In step 210, partition information data, prediction information data, and a quantized transformed coefficient block (or quantized residual block) are obtained by entropy decoding the bitstream of the encoded video picture data. For example, this bitstream has been generated according to method 100.
[0125] Other information data may also be entropy decoded from the bitstream for decoding the current block of the video picture VP.
[0126] In step 220, the reconstructed picture is partitioned into current blocks based on the partitioning information. Each current block is entropy decoded from the bitstream within the decoding loop. Each decoded current block is either a quantized transform coefficient block or a quantized prediction residual block.
[0127] In step 230, the current block is dequantized and possibly inverse transformed (step 240) to obtain a decoded prediction residual block.
[0128] On the other hand, prediction information data is used to predict the current block. A predicted block is obtained through its intra prediction (step 250) or its motion-compensated temporal prediction (step 260). The prediction processing performed at the decoding side is the same as the prediction processing performed at the encoding side.
[0129] Next, the decoded prediction residual block and the predicted block are then combined, typically by summing, which provides the reconstructed block.
[0130] In step 270, a loop filter can be applied to the reconstructed picture (including the reconstructed blocks), and the reconstructed blocks or the filtered reconstructed blocks form a reference picture, which can be stored in the decoded picture buffer (DPB), as discussed above ( Figure 1 ).
[0131] Generally, the present application relates to the rendering of AR content and to data processing for implementing the rendering of AR content. In particular, the present application relates to an AR device for performing the rendering of AR content and a processing device for implementing the rendering of AR content, wherein the AR device and the processing device can cooperate with each other to implement the rendering of one or more AR video pictures.
[0132] The concept of the present application is to incorporate a relighting effect into the AR video pictures rendered by an AR device to improve the realism and rendering quality of the rendered AR content. To this end, light estimation is performed by an external processing device (i.e., a processing device external to the AR device) to determine light information based on the video pictures. The light information is transmitted to the AR device such that the AR device can aggregate the video pictures and volume information into an aggregated video picture, and obtain an AR video picture by incorporating a relighting effect based on the received light information into the aggregated video picture. In particular, based on the light information, a relighting effect associated with one or more virtual objects defined by the volume information can be applied within the aggregated video picture.
[0133] In an embodiment, there is provided a method for performing AR content rendering implemented by an AR device, the method including: transmitting encoded source video frame data representing a video frame of a real-world scene to a processing device external to the AR device; receiving lighting information from the processing device; and generating an AR video frame by: obtaining an aggregated video frame by aggregating the video frame and volume information representing at least one virtual object; and combining a relighting effect associated with the at least one virtual object into the aggregated video frame based on the received lighting information. The AR device may further render (e.g., display) the AR video frame.
[0134] In an embodiment, there is provided a method for implementing AR content rendering implemented by a processing device, the method including: receiving encoded video frame data representing a video frame of a real-world scene from an AR device external to the processing device; obtaining the video frame of the real-world scene by decoding the received encoded video frame data; determining lighting information by performing lighting estimation based on the video frame; and transmitting the lighting information to the AR device for implementing AR rendering of the video frame.
[0135] In the present application, therefore, lighting estimation that requires a large amount of processing resources is entrusted to a processing device external to the AR device responsible for rendering the final AR video frame, or remotely executed by the processing device.
[0136] The present application allows virtual objects to be gracefully and smoothly integrated into the video frame of a real-world scene, thereby providing a realistic and immersive AR experience for the user, regardless of the resources available at the AR device level. Thus, even with an AR device having limited processing resources (such as a mobile device (e.g., AR on Google, etc.)), optimal rendering quality of AR content can be achieved. In particular, for realism, realistic lighting can be efficiently reconstructed in the AR video frame.
[0137] For example, accurate lighting estimation can be achieved in real time without latency or with limited latency. In particular, only part of the processing can be entrusted to be executed external to the AR device, i.e., only part of the processing operations required for rendering AR content can be entrusted to an external processing device, including lighting estimation for determining lighting information. However, other processing operations (such as aggregation of the video frame and volume information and aggregation of the relighting effect (as discussed above)) can be performed locally by the AR device to limit the communication latency between the AR device and the processing device, thereby maximizing AR rendering efficiency.
[0138] Thus, even though the available bandwidth for communication and the local resources in the AR device are limited, an optimal compromise can be achieved through this hybrid processing distribution between the AR device and the processing device to allow for efficient and realistic AR content rendering.
[0139] Figure 4 A first method is illustrated by way of example, in which the AR device 400 performs a lighting estimation (402) based on a video frame PR1 (also referred to as the source video frame) of a real-world scene (e.g., a physical object OB1 in any form). In this real-world scene, one or more physical light sources SC1 can project light that changes the physical object OB1 represented by the video frame PR1. For example, the video frame PR1 can represent a physical object OB1 associated with a real shadow SH1 as part of a real-world scene, where the real shadow SH1 is caused by light from the light source SC1 being projected onto the physical object OB1. Variants are also possible, where the video frame PR1 represents a physical object OB1 without an apparent real shadow SH1.
[0140] As Figure 4 shown, the AR device 400 can obtain lighting information (or light information) by performing the lighting estimation 402 based on the video frame PR1. Based on the lighting information and the video frame PR1, the AR device 400 can further perform a lighting reconstruction (404) to obtain an AR video frame PR2. More specifically, the AR device 400 can obtain an aggregated video frame by aggregating the volume information representing the virtual object OB2 with the video frame PR1 (such that the virtual object OB2 is inserted into the source video frame PR1), and perform the lighting reconstruction 404 by combining the relighting effect (or lighting effect) associated with the virtual object OB2 into the aggregated video frame based on the lighting information obtained in 402, thereby obtaining the AR video frame PR2. Due to the lighting reconstruction 404, the AR video frame PR2 can, for example, include a virtual shadow SH2 of the virtual object OB2.
[0141] It can be seen that Figure 4 all the processing steps in the method are performed locally by the AR device 400, and all these processing steps include the lighting estimation 402, the aggregation of the volume information with the video frame PR1, and the combination of the relighting effect. However, as described previously, this method may be difficult to implement or may produce poor results, especially when the AR device 400 has limited local (or embedded) resources, or at least when the local resources it contains are insufficient to perform an appropriate lighting estimation.
[0142] Figure 5An embodiment of the present application is shown, in which an AR device 500 configured to perform AR content rendering cooperates with an external processing device 502 configured to enable the AR device 500 to perform AR content rendering. It can be seen that Figure 5 The embodiment of Figure 4 differs from the method of Figure 5 in that the light estimation 404 in
[0143] is delegated to the external processing device 502, that is, remotely executed by the external processing device 502. Figure 4 More particularly, the AR device 500 obtains a video frame PR1 representing a real-world scene in any suitable manner (as previously described with reference to Figure 4 ), and transmits the video frame PR1 to the processing device 502, which determines the light information by performing light estimation 402 based on the video frame PR1. Then, the processing device 502 transmits the light information to the AR device 500, which performs light reconstruction 404 as previously described based on the light information. In particular, the AR device 500 can obtain an aggregated video frame by aggregating the video frame PR1 and the volume information representing at least one virtual object OB2, and generate an AR video frame PR3 by combining the relighting (or lighting) effect associated with the virtual object OB2 into the aggregated video frame based on the received light information.
[0144] As will be described in more detail in the following embodiments, the relighting effect combined into the aggregated video frame to obtain the AR video frame PR3 can be adjusted according to each case.
[0145] Other aspects and advantages of the present application will be described in specific embodiments with reference to the accompanying drawings.
[0146] It should be understood that the present application can be applied in the same way to one or more source video frames PR1 to obtain and deliver one or more AR video frames PR3. In addition, the real-world scene represented by the source video frame PR1 can have various properties and can include one or more physical objects OB1 with or without (one or more) real shadows SH1, or may not include any physical objects themselves (e.g., only the background). Similarly, the AR video frame PR3 can be obtained by aggregating the (source) video frame with the volume information representing one or more virtual objects OB2 (also referred to as volumetric objects OB2).
[0147] Now, specific embodiments of the present application will be described with reference to Figures 6 to 11 The embodiments described further below can be Figure 5Exemplary implementation of an embodiment. In the following embodiments, it is assumed that the AR video frame PR3 is obtained based on the source video frame PR1 (also referred to as video frame PR1) of a real-world scenario. The source video frame PR1 represents a real-world scenario and may include one or more physical (real) objects of the scenario. However, it should be understood that the present application can be applied in the same manner to multiple source video frames PR1, such as the video frames PR1 of a video data stream. In particular, the following embodiments can be applied to rendering an AR video frame as a video sequence.
[0148] Figure 6 FIG. shows a schematic block diagram of steps S2 to S18 of method 600 for performing AR content rendering implemented by the AR device DV1 and steps S40 to S46 of method 640 for implementing AR content rendering implemented by the processing device DV2.
[0149] The AR device DV1 and the processing device DV2 can take various forms as further described below and can respectively correspond to, for example Figure 5 devices 500 and 502. Exemplary implementations of the AR device DV1 and the processing device DV2 are described below with reference to Figures 10 to 11 this.
[0150] The AR device DV1 and the processing device DV2 cooperate with each other to allow AR content rendering. To this end, a communication channel CN1 can be established between the AR device DV1 and the processing device DV2. In other words, the processing device DV2 can be configured as an external tethered device relative to the AR device DV1.
[0151] In the obtaining step S4 ( Figure 6 ), the AR device DV1 obtains the source video frame PR1 representing the real-world scenario (or real-world environment). The real-world scenario can have various properties depending on each case, and the source video frame PR1 can have any appropriate format.
[0152] For example, the source video frame PR1 can be obtained by capturing (S4) the real-world scenario as the source video frame PR1 by means of at least one capture sensor embedded (or locally included) in the AR device DV1. For example, a camera can be used to capture the video frame PR1. In a variant, the AR device DV1 can receive (S4) the source video frame PR1 from a capture device DV3 ( Figure 10 ) that captures the real-world scenario as the source video frame PR1 (e.g., by means of a camera or any other capture sensor). In a variant, the AR device DV1 retrieves (S4) the source video frame PR1 from a local memory (i.e., a memory included in the AR device DV1).
[0153] In encoding step S6, the AR device DV1 encodes the source video frame PR1 of the real-world scene into encoded video frame data DT2. For this purpose, the AR device DV1 can use any suitable encoding technique, such as the encoding method (Method 100) described with reference to Figure 2 the encoding method (Method 100) described with reference to
[0154] In transmission step S8, the AR device DV1 transmits the encoded video frame data DT2 to the processing device DV2. In this example, the transmission S8 is performed in the form of an encoded bitstream BT1, for example, via a communication channel (or communication link) CN1 that can be a 5G link or any other suitable channel, such as using a communication network. 5G provides low-latency communication.
[0155] In an embodiment, in encoding step S6, the AR device DV1 also encodes the spatial mapping information DT3 together with the encoded video frame data DT2. In this case, therefore, the spatial mapping information DT3 is sent (S8) to the processing device DV2 together with or along with the encoded video frame data DT2. The spatial mapping information DT3 defines a 3D coordinate system and can be used such that the AR device DV1 and the processing device DV2 share the same 3D spatial coordinate system, thus allowing for consistent light estimation.
[0156] In receiving step S40, the processing device DV2 receives the encoded video frame data DT2. Then, the processing device DV2 obtains (S42) the source video frame PR1 by decoding the received encoded video frame data DT2. The decoding can be performed in any suitable manner, such as by the decoding method (Method 200) described with reference to Figure 3 the decoding method (Method 200) described with reference to
[0157] The processing device DV2 may also obtain (S42) the spatial mapping information DT3 (if any) by decoding the encoded source video frame data DT2.
[0158] In determination step S44, the processing device determines the lighting information (or light information) DT4 by performing light estimation based on the source video frame PR1. The light estimation can be performed according to any known method, such as by according to (Google Play Services for AR) or implementing a light estimation algorithm. For example, a light estimator unit can be used, such as (Generative Adversarial Network for Shadow Generation), (Illumination Estimation Framework for Real-Time Photorealistic Augmented Reality on Mobile Devices), (Light Source Estimation for Augmented Reality Using Deep Learning), etc.
[0159] The illumination information DT4 determined in S44 defines the characteristics of the light to which the real-world scene represented by the source video frame PR1 is exposed. In other words, the illumination information DT4 characterizes the light emitted from one or more physical light sources of the real-world scene.
[0160] Figure 7 is a schematic diagram showing by way of example a physical light source SC1 (i.e., a light source existing in the real world), which projects (real) light onto a physical object OB1 existing in the real-world scene captured by the source video frame PR1. By performing the illumination estimation S44, one or more characteristics of the light emitted from the physical light source SC1 can be determined.
[0161] The illumination information DT4 determined in S44 may include at least one (or any combination) of the following light characteristics: illumination direction, illumination intensity, illumination position, color temperature, spectrum, illumination level, and ambient spherical harmonics.
[0162] The illumination information DT4 may be determined in S44 based on the spatial mapping information DT3 (if any) received from the AR device DV1 as part of the encoded video frame data DT2. In particular, as Figure 7 shown, the illumination position 702 of the physical light source SC1 may be defined in 3D spatial coordinates (representing the real-world scene), for example, using the 3D spatial coordinates defined by the spatial mapping information DT3 (if any) received (S40) from the AR device DV1. The physical light source SC1 may be part of an illumination device 700 (FIG. 700), which illuminates the real-world scene defined by the source video frame PR1.
[0163] Ambient spherical harmonics may represent the spherical harmonic coefficients for ambient illumination.
[0164] For example, the illumination information DT4 may define a main directional light representing the main light source SC1 in the real-world scene and ambient spherical harmonics representing the remaining ambient light energy in the real-world scene. Determining the main directional light may be used, for example, for casting shadows, as described further below.
[0165] In an embodiment, in the determination step S44, the processing device DV2 also determines ambient map data DT5 based on the source video frame PR1. These ambient map data DT5 define a 3D view of the estimated illumination as an image. However, embodiments in which such ambient map data are not used are also possible.
[0166] More specifically, the environmental map data DT5 defines an environmental map, that is, a mapped image representing an omnidirectional view of the environmental illumination of a 3D scene seen from a specific 3D position (representing the 3D scene in the source video frame PR1). Each pixel in the mapped image corresponds to a 3D direction, and the data stored in the pixel represents the amount of light arriving from this direction. In 3D rendering applications, the environmental map can be used for image-based lighting techniques that approximate how an object is illuminated by its surrounding environment.
[0167] An environmental map having sufficient dynamic range to represent even the brightest light sources in the environment is sometimes referred to as a "light probe image".
[0168] For example, the environmental map data DT5 can represent the mapped image as a cubemap (or environmental cubemap, sometimes also referred to as an HDR (high dynamic range) cubemap), which is a way of encoding 360-degree light information. In a cubemap, the environmental illumination is projected onto the six faces of an axis-aligned cube. Then, the individual faces of the cube are arranged in a 2D image.
[0169] For example, (Google Play services for AR) uses a cubemap to represent environmental illumination. By using a cubemap, good performance for efficient encoding and decoding of 360 video sequences can be achieved. However, any other 360 projection can be used, such as equirectangular projection, fisheye projection, cylindrical projection, etc. For example, the environmental map can also be represented as a "latitude-longitude map", also known as equirectangular projection in the 360 video domain.
[0170] The environmental map data DT5 in the form of a cubemap as described above can be advantageously used to recreate realistic lighting effects on virtual objects. It can be particularly used to render the reflections of shiny metal objects in AR content, as described further below.
[0171] In an embodiment, the lighting information DT4 determined in S44 includes a main directional light representing the main light source SC1 in the real-world scene, environmental spherical harmonics representing the remaining environmental light energy in the real-world scene, and the environmental map data DT5 as previously described. Figure 8 The effects of these three components are illustrated in a specific example.
[0172] To maximize the AR content rendering quality, each time a new source video frame PR1 (or video frame) is provided by the AR device DV1, the processing device DV2 can refresh (determine) the above three components. Therefore, the light direction information and the environmental spherical harmonics can become a timed sequence of metadata, while the environmental map data DT5 can become a timed sequence of textures, i.e., a video sequence.
[0173] In the transmission step S46, the processing device DV2 transmits the lighting information DT4 and also possibly the environmental map data DT5 (if any) to the AR device DV1 to enable AR content rendering. This transmission can be performed, for example, via a communication channel CN1 established between the AR device DV1 and the processing device DV2 (e.g., via a 5G link).
[0174] In an embodiment, the transmission S46 is performed as a bitstream BT2 of encoded video frame data in an encoded form, which bitstream BT2 includes the lighting information DT4 and may also include the environmental map data DT5 (if any). The encoding can be performed in any suitable manner, such as the encoding method (Method 100) described previously with reference to Figure 2 the description.
[0175] In an embodiment, before transmitting S46 to the AR device DV1, the processing device DV2 inserts (S46) the lighting information DT4 as metadata into the bitstream BT2 of the encoded video frame data. For example, the lighting information DT4 as metadata is carried in the SEI (Supplemental Enhancement Information) message of the bitstream BT2. The specific implementation of carrying the lighting information DT4 in the bitstream BT2 will be further described below.
[0176] In an embodiment, the environmental map data DT5 determined in S44 is also encoded (S46) as an encoded image in the encoded bitstream BT2 of the video frame data before being transmitted in S46 to the AR device DV1.
[0177] It should be understood that since the AR device DV1 already has the source video frame PR1 (acquisition step S4), the processing device DV2 does not return the source video frame PR1 to the AR device DV1 to avoid unnecessary data exchange and communication latency. When waiting to receive the bitstream BT2, the AR device DV1 can store the source video frame PR1 in its local memory.
[0178] In the receiving step S10, the AR device DV1 receives illumination information DT4 carried, for example, by an encoded bitstream BT2 of encoded video frame data, and possibly also environmental map data DT5 (if any). More specifically, the AR device DV1 can obtain the illumination information DT4 and the environmental map data DT5 by decoding the bitstream BT2 of the encoded video frame data received from the processing device DV2. The decoding can be performed in any suitable manner, for example, according to the decoding method (method 200) described in the previous reference Figure 3 as described.
[0179] In the generating step S12, the AR device DV1 generates an AR video frame PR3 by performing an aggregating step S14 and a combining step S16, which will be further described in the following embodiments.
[0180] More specifically, in the aggregating step S14, an aggregated video frame PR2 is obtained by aggregating a source video frame PR1 and volume information DT1 representing at least one virtual (or volumetric) object OB1. To this end, the processing device DV2 can parse the illumination information DT4 (and any other possible information provided by the processing device DV2) to determine how the source video frame PR1 of the real-world scene is illuminated.
[0181] In an embodiment of the present application, it is assumed that the volume information DT1 represents a virtual object OB2 ( Figure 6 ), and this virtual object OB2 is inserted (or integrated) into the source video frame PR1 to form the aggregated video frame PR2. The virtual object OB2 can be defined in any suitable volume format by the volume information DT1.
[0182] The aggregation S14 causes the virtual object OB2 to be superimposed on (or integrated into) the source video frame PR1, thereby generating the aggregated video frame PR2 that forms the AR content. In particular, the virtual object OB2 constitutes enhanced content superimposed on top of the source video frame PR1 representing the real-world environment perceived by the user.
[0183] To this end, the AR device DV1 can obtain (S2) the volume information DT1 in any suitable manner. In an embodiment, the AR device DV1 receives (S2) the volume information DT1, for example, via a second communication CN2 independent of the first communication channel CN1, independent of (without) the illumination information DT4 (e.g., independent of the encoded bitstream BT2), from a content server DV4 ( Figure 6 and 10 ). The volume information DT1 can be received (S2), for example, from a content server DV4 different from the processing device DV2.
[0184] In an embodiment, the AR device DV1 receives (S2) volume information DT1 and lighting information DT4 which is part of the bitstream BT2 of the encoded video frame data received in S10. In this case, it is thus the processing device DV2 that transfers the volume data DT1 to the AR device DV1.
[0185] In an embodiment, the content server DV4 transfers the volume data DT1 to the processing device DV2, and then the processing device DV2 encodes the received volume data DT1 and lighting information DT4 (and possibly also environmental map data DT5, if any) into the bitstream BT2 before transferring it to the AR device DV1. In this variant, the processing device DV2 thus acts as a relay for transferring the volume information DT1 to the AR device DV1.
[0186] In an embodiment, the AR device DV1 retrieves (S2) volume information DT1 from the local memory of the AR device DV1.
[0187] Still as Figure 6 shown, in the above-mentioned combining step S16, the AR device DV1 combines (or integrates) a relighting effect (or lighting effect, or reconstructed lighting) associated with the virtual object OB2 into the aggregated video frame PR2 based on the received lighting information DT4 and possibly also based on environmental map data DT5 (if any), thereby obtaining the AR video frame PR3. In other words, the AR video frame PR3 is obtained by applying (S16) a relighting effect to the aggregated video frame PR2. These relighting effects are intended to change the lighting within the AR content of the aggregated video frame PR2 to provide a more immersive and realistic experience to the user.
[0188] The relighting effect can have different types and can particularly include at least one of the following (or any combination thereof): shadow effect; ambient lighting illumination; specular highlights; and light reflection.
[0189] The relighting effect combined (S16) into the aggregated video frame PR2 is determined by applying the lighting information DT4 and possibly environmental map data DT5 (if any) to the virtual object OB2 inserted into the aggregated video frame PR2. The relighting effect is associated with the virtual object OB2 in the sense that it changes the lighting of the virtual object OB2 itself and / or changes the lighting of another area of the aggregated video frame PR2 due to the emission of virtual shadows from the virtual object OB2.
[0190] In an embodiment, the combination S16 of the relighting effect includes any one (or both) of the following: superimposing the reconstructed lighting on the virtual object OB2 inserted into the aggregated video frame PR2; and superimposing the shadow effect of at least one reconstructed shadow SH2 representing the virtual object OB2 on the aggregated video frame PR2. The reconstructed lighting and shadow effects are determined based on the lighting information DT4 (and possibly also based on the environmental map data DT5, if any).
[0191] Figure 9 Illustrated is the generation of the AR video frame PR3 according to an embodiment. The virtual object OB2 (rocket) is inserted into the source video frame PR1, and the relighting effect is applied to the inserted virtual object OB2, i.e., light reflections, etc. are added to the virtual object (e.g., on the upper end of the rocket), and the virtual shadow SH2 emitted from the virtual object OB2 is added to the background of the video frame. The virtual shadow SH2 represents the virtual exposure of the virtual object OB2 to the real light within the real-world scene, i.e., the light projected by the physical source SC1.
[0192] In the rendering step S18, the AR device DV1 renders the AR video frame PR3 by any suitable means. The rendering can be performed by displaying the AR video frame PR3 (or by using any 3D projection, etc.). In an embodiment, the AR device DV1 uses the local display unit embedded in the AR device DV1 to display the AR video frame PR3. The display unit can take any suitable form, such as AR Google Glass, a head-mounted headset for rendering AR content, a screen, etc.
[0193] The AR device DV1 can perform any suitable calculations to render the AR video frame PR3 as AR content. In an embodiment, the AR device DV1 can send the AR video frame PR3 to an external display unit for display.
[0194] This application may cover various adaptations of methods 600 and 640, some of which will be described by way of examples below.
[0195] In an embodiment, the AR device DV1 obtains (S4, Figure 6 ) the lighting information DT6, which defines the lighting (illumination) applied by the lighting device to capture the real-world scene as the source video frame PR1. More particularly, a lighting device (lamp or other light source) can be used to physically illuminate the real-world scene captured as the source video frame PR1. The lighting information DT6 characterizes how the real-world scene is illuminated by one or more physical light sources of the lighting device. The lighting device can be Figure 7 the illustrated lighting device 702, or a part of the capture device DV3 further described below with reference to Figure 10 the capture device DV3.
[0196] The AR device DV1 can further encode the illumination information DT6 and the source video frame PR1 (S6, Figure 6 ), and then transmit (S8) the encoded source video frame data DT2 including the encoded source video frame PR1 together with the encoded illumination information DT6 to the processing device DV2, for example, as part of an encoded bitstream BT1. In other words, the illumination information DT6 can be inserted into the bitstream BT1 before being transmitted to the processing device DV2. In a variant, the illumination information DT6 is received by the processing device DV2 independently of the encoded bitstream BT1. In any case, the processing device DV2 can obtain the illumination information DT6, for example, by decoding (S42) the encoded bitstream BT1 received from the AR device DV1, and determine the illumination information DT4 based on the illumination information DT6 (and the source video frame PR1 and possibly also the spatial mapping information DT3, if any) in the light estimation step S44.
[0197] By using the illumination information DT6 in the light estimation step S44, more accurate and efficient light estimation can be achieved.
[0198] This application allows for an improvement in the realism and rendering quality of the rendered AR content by only delegating (or externalizing) part of the processing work to the processing device DV2, particularly for light estimation which requires a large amount of processing and power resources and is not always locally supported by the AR device. However, to avoid latency and ensure efficient AR content rendering, a compromise is achieved by locally performing as much of the required processing work as possible (including the aggregation of real and virtual content and relighting reconstruction) at the AR device level. While the solution might lie in remotely performing all processing operations (including aggregation and relighting) by an external processing device DV2, this could compromise the quality of the AR content rendering, significantly because a large amount of data (corresponding to the aggregated video frames including the relighting effects) must be transmitted from the processing device DV2 to the AR device DV1, thus requiring a large amount of bandwidth and communication resources. Additionally, the additional processing also requires additional latency. Due to this application, since the processing device DV2 does not need to return data corresponding to the source video frame PR1, a limited amount of data is transmitted from the processing device DV2 to the AR device DV1.
[0199] In an embodiment, the AR device DV1 determines its ability to communicate with the processing device DV2, for example, via a communication channel CN1 ( Figure 6 and 10)。If the communication capacity (e.g., available bandwidth) is higher than a threshold, the AR device DV1 implements method 600 to delegate the light estimation to the processing device DV2. Otherwise, method 600 is not executed. If the communication capacity is lower than the threshold, the AR device DV1 can perform the light estimation S44 locally, as previously described ( Figures 6 to 9 ).
[0200] In addition, as previously described, according to an embodiment of the present application, the light information DT4 can be transmitted (S46, Figure 6 ) from the processing device DV2 to the AR device DV1 to enable AR content rendering. The light information DT4 can be signaled as metadata into the bitstream BT2 of the encoded video frame data transmitted to the AR device DV1. An exemplary syntax for signaling the light information may be as follows:
[0201]
[0202]
[0203] Where:
[0204] li_lighting_id specifies an identification number that can be used to identify a light or a light group.
[0205] li_sources_count_minus1 plus 1 represents the number of light sources in the current SEI message.
[0206] li_light_info_present_flag[i] being equal to 1 indicates that the information of the i-th light source exists in the bitstream. li_light_info_present_flag[i] being equal to 0 indicates that the information of the i-th light source does not exist in the bitstream.
[0207] li_light_type_id[i] represents the light type of the i-th light source, as listed in the following table. The value of li_light_type_id[i] should be in the range of 0 to 5, including the endpoint values. Other values are reserved for future use.
[0208]
[0209] li_color_temperature_idc[i], when in the range of 800 to 12200 (including the endpoint values), specifies the color temperature of the i-th light source in Kelvin. When li_color_temperature_idc[i] is not in the range of 800 to 12200 (including the endpoint values), the color temperature is unknown or not specified, or is specified in other ways.
[0210] In an embodiment, the intensity of the light source is indicated, expressed, or represented in lumens or lumens per watt or lux or watts per square meter.
[0211] In an embodiment, the spectrum of the light source is represented as a 2D table of quantized wavelengths (in nm) and normalized intensities.
[0212] In an embodiment, the corresponding white point and primary colors of the light source are indicated (these data represent the information characteristics of the light source, such as those represented on a display).
[0213] In an embodiment, the color temperature is not provided as such, but rather in the form of a color transformation representing the white point variation as a 3×3 matrix with 9 coefficients indicated in the bitstream, or an indicator specifying the transformation coefficients or an external document specifying the transformation to be applied. As a variant, instead, the color temperature is signaled by the white point coordinates (x,y), which, according to the CIE 1931 definition of x and y in ISO 11664-1, are indicated by the normalized x and y chromaticity coordinates (li_white_point_x, li_white_point_y) of the white point of the light source.
[0214] The li_position[i][d], li_rotation_qx[i], li_rotation_qy[i], li_rotation_qz[i], li_center_view_flag[i], li_left_view_flag[i] syntax elements have the same semantics as the matching elements in the viewport position SEI message of V-PCC, except that it defines the light source position (instead of the viewport). Adopting similar syntax and semantics improves the readability and understandability of the specification.
[0215] In an embodiment, the rotation syntax elements (li_rotation_qx / qy / qz) that can indicate direction do not exist, but only the 3D coordinates (li_position) exist.
[0216] In an embodiment, the light source information does not exist.
[0217] In an embodiment, the light source position does not exist.
[0218] In an embodiment, the position is represented in the AR glasses (2D) viewport coordinate system.
[0219] A process (such as a renderer in an AR device) can use the lighting information data to re-light and / or add shadows to volumetric objects inserted into the captured scene.
[0220] In an embodiment, the environmental spherical harmonic coefficients are also carried in the bitstream BT2 of the encoded video picture data (S46, Figure 6 ). This may consist of an array of 9 sets of 3 RGB coefficients (i.e., 27 values).
[0221]
[0222]
[0223] Where:
[0224] li_ambient_spherical_harmonics_coeffs_present_flag being equal to 1 indicates that the environmental spherical harmonic coefficients are present in the bitstream. When equal to 0, it indicates that the environmental spherical harmonic coefficients are not present in the bitstream, and the value li_ash_coeffs[i][j] is inferred to be equal to 0.
[0225] When in the range of 0 to 1,000,000,000 (including the endpoint values), li_ash_coeffs[i][j] specifies the environmental spherical harmonic coefficients of the i-th group of the j-th component (where j = 0, 1, or 2 represents R, G, or B respectively) in increments of 0.000000001 units.
[0226] In an embodiment, the environmental HDR CubeMap is also carried in the bitstream BT2 of the encoded video picture data (S46, Figure 6 , e.g., encoded with VVC).
[0227] In computer graphics, the environmental HDR CubeMap can be represented in the RGB linear light (i.e., the color components are not gamma-corrected / the OETF is not applied) color space (possibly RGBA, where A is the alpha channel), and the values are represented by 16-bit floating point (also known as half-floating point). However, this format is not convenient for the distribution network, and after calculation / determination by a powerful computing device (such as an edge server or a high-end smartphone), the environmental HDR Cubemap that may be represented in RGB light linear half-floating point should be converted to the HDR 10-bit format, such as the SMPTE ST 2084 R’G’B’ or ST2084 Y’CbCr color space with 10-bit quantization. Therefore, the 10-bit HDR Cubemap can be encoded and decoded by a codec that supports signaling of the Cubemap format (such as HEVC or VVC), where the accompanying cube map information is encoded and decoded in the Cubemap projection SEI message, and gcmp_face_index[i] may be assigned as follows:
[0228] gmcp_face_index[0] = 0 (positive X)
[0229] gmcp_face_index[1] = 1 (negative X)
[0230] gmcp_face_index[2] = 5 (positive Y)
[0231] gmcp_face_index[3] = 4 (negative Y)
[0232] gmcp_face_index[4] = 3 (positive Z)
[0233] gmcp_face_index[5] = 2 (negative Z)
[0234] This corresponds to the normal indices of the environmental HDR cubemap.
[0235] The transfer function (ST 2084) and color space (RGB or Y’CbCr) are recorded in the HDR Cubemap stream, e.g., in the visual usability information or in the VUI with syntax elements such as transfer_characteristics and matrix_coeffs.
[0236] The previously described lighting information SEI message may also accompany the HDR Cubemap stream.
[0237] Furthermore, the present application also relates to an AR device DV1 and a processing device DV2 that are respectively configured to execute methods 600 and 640 Figures 6 to 9 ), respectively. The AR device DV1 and the processing device DV2 may include suitable tools (or units or modules) configured to respectively execute the respective steps of methods 600 and 640 according to any one of the embodiments of the present application. These devices may respectively correspond to, for example, Figure 5 the devices 500 and 502 as shown.
[0238] Specific embodiments of the AR device DV1 and the processing device DV2 are described below.
[0239] Figure 10 An exemplary embodiment of a system SY1 is shown, which system SY1 includes an AR device DV1 and a processing device DV2 configured to cooperate with each other as previously described to allow AR content rendering. With particular reference to Figures 5 to 9 the methods of the different embodiments shown in.
[0240] According to any embodiment of the present application, the AR device DV1 can be or include, for example, AR Google, a smartphone, a computer, a tablet, a head-mounted headset for rendering AR content, etc., or any other suitable device or apparatus configured to perform the steps of a method for rendering AR content.
[0241] The AR device DV1 may include a frame acquisition unit 1002 for acquiring a source video frame PR1 (e.g., acquired from a capture device DV3), and a video encoder 1004 for encoding the source video frame PR1 (possibly together with illumination information DT1 and / or spatial mapping information DT3) into an encoded bitstream BT1.
[0242] The AR device DV1 may further include a volume information acquisition unit 1006 for acquiring volume information DT1 (i.e., by receiving the DT1 from a content server DV4, and possibly by decoding it before aggregating the DT1 with PR1 in an AR aggregator 1010), an AR generation unit 1008 for generating an AR video frame PR3, and an AR rendering unit 1014 for rendering the AR video frame PR3. In particular, the AR generation unit 1008 may include an aggregator (or aggregation unit) 1010 for aggregating the source video frame PR1 with the volume information DT1, and a relighting unit 1012 for obtaining the AR video frame PR3 by combining a relighting effect associated with a virtual object OB2 into the aggregated video frame PR2 based on lighting information DT4.
[0243] According to any embodiment of the present application, still as Figure 10 shown, the processing device DV2 can be a high-end workstation or a mobile device, an edge processing unit, a server, etc., or any other suitable device or apparatus configured to perform the steps of a method for implementing AR content rendering.
[0244] The processing device DV2 may include a video decoder 1040 for decoding the encoded bitstream BT1 received from the AR device DV1, and a lighting estimation unit 1042 for performing lighting estimation based on the source video frame PR1.
[0245] The capture device DV3 is configured to capture the source video frame PR1 by means of, for example, any suitable capture sensor. In an embodiment, the AR device DV1 and the capture device DV3 form the same device.
[0246] The content server DV4 may include a storage means for obtaining / storing the volume information DT1.
[0247] Figure 11 A schematic block diagram illustrating an example of a system 1100 in which various aspects and embodiments are implemented is shown.
[0248] System 1100 may be embedded as one or more devices, including various components described below. In various embodiments, System 1100 may be configured to implement one or more aspects described in this application. For example, System 1100 is configured to execute a method for rendering AR content or a method for implementing AR content rendering according to any one of the previously described embodiments. Thus, System 1100 may constitute an AR device or a processing device in the sense of this application.
[0249] Examples of equipment that may constitute all or part of System 1100 include personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital TV receivers, personal video recording systems, connected household appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from video decoders, pre-processors that provide input to video encoders, web servers, video servers (such as broadcast servers, video-on-demand servers or network servers), static or video cameras, encoding or decoding chips or any other communication devices. The elements of System 1100 may be implemented singly or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of System 1100 may be distributed across multiple ICs and / or discrete components. In various embodiments, System 1100 may be communicatively coupled to other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports.
[0250] System 1100 may include at least one processor 1110 configured to execute instructions loaded therein for implementing, for example, various aspects described in this application. Processor 1110 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 1100 may include at least one memory 1120 (e.g., volatile memory devices and / or non-volatile memory devices). System 1100 may include a storage device 1140, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 1140 may include internal storage devices, attached storage devices, and / or network-accessible storage devices.
[0251] System 1100 may include an encoder / decoder module 1130, which is configured to process data, for example, to provide encoded / decoded video picture data, and the encoder / decoder module 1130 may include its own processor and memory. The encoder / decoder module 1130 may represent one or more modules or one or more units that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 1130 may be implemented as a separate element of the system 1100 or may be incorporated into the processor 1110 as a combination of hardware and software known to those skilled in the art.
[0252] The program code to be loaded onto the processor 1110 or the encoder / decoder 1130 to perform the various aspects described in this application may be stored in the storage device 1140 and subsequently loaded onto the memory 1120 for execution by the processor 1110. According to various embodiments, during the execution of the processes described in this application, one or more of the processor 1110, the memory 1120, the storage device 1140, and the encoder / decoder module 1130 may store one or more of various items. Such stored items may include, but are not limited to, video picture data, information data for encoding / decoding video picture data, bitstreams, matrices, variables, and intermediate or final results of equations, formulas, operations, and operation logic processing.
[0253] In several embodiments, the memory internal to the processor 1110 and / or the encoder / decoder module 1130 may be used to store instructions and provide a working memory for the processing that may be performed during encoding or decoding.
[0254] However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 1110 or the encoder / decoder module 1130) is used for one or more of these functions. The external memory may be the memory 1120 and / or the storage device 1140, for example, dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM may be used as a working memory for video encoding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 video), AVC, HEVC, EVC, VVC, AV1, etc.
[0255] As indicated in block 1190, input to the elements of system 1100 can be provided via a variety of input devices. Such input devices include, but are not limited to, (i) an RF portion that can receive RF signals transmitted over the air, for example, by a broadcast device, (ii) composite input terminals, (iii) USB input terminals, (iv) HDMI input terminals, and (v) when the present invention is implemented in the automotive field, buses such as CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data Rate), FlexRay (ISO17458), or Ethernet (ISO / IEC 802-3).
[0256] In various embodiments, the input devices of block 1190 have associated respective input processing elements, as is known in the art. For example, the RF portion can be associated with elements necessary for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a band), (ii) down-converting the selected signal, (iii) band-limiting the band again to a narrower band to select a signal band that can be referred to as a channel, for example, in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) de-multiplexing to select a desired data packet stream. The RF portion of various embodiments can include one or more elements that perform these functions, for example, a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a de-multiplexer. The RF portion can include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or baseband.
[0257] In one set-top box embodiment, the RF portion and its associated input processing elements can receive an RF signal transmitted over a wired (e.g., cable) medium. The RF portion can then perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band.
[0258] Various embodiments reorder the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.
[0259] Adding elements can include inserting elements between existing elements, such as, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF portion can include an antenna.
[0260] In addition, the USB and / or HDMI terminals may include respective interface processors for connecting the system 600 to other electronic devices via the USB and / or HDMI connections. It should be understood that various aspects of the input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within the processor 1110 when necessary. Similarly, various aspects of the USB or HDMI interface processing may be implemented within a separate interface IC or within the processor 1110 when necessary. The demodulated, error-corrected, and demultiplexed stream may be provided to various processing elements, including, for example, the processor 1110 and the encoder / decoder 1130, which operate in conjunction with memory and storage elements to process the data stream, when necessary, for presentation on an output device.
[0261] The various elements of the system 1100 may be provided within an integrated housing. Within the integrated housing, suitable connection arrangements 1190, such as internal buses (including I2C buses), wirings, and printed circuit boards known in the art, may be used to interconnect the various elements and transfer data between them.
[0262] The system 1100 may include a communication interface 1150 that enables communication with other devices via a communication channel 1151. The communication interface 1150 may include, but is not limited to, a transceiver configured to send and receive data on the communication channel 1151. The communication interface 1150 may include, but is not limited to, a modem or a network card, and the communication channel 1151 may be implemented, for example, within a wired and / or wireless medium.
[0263] In various embodiments, data may be streamed to the system 1100 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments may be received via the communication channel 1151 and the communication interface 1150 suitable for Wi-Fi communication. The communication channel 1151 of these embodiments may generally be connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other over-the-top communications.
[0264] Other embodiments may use a set-top box to provide streamed data to the system 1100, which delivers the data via an HDMI connection of the input block 1190.
[0265] Still other embodiments may use an RF connection of the input block 1190 to provide streamed data to the system 1100.
[0266] The data being streamed can be used as a way of signaling information used by system 1100, such as component conversion information DT1 (as described previously). The signaling information can include a bitstream B and / or information such as the number of video picture pixels and / or any encoding / decoding setting parameters.
[0267] It should be recognized that signaling can be implemented in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc. can be used to signal information to the corresponding decoder.
[0268] System 1100 can provide output signals to various output devices, including display 1161, speaker 1171, and other peripheral devices 1181. In various examples of the embodiments, other peripheral devices 1181 can include one or more of a standalone DVR, disc player, stereo system, lighting system, and other devices based on the output providing functions of system 1100.
[0269] In various embodiments, control signals can be communicated between system 1100 and display 1161, speaker 1871, or other peripheral devices 1181 using signaling such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that enable device-to-device control with or without user intervention.
[0270] The output devices can be communicatively coupled to system 1100 via dedicated connections through respective interfaces 1160, 1170, and 1180.
[0271] Optionally, the output devices can be connected to system 1100 via communication channel 1151 using communication interface 1150. Display 1161 and speaker 1171 can be integrated with other components of system 1100 in a single unit in an electronic device such as, for example, a television.
[0272] In various embodiments, display interface 1160 can include a display driver, such as, for example, a timing controller (TCon) chip.
[0273] For example, if the RF portion of input terminal 1190 is part of a separate set-top box, then display 1161 and speaker 1171 can optionally be separate from one or more of the other components. In various embodiments where display 1161 and speaker 1171 can be external components, output signals can be provided via dedicated output connections including, for example, HDMI ports, USB ports, or COMP outputs.
[0274] In Figures 1 to 9Among them, various methods are described herein, and each method includes one or more steps or actions to implement the described method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions can be modified or combined.
[0275] Some examples are described with respect to block diagrams and / or operational flowcharts. Each block represents a circuit element, module, or portion of code that includes one or more executable instructions for implementing the specified logical function(s). It should also be noted that in other embodiments, the function(s) labeled in the blocks may not occur in the order indicated. For example, depending on the functions involved, two blocks shown in succession may actually be executed substantially concurrently, or sometimes the blocks may be executed in the reverse order.
[0276] The embodiments and aspects described herein can be implemented in, for example, a method or process, apparatus, computer program, data stream, bit stream, or signal. Even if discussed only in the context of a single form of embodiment (e.g., only as a method), the embodiments of the features discussed can be implemented in other forms (e.g., an apparatus or a computer program).
[0277] A method can be implemented, for example, in a processor, which generally refers to a processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device.
[0278] In addition, a method can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the embodiments) can be stored on a computer-readable storage medium (such as, for example, storage device 1140( Figure 11 ). The computer-readable storage medium can take the form of a computer-readable program product implemented in one or more computer-readable media and having computer-readable program code executable by a computer implemented thereon. Considering the inherent ability to store information therein and the inherent ability to retrieve information therefrom, a computer-readable storage medium as used herein can be considered a non-transitory storage medium. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. It should be recognized that the following, although providing more specific examples of computer-readable storage media to which this embodiment can be applied, is merely illustrative and not an exhaustive list: a portable computer floppy disk; a hard disk; a read-only memory (ROM); an erasable programmable read-only memory (EPROM or flash memory); a portable compact disc read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination of the foregoing.
[0279] Instructions can form an application program tangibly implemented on a processor-readable medium.
[0280] For example, the instructions can be in hardware, firmware, software, or a combination. For example, the instructions can be found in an operating system, a separate application, or a combination of both. Thus, a processor can be characterized as, for example, a device configured to execute a process and a device including a processor-readable medium (such as a storage device) having instructions for executing the process. Additionally, in addition to or instead of the instructions, the processor-readable medium can store data values generated by an implementation.
[0281] The apparatus can be implemented in, for example, suitable hardware, software, and firmware. Examples of such apparatus include personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, a post-processor that processes the output from the video decoder, a pre-processor that provides input to the video encoder, web servers, set-top boxes, and any other device for processing video frames, or other communication devices. It should be clear that the equipment can be mobile and even installed in a moving vehicle.
[0282] The computer software can be implemented by processor 1110 or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can also be implemented by one or more integrated circuits. The memory 1120 can be of any type suitable for the technical environment and can be implemented using any appropriate data storage technology (such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples). The processor 1110 can be of any type suitable for the technical environment and can encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture, as non-limiting examples.
[0283] Based on the present application, as will be apparent to those of ordinary skill in the art, the embodiments can generate various signals that are formatted to carry information such as can be stored or transmitted. The information can include, for example, instructions for performing a method or data generated by one of the described embodiments. For example, the signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, the signal can be transmitted via various different wired or wireless links. The signal can be stored on a processor-readable medium.
[0284] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an" and "the" may also be intended to include the plural forms, unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "include" and / or "comprise" may specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Also, when an element is referred to as being "responsive" or "connected" or "associated" to another element, it can be directly responsive or connected to or associated with the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to another element or "directly associated" with another element, there are no intervening elements.
[0285] It should be recognized that, for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", the use of any one of the symbols / terms " / ", "and / or", and "at least one" can be intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to as many items as are listed.
[0286] A variety of numerical values can be used in the present application. Specific values can be used for illustrative purposes and the aspects described are not limited to these specific values.
[0287] It will be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the teachings of the present application, the first element can be referred to as the second element, and similarly, the second element can be referred to as the first element. No order is implied between the first element and the second element.
[0288] References to "an embodiment" or "embodiments" or "an implementation" or "implementations" and other variations are frequently used to convey that a particular feature, structure, characteristic, etc. (described in connection with the embodiment / implementation) is included in at least one embodiment / implementation. Thus, the appearances of the phrases "in an embodiment" or "in embodiments" or "in an implementation" or "in implementations" and any other variations that appear throughout the present application do not necessarily all refer to the same embodiment.
[0289] Similarly, references in this document to "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" and other variations thereof are frequently used to convey that a particular feature, structure, or characteristic (described in connection with an embodiment / example / implementation) may be included in at least one embodiment / example / implementation. Thus, the phrases "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" that appear throughout this application do not necessarily refer to the same embodiment / example / implementation, nor do individual or alternative embodiments / examples / implementations necessarily exclude each other from other embodiments / examples / implementations.
[0290] The reference signs that appear in the claims are for illustration purposes only and have no limiting effect on the scope of the claims. Although not explicitly described, the embodiments / examples and variations thereof may be employed in any combination or sub - combination.
[0291] When a figure is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0292] Although some figures include arrows on communication paths to indicate the main direction of communication, it should be understood that communication can occur in the direction opposite to the depicted arrows.
[0293] Various embodiments relate to decoding. As used in this application, "decoding" can cover, for example, all or part of the process of performing on a received video picture (which may include a received bitstream encoding one or more video pictures) to produce a final output suitable for display or further processing in a reconstructed video domain. In various embodiments, such processes include one or more of the processes typically performed by a decoder. In various embodiments, for example, such processes also or optionally include the processes performed by the decoders of the various embodiments described in this application.
[0294] As a further example, in one embodiment "decoding" may refer only to de - quantization, in one embodiment "decoding" may refer to entropy decoding, in another embodiment, "decoding" may refer only to differential decoding, and in another embodiment, "decoding" may refer to a combination of de - quantization, entropy decoding, and differential decoding. Based on the context of the specific description, whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear and is believed to be well understood by those skilled in the art.
[0295] Various embodiments relate to encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in the present application may cover, for example, all or part of the process performed on an input video picture to produce an output bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder. In various embodiments, such processes also include or optionally include the processes performed by the encoders of the various embodiments described in the present application.
[0296] As a further example, in one embodiment "encoding" may refer only to quantization, in one embodiment "encoding" may refer only to entropy encoding, in another embodiment, "encoding" may refer only to differential encoding, and in another embodiment, "encoding" may refer to a combination of quantization, differential encoding, and entropy encoding. Based on the context of a particular description, it will be clear whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process, and it is believed that those skilled in the art will well understand this.
[0297] Furthermore, the present application may refer to "obtaining" various information. Obtaining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from a memory, processing information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0298] Furthermore, the present application may refer to "receiving" various information. Receiving information may include, for example, one or more of accessing information or receiving information from a communication network.
[0299] Moreover, as used herein, the word "signal" especially indicates something to a corresponding decoder, etc. For example, in certain embodiments, the encoder signals specific information, such as encoding parameters or encoded video picture data. In this way, in an embodiment, the same parameters can be used on the encoder side and the decoder side. Thus, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, then signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select the specific parameters. By avoiding transmitting any actual functionality, bit savings are achieved in various embodiments. It should be recognized that signaling can be accomplished in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the foregoing relates to the verb form of the word "signal", the word "signal" can also be used as a noun herein.
[0300] Multiple embodiments have been described. However, it should be understood that various modifications can be made. For example, elements of different embodiments can be combined, supplemented, modified, or removed to yield other embodiments. Additionally, those of ordinary skill in the art will understand that other structures and processes can be substituted for the disclosed structures and processes, and the resulting embodiments will perform at least substantially the same (one or more) functions in at least substantially the same (one or more) ways to achieve at least substantially the same (one or more) results as the disclosed embodiments. Accordingly, this application contemplates these and other embodiments.
Claims
1. An augmented reality (AR) device (500; DV1) A method (600) for performing AR content rendering, the method comprising: - Transmitting (S8) encoded source video frame data (DT2) of a video frame (PR1) representing a real-world scene to a processing device (DV2) external to the AR device; - Receiving (S10) lighting information (DT4) from the processing device; - Generating (S14) an AR video frame (PR3) by: ο Obtaining (S16) an aggregated video frame (PR2) by aggregating the video frame and volume information (DT1) representing at least one virtual object (OB2); And ο Combining (S18) a relighting effect associated with the at least one virtual object into the aggregated video frame based on the received lighting information (DT4); and - Rendering (S18) the AR video frame.
2. The method according to claim 1, wherein the lighting information (DT4) and environmental map data (DT5) are obtained (S10) by decoding a bitstream of the encoded video frame data received from the processing device; wherein the relighting effect combined into the aggregated video frame is determined (S18) by applying the lighting information and the environmental map data to the at least one virtual object (OB2) inserted into the aggregated video frame (PR2).
3. The method according to claim 1 or 2, wherein combining (S16) the relighting effect comprises at least one of the following: - Overlaying a reconstructed lighting determined based on the lighting information on the at least one virtual object (OB2) inserted into the aggregated video frame; and - Overlaying a shadow effect representing at least one reconstructed shadow (SH2) of the at least one virtual object (OB2) on the aggregated video frame (PR2), the shadow effect being determined based on the lighting information.
4. The method according to any one of claims 1 to 3, wherein the volume information is obtained (S2) by any one of the following: - Receiving the volume information from a content server independently of the lighting information received from the processing device; and - Receiving the volume information and the lighting information as part of a bitstream of encoded video frame data.
5. A processing device (502; DV2) A method (640) for implementing augmented reality AR content rendering, the method comprising: - Receiving (S40) encoded video frame data (DT2) of a video frame (PR1) representing a real-world scene from an AR device (500; DV1) external to the processing device; - Obtaining (S42) a video frame of the real-world scene by decoding the received encoded video frame data (DT2); - Determining (S44) lighting information (DT4) by performing lighting estimation based on the video frame; and - Transmitting (S46) the lighting information (DT4) to the AR device for implementing AR rendering of the video frame.
6. The method according to claim 5, further comprising: - Obtain (S42) spatial mapping data (DT3) included in the encoded video frame data (BT1) by decoding the encoded video frame data, where the spatial mapping data defines a 3D coordinate system; wherein the lighting information (DT4) is determined based on the spatial mapping data (DT3) such that the processing device and the AR device share the 3D coordinate system.
7. The method according to claim 5 or 6, further comprising: - Obtain (S44) environmental map data (DT5) based on the source video frame (PR1), where the environmental map data defines a 3D view of the estimated lighting as an image; and - Encode (S46) the environmental map data in the bitstream transmitted to the AR device as an encoded image.
8. The method according to any one of claims 5 to 7, further comprising: - Obtain (S42) lighting information (DT6) that defines the lighting applied by the lighting device (700) to the real-world scene to capture the source video frame; wherein the lighting information (DT4) is determined based on the lighting information (DT6).
9. The method according to any one of claims 5 to 8, further comprising: - Insert (S46) the lighting information (DT4) as metadata in the bitstream of the encoded video frame data transmitted to the AR device.
10. A bitstream (BT2) of encoded processed data, formatted to include lighting information (DT4) obtained from one of the methods according to any one of claims 5 to 9.
11. The bitstream (BT2) according to claim 10, wherein the lighting information is formatted in a supplementary enhancement information message.
12. An AR device (500; DV1) for AR content rendering, the device including means for performing one of the methods (600) according to any one of claims 1 to 4.
13. A processing device (502; DV2) for implementing AR content rendering by an AR device, the processing device including means for performing one of the methods (640) according to any one of claims 5 to 9.
14. A computer program product, including instructions that, when the program is executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 9.
15. A non-transitory storage medium carrying instructions for program code for performing the method according to any one of claims 1 to 9.