Rendering Augmented Reality Content

The integration of physical and virtual lighting information for AR content rendering addresses the challenge of limited processing resources, achieving realistic and immersive AR experiences by efficiently relighting AR video images.

JP7811280B2Active Publication Date: 2026-02-04BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024559341
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-04-07
Filing Date
2023-02-17
Publication Date
2026-02-04
Estimated Expiration
2043-02-17

AI Technical Summary

Technical Problem

Existing AR devices struggle with accurate lighting estimation, leading to unrealistic rendering effects and reduced immersion due to limited processing resources and computational intensity, especially in mobile devices.

Method used

A method for AR content rendering that integrates physical and virtual lighting information to relight AR video images, using virtual and physical light sources, allowing for efficient relighting even on devices with limited resources.

Benefits of technology

Enables realistic and immersive AR experiences by accurately relighting AR content in real-time, optimizing rendering quality across various devices regardless of resource limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007811280000009
    Figure 0007811280000009
  • Figure 0007811280000010
    Figure 0007811280000010
  • Figure 0007811280000011
    Figure 0007811280000011
Patent Text Reader

Abstract

The present application relates to a method (400) for performing AR content rendering, the method including the steps of obtaining (410) a virtual relighting effect based on virtual lighting information (DT6) representing at least one virtual light source (SC2), obtaining (412) a mixed relighting effect based on the virtual lighting information (DT6) and physical lighting information (DT4) representing at least one physical light source (SC1), and generating (408) an AR video image (PR3) by aggregating a real-world scene relighted based on the virtual relighting effect (RL1) and at least one virtual object relighted based on the mixed relighting effect (RL2). Then, the AR video image (PR3) can be rendered.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates generally to augmented reality (AR) and to devices and methods for rendering or achieving the rendering of AR video images. In particular, but not exclusively, this application relates to lighting effects applied to AR video images and the processes involved in achieving such lighting effects. [Background technology]

[0002] This section is intended to introduce the reader to various aspects of the art, which may be related to various aspects of at least one embodiment of the present application, as described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to better understand all aspects of the present application. As such, it should be understood that these statements are to be read in this light, and not as admissions of related art.

[0003] Augmented reality (hereinafter abbreviated as "AR") can be defined as an experience in which three-dimensional (3D) information (the "augmented" portion) representing virtual objects is overlaid with video images (the "real" portion) representing the real-world environment as perceived by the user (3D). The visual AR information generated by the combination of the video images and the virtual information can be displayed by a variety of known AR devices, such as AR glasses, headphones, and mobile devices.

[0004] Game "Pokemon GO TM " may be considered one of the first large-scale adoptions of consumer AR services (2016). Since then, many AR services have been widely distributed across various sectors, including travel, education, healthcare, navigation systems, architecture, and retail.

[0005] To support various AR functions and services, AR devices may include multiple sensors (cameras, trackers, gyros, accelerometers, depth sensors, etc.) and processing modules (codecs, visual engines, renderers, etc.). For example, according to the 3GPP standardization committee, which standardizes and deploys AR services over 5G, AR devices can be divided into four types: 5G-independent AR UE, 5G EDGe-dependent AR UE, 5G tethered AR UE, and 5G wired tethered AR UE (XR over 5G presentation to VR-IF: https: / / www.vr-if.org / wp-content / uploads / VRIF-April-2021-Workshop-3GPP-SA4-presentation.pdf). Each AR device architecture can be applied to specific use cases or scenarios recognized by 3GPP.

[0006] For decades, various delivery technologies for 3D-based videos and models (e.g., H.264 / MVC, H.265 / MV-HEVC, 3D-HEVC, Google Draco) have been researched and standardized. However, due to different integration requirements, augmented reality, virtual reality, and mixed reality (AR, VR, and MR, respectively) have now become mainstream services and products in the market (e.g., AR glasses, headphones, etc.). 5G is also considered the primary carrier-based deployment of XR (XR = AR / VR / MR) services in consumers' daily lives. The maturity of 3D capture technologies is a major driver of this emergence, especially point cloud capture and multi-view + depth-based technologies.

[0007] One concept of AR technology is the insertion of visual (augmented) information into a captured real-world environment. In particular, AR video images can be obtained by inserting virtual objects into video images of real-world scenes. Therefore, lighting is one of the key aspects to providing a realistic experience to the user. In fact, overlaying virtual objects in an environment with inappropriate lighting and shading can break the illusion of immersion, making the objects appear to be floating in the air or not actually part of the scene.

[0008] 1 shows, for example, an AR video image 10 in which a virtual object 14 is inserted into a video image 12 of a real-world scene. The virtual object 14 may give the undesirable impression of floating in its real-world surroundings, especially in poor lighting.

[0009] Therefore, correctly lighting inserted virtual objects is crucial to providing a proper user experience. Non-relighted objects without consistent shading can reduce the immersion and interest of a business AR experience.

[0010] In the last few years, known lighting estimation methods have been developed to determine the lighting information of a perceived or captured scene into which a virtual object is inserted, allowing specific effects to be added to the virtual object for an elegant and smooth integration in the real-world environment. With the advent of artificial intelligence, learning-based methods have also emerged for this lighting estimation.

[0011] However, lighting estimation does not always produce the best results, and there is a need to perform more effective relighting of AR content to improve realism and enhance the user experience.

[0012] Another problem is that lighting estimation is not always supported by AR devices, or is supported in a limited and unfavorable way, which may result in a lack of realistic rendering effects. It is particularly difficult to achieve a smooth experience with real-time processing for AR devices (especially those with limited power resources, such as mobile devices, tablets, AR devices implemented as glasses, or headphones). Even in the case of pre-trained deep learning-based solutions, lighting estimation algorithms may require powerful resources. The built-in processing power of handheld devices may be limited, and batteries may be rapidly depleted (autonomous loss) due to computationally intensive tasks (e.g., decoding, 3D scene processing, and rendering).

[0013] This results in a trade-off algorithm that may have unrealistic effects that are detrimental to the user experience rather than enhancing the immersive scene.

[0014] Therefore, accurate relighting of realistic AR content, especially AR video images, is necessary to ensure an immersive user experience. Optimal rendering quality of AR content is expected from all AR devices. Summary of the Invention

[0015] The following section presents a simplified summary of at least one embodiment in order to provide a basic understanding of some aspects of the present application. This summary is not an extensive overview of the embodiments. It is not intended to identify key or critical elements of the embodiments. The following presents some aspects of at least one embodiment in a simplified form as a prelude to the more detailed description presented elsewhere herein.

[0016] According to a first aspect of the present application, there is provided a method for performing AR content rendering, the method comprising: - obtaining a virtual relighting effect based on virtual lighting information representing at least one virtual light source; - obtaining a mixed relighting effect based on virtual lighting information and physical lighting information representing at least one physical light source; aggregating the real-world scene relit based on the virtual relighting effect and the at least one virtual object relit based on the blended relighting effect to generate an AR video image; - rendering the AR video image.

[0017] In an embodiment, said generating an AR video image comprises: - applying said virtual relighting effect based on a video image representing a real-world scene to obtain a relighted real-world scene; - obtaining at least one relit virtual object by applying the blended relighting effect based on three-dimensional (3D) information representing the at least one virtual object; and aggregating the relit real-world scene and the at least one relit virtual object into an AR video image.

[0018] In an embodiment, the at least one physical light source and the at least one virtual light source are respectively controlled by physical lighting information and virtual lighting information: -Directional light source; -point light source, - a spot light source; is defined as any type of ambient light source and / or punctual light source.

[0019] In an embodiment, the method comprises: - obtaining a scene description document including first lighting information associated with a first indicator; - recognizing the first lighting information as physical lighting information based on the first indicator; - acquiring virtual lighting information through an API of the AR content rendering application.

[0020] In an embodiment, the method comprises: obtaining a scene description document including at least one of first lighting information associated with the first indicator and second lighting information associated with the second indicator, Here, the method - determining, based on a first indicator, to use first lighting information as physical lighting information; - determining, based on the second indicator, to use the second lighting information as virtual lighting information.

[0021] In an embodiment, the method comprises: obtaining a scene description document containing the descriptors; - accessing the virtual lighting information and the physical lighting information contained in the data stream based on the descriptors.

[0022] In an embodiment, the method comprises: - detecting, based on at least one indicator included in the received bitstream, that the received bitstream carries said virtual lighting information and physical lighting information as metadata; - extracting the virtual lighting information and physical lighting information from the received bitstream.

[0023] In an embodiment, the method comprises: - obtaining 3D coordinates associated with the virtual lighting information and the physical lighting information, respectively; - determining a spatial location where the virtual lighting information and the physical lighting information are to be applied relative to the real-world scene based on the 3D coordinates; Here, the virtual relighting effect and the mixed relighting effect are determined based on the spatial positions of the virtual lighting information and the physical lighting information.

[0024] According to a second aspect of the present application, there is provided a method for realizing AR content rendering, the method comprising: - obtaining physical lighting information representative of at least one physical light source by performing a first lighting estimation based on a video image of a real-world scene; - obtaining virtual lighting information representing at least one virtual light source emitting virtual light in a real-world scene; - transmitting physical lighting information and virtual lighting information to realize AR rendering of the AR video image.

[0025] In an embodiment, physical lighting information and virtual lighting information are transmitted in a scene description document to realize AR rendering of an AR video image, and the scene description document includes at least one syntax element representing the virtual lighting information and the physical lighting information, and at least one indicator indicating the presence of the virtual lighting information and / or the physical lighting information in the scene description document.

[0026] According to a third aspect of the present application, there is provided a scene description document obtained by the method of the second aspect of the present application, the scene description document being formatted to include at least one syntax element representing virtual lighting information and physical lighting information, and at least one indicator indicating the presence of virtual lighting information and / or physical lighting information in the scene description document.

[0027] According to a fourth aspect of the present application, there is provided an AR apparatus (or AR device) for AR content rendering, the AR apparatus comprising means for performing any of the methods of the first aspect of the present application.

[0028] According to a fifth aspect of the present application, there is provided a processing apparatus (or processing device) for realizing AR content rendering, the processing apparatus comprising means for performing any of the methods of the second aspect of the present application.

[0029] According to a sixth aspect of the present application, there is provided a computer program product comprising instructions, which when executed by one or more processors, cause the one or more processors to perform the method of the first aspect of the present application.

[0030] According to a seventh aspect of the present application, there is provided a non-transitory storage medium (or storage medium), the non-transitory storage medium (or storage medium) carrying program code instructions for performing the method of the first aspect of the present application.

[0031] According to an eighth aspect of the present application, there is provided a computer program product comprising instructions, which when executed by one or more processors, cause the one or more processors to perform the method of the second aspect of the present application.

[0032] According to a ninth aspect of the present application, there is provided a non-transitory storage medium (or storage medium), the non-transitory storage medium (or storage medium) carrying program code instructions for performing the method of the second aspect of the present application.

[0033] In this application, virtual objects can be elegantly and smoothly integrated into video images of real-world scenes, providing users with a realistic and immersive AR experience. By considering all physical and virtual light sources that may affect each part of the AR video image (i.e., the part representing the real-world scene and the part representing the virtual object, respectively), efficient relighting of the AR video image can be achieved.

[0034] Furthermore, the present application allows for the generation of realistic and immersive AR content regardless of the resources available at the AR device level and / or the resources for communication between the AR device and the outside. Therefore, optimal rendering quality of AR content can be achieved even by AR devices with limited processing resources (e.g., mobile devices (e.g., AR Google)). In particular, for realism, realistic lighting can be effectively reconstructed in AR video images. For example, accurate lighting estimation with no or limited delay can be achieved in real time.

[0035] The particular nature of at least one embodiment, as well as other objects, advantages, features, and applications of said at least one embodiment, will become more apparent from the following description of the example taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0036] The drawings of the present embodiment are given here by way of example. [Figure 1] FIG. 1 illustrates an AR video image generated by a related technique. [Figure 2] 1 is a schematic block diagram illustrating steps of a method 100 for encoding video images according to the related art; [Figure 3] 2 is a schematic block diagram illustrating steps of a method 200 for decoding video images according to the related art; [Figure 4] FIG. 1 illustrates a method for rendering AR content in one embodiment of the present disclosure. [Figure 5] 1A and 1B are diagrams illustrating a method for rendering AR content and a method for realizing AR content rendering according to an embodiment of the present disclosure. [Figure 6] FIG. 1 illustrates a real-world scene and illumination estimation according to an embodiment of the present application. [Figure 7] 1 illustrates various components of a lighting system according to an embodiment of the present application. [Figure 8A]FIG. 1 illustrates a portion of a method for rendering AR content according to an embodiment of the present application. [Figure 8B] FIG. 1 illustrates a portion of a method for rendering AR content according to an embodiment of the present application. [Figure 9] FIG. 1 illustrates an AR video image captured by an embodiment of the present application. [Figure 10] FIG. 1 is a schematic block diagram illustrating an example of a system in which various aspects and embodiments may be implemented. [Figure 11] 1 is a schematic block diagram illustrating an example of a system in which various aspects and embodiments may be implemented, with similar or identical components being numbered the same. DETAILED DESCRIPTION OF THE INVENTION

[0037] At least one embodiment will be described more fully hereinafter with reference to the accompanying drawings, in which at least one example embodiment is shown. However, embodiments may be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Accordingly, it should be understood that it is not intended to limit the embodiments to the particular forms disclosed. Instead, this application is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of this application.

[0038] At least one of these aspects relates generally to encoding and decoding video images, another relates generally to providing or transmitting encoded bitstreams, and another relates generally to receiving / accessing decoded bitstreams.

[0039] At least one of these embodiments is described as encoding / decoding a video image, but extends to encoding / decoding a video image (a sequence of images) since each video image can be sequentially encoded / decoded as follows:

[0040] It should be noted that at least one embodiment of this invention may be implemented using a video coding standard such as AVC (ISO / IEC 14496-10, Advanced Video Coding for generic audio-visual services, ITU-T Recommendation H.264, https: / / www.itu.int / rec / T-REC-H.264-202108-P / en), EVC (ISO / IEC 23094-1 Basic Video Codec), HEVC (ISO / IEC 23008-2 Efficient Video Codec, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en), or VVC (ISO / IEC 23090-3 Generic Video Coding, ITU-T Recommendation The present invention is not limited to MPEG standards such as H.266 (https: / / www.itu.int / rec / T-REC-H.266-202008-I / en), but may also apply to other standards and proposals, such as AV1 (AOMedia Video 1, http: / / aomedia.org / av1 / specification / ). At least one embodiment may apply to existing or future developments and extensions of such standards and proposals. Unless otherwise specified or technically excluded, aspects described herein may be used alone or in combination.

[0041] A pixel corresponds to the smallest display unit on the screen and can consist of one or more light sources (one for monochrome screens, three or more for color screens).

[0042] A video image (also called a frame or image frame) comprises at least one component (also called an image component or channel) determined by a particular image / video format, which specifies all information related to pixel values ​​and all information that a display unit and / or any other device can use to display and / or decode video image data related to said video image.

[0043] A video image comprises at least one component that is usually represented as an array of samples.

[0044] A monochrome video image contains a single component, and a color video image may contain three components.

[0045] For example, if the image / video format is the well-known (Y,Cb,Cr) format, a color video image may include a lightness (or luminance) component and two chroma components, or if the image / video format is the well-known (R,G,B) format, a color video image may include three color components (one for red, one for green, and one for blue).

[0046] Each component of a video image may include a number of samples that is related to the number of pixels of the screen on which the video image is intended to be displayed. For example, the number of samples included in a component may be the same as the number of pixels of the screen on which the video image is intended to be displayed, or may be a multiple (or fraction) of that number.

[0047] The number of samples contained in a component may be a multiple (or fraction) of the number of samples contained in another component of the same video image.

[0048] For example, if a video format includes a luminance component and two chroma components (e.g., a (Y,Cb,Cr) format), the chroma components may include half as many samples in width and / or height as the luminance component, depending on the color format considered.

[0049] A sample is the smallest visual information unit of a component that makes up a video image. A sample value may be, for example, a brightness or chroma value, or a color value in (R,G,B) format.

[0050] A pixel value is the value of a screen pixel. For monochrome video images, a pixel value can be represented by a single sample, while for color video images, a pixel value can be represented by multiple co-located samples. The co-located samples associated with a pixel are the samples that correspond to the pixel's location on the screen.

[0051] A video image is typically viewed as a set of pixel values, with each pixel represented by at least one sample.

[0052] A block of a video image is a set of samples of one component of the video image, which may be considered a block of at least one luminance sample or a block of at least one chroma sample if the image / video format is a known (Y,Cb,Cr) format, or a block of at least one color sample if the image / video format is a known (R,G,B) format.

[0053] This at least one embodiment is not limited to any particular image / video format.

[0054] 2 and 3 provide an overview of video encoding / decoding methods used in current video standard compression systems (e.g., VVC). As further pointed out below, these video encoding / decoding techniques or any suitable variations thereof may be used for encoding / decoding purposes in the present application. However, the present application is not limited to these examples.

[0055] FIG. 2 is a block diagram that illustrates in outline the steps of a method 100 for encoding a video image VP according to the related art.

[0056] In step 110, the video image VP is divided into sample blocks and partition information data is signaled in the bitstream. Each block contains samples of one component of the video image VP. Thus, these blocks contain samples that define each component of the video image VP.

[0057] For example, in HEVC, an image is divided into coding tree units (CTUs). Each CTU can be further subdivided using quadtree partitioning, where each leaf of the quadtree represents a coding unit (CU). The partitioning information data can then include data defining the CTUs and the quadtree subdivision of each CTU.

[0058] Then, each block of samples (CU) (abbreviated as block) is coded in a coding loop using an intra or inter prediction coding mode. Hereinafter, the term "in a loop" may refer to a step, function, etc. implemented in a loop (i.e., a coding loop in the coding stage or a decoding loop in the decoding stage).

[0059] Intra prediction (step 120) involves predicting the current block by predictive blocks based on coded, decoded and reconstructed samples located around the current block in the image (usually located above and to the left of the current block). Intra prediction is performed in the spatial domain.

[0060] In inter-prediction mode, motion estimation (step 130) and motion compensation (135) are performed. Motion estimation searches for candidate reference blocks as good predictors of the current block within one or more reference video images for predictively encoding the current video image. For example, a good predictor of the current block is a predictor similar to the current block. The output of motion estimation step 130 is one or more motion vectors and one or more reference image indexes associated with the current block. Next, motion compensation (step 135) obtains a prediction block from the motion vector(s) and one or more reference image indexes determined in motion estimation step 130. Essentially, a block belonging to a selected reference image and pointed to by a motion vector can be used as a prediction block for the current block. Note that, since motion vectors are expressed as fractions of integer pixel positions (called sub-pixel accuracy motion vector representation), motion compensation generally involves spatial interpolation of several reconstructed samples of reference images to calculate a prediction block sample.

[0061] Prediction information data is signaled in the bitstream, where the prediction information may include prediction mode, prediction information coding mode, intra prediction mode, or motion vector(s) and reference image index(es), and other information for obtaining the same prediction block at the decoding side.

[0062] The method 100 optimizes the rate-distortion trade-off and selects one of the intra or inter coding modes, taking into account, for example, the coding of a prediction residual block calculated by subtracting a candidate prediction block from a current block and the signaling of prediction information data required to determine the candidate prediction block at the decoding side.

[0063] Typically, the optimal prediction mode is given as the prediction mode of the optimal coding mode p* of the current block, and is given by:

number

number

[0064] Typically, the current block is coded based on the prediction residual block PR. More precisely, the prediction residual block PR is calculated, for example, by subtracting the best prediction block from the current block. Then, the prediction residual block PR is transformed (step 140), for example, by using a transform of the DCT (Discrete Cosine Transform) or DST (Discrete Sine Transform) type or any other suitable transform, and the obtained transform coefficient block is quantized (step 150).

[0065] In a variant, the method 100 may skip the transformation step 140 and apply quantization directly to the prediction residual block PR (step 150), according to a so-called transform-skip coding mode.

[0066] The quantized transform coefficient block (or quantized prediction residual block) is entropy coded into a bitstream (step 160).

[0067] As part of the encoding loop, the quantized transform coefficient block (or quantized residual block) is then inverse quantized (step 170) and inverse transformed (180) (or not) to obtain a decoded prediction residual block, which is then combined, typically summed, to provide a reconstructed block.

[0068] Other information data may be entropy coded in step 160 to code the current block of video image VP.

[0069] A loop filter (step 190) can be applied to the reconstructed image (including the reconstructed blocks) to reduce compression artifacts. After all image blocks have been reconstructed, a loop filter can be applied, including, for example, a deblocking filter, a sample self-adaptive offset (SAO), or a self-adaptive loop filter.

[0070] The reconstructed block or the filtered reconstructed block forms a reference picture, which is stored in a decoded picture buffer (DPB) and can be used as the coded reference picture for the next current block of the video picture VP or the next video picture to be coded.

[0071] FIG. 3 is a block diagram that illustrates in outline the steps of a method 200 for decoding a video image VP according to the related art.

[0072] In step 210, a bitstream of coded video image data is entropy decoded to obtain partition information data, prediction information data, and quantized transform coefficient blocks (or quantized residual blocks), for example, generated by method 100.

[0073] Other information data may also be entropy decoded from the bitstream to decode the current block of the video image VP.

[0074] In step 220, the reconstructed image is divided into current blocks based on the division information. Each current block is entropy decoded from the bitstream in the decoding loop. Each decoded current block is a quantized transform coefficient block or a quantized prediction residual block.

[0075] In step 230, the current block is inverse quantized and possibly inverse transformed (step 240) to obtain a decoded prediction residual block.

[0076] In another embodiment, prediction information data is used for the current block to obtain a predicted block by its intra prediction (step 250) or its motion-compensated temporal prediction (step 260). The prediction process performed on the decoding side is the same as the prediction process performed on the encoding side.

[0077] The decoded prediction residual block and the prediction block are then combined, typically summed, to provide a reconstructed block.

[0078] In step 270, a loop filter is applied to the reconstructed image (including the reconstructed block), and the reconstructed block or the filtered reconstructed block forms a reference image, which may be stored in a decoded picture buffer (DPB) as described above (FIG. 1).

[0079] Generally, the present application relates to rendering of augmented reality (AR) content and data processing for realizing the rendering of the AR content. One aspect of the present application relates to an AR device for performing AR content rendering and a processing device for realizing the AR content rendering, where the AR device and the processing device can cooperate with each other to realize the rendering of one or more AR video images. In an embodiment, the AR device and the processing device form the same device.

[0080] As will be further explained below, an AR video image is obtained by aggregating a real-world scene defined by a (source) video image and at least one virtual object (or virtual scene) defined by 3D information. In this application, to enhance the rendering of the acquired AR video image, the aggregated real-world scene and at least one virtual object are relighted (e.g., before or after aggregation) based on their respective lighting information. As can be seen, this application allows for relighting of AR video image content to reflect (or render) physical and virtual light projected by physical and virtual light sources, respectively, in the AR video image.

[0081] A physical light source is a light source that emits physical light (i.e., actual light that affects the appearance of a real-world scene captured in a source video image by any suitable capture sensor (e.g., a camera, etc.)). Physical light actually exists in a real-world scene and can be captured as part of a source video image of the real-world scene.

[0082] A virtual light source is a light source that emits virtual light (i.e., artificially reconstructed light that does not physically exist in the real-world scene but is integrated into the AR video image by a computational relighting process). Virtual light may originate from virtual objects that are not captured as part of the source video image of the real-world scene but are aggregated and integrated into the AR video image (acting as a virtual light source).

[0083] As will be further explained below, physical and virtual light can have different properties, such as punctual and / or ambient light. This also applies to physical and virtual light sources. In the following, the embodiments of the present application will refer to different types of physical and virtual light sources.

[0084] A directional light source (or directional light) is a light source that behaves as if it were at infinity, emitting light in the direction of the local z-axis. This light type inherits the orientation of the node it belongs to. Position and ratio are negligible, except for the inherited influence on the node's orientation. Because it is at infinite distance, the light is not attenuated. Its intensity can be defined in lumens per square meter or lux (lm / m2).

[0085] A point light source (or point light) emits in all directions from a position in space, and rotation and scale can be ignored as long as they do not affect the position of the inherited node. As the distance from the light's position increases, the brightness of the light falls off in a physically correct manner (i.e., brightness is inversely proportional to the square of the distance). The intensity of a point light can be defined in candela, which is lumens per square radian (lm / sr).

[0086] A spot light source (or spot light) emits a cone of light in the direction of the local z-axis. The cone angle and falloff are defined by two numbers, "innerConeAngle" and "outerConeAngle". Like point lights, the intensity also falls off in a physically correct manner with increasing distance from the light source position (i.e., intensity is inversely proportional to the square of the distance). The intensity of a spot light is the brightness within the innerConeAngle (and at the light position) and is defined in candelas, which are lumens per square radian (lm / sr). Engines that do not support two spot light angles should use the outerConeAngle as the spot light angle (which implicitly sets innerConeAngle to 0).

[0087] One aspect of the present application is to improve the realism and rendering quality of AR video images by relighting the AR video images by taking into account both physical and virtual light sources. To this end, an AR video image can be generated by aggregating a real-world scene relighted based on a virtual relighting effect and at least one virtual object relighted based on a blended relighting effect. The virtual relighting effect can be obtained based on virtual lighting information representing at least one virtual light source, and the blended relighting effect can be obtained based on the virtual lighting information and physical lighting information representing at least one physical light source.

[0088] That is, the AR video image can be generated based on a video image (or source video image) representing a real-world scene and 3D information representing at least one virtual object. The video image of the real-world scene can be acquired or captured by any suitable means. The 3D information may include volumetric information that defines at least one virtual object in space as a non-zero volume in space, although variations are possible in which the 3D information defines the virtual object as a point(s) or a 2D object.

[0089] Note that, before or after aggregating, a relighting effect may be applied to the real-world scene defined by the source video images and the at least one virtual object defined by the 3D information, respectively. More specifically, a virtual relighting effect may be applied to the real-world scene based on virtual lighting information representing at least one virtual light source, and a mixed relighting effect may be applied to the at least one virtual object based on the virtual lighting information and physical lighting information representing at least one physical light source. By aggregating the real-world scene relit based on the virtual lighting effect and the at least one virtual object relit based on the mixed relighting effect, an AR video image with improved quality may be obtained.

[0090] In an embodiment, a method for performing AR content rendering is provided, the method comprising: - obtaining a virtual relighting effect based on virtual lighting information representing at least one virtual light source; - obtaining a mixed relighting effect based on virtual lighting information and physical lighting information representing at least one physical light source; - aggregating the real-world scene relit based on the virtual relighting effect and at least one virtual object relit based on the mixed relighting effect to generate an AR video image.

[0091] The generated AR video image can then be rendered.

[0092] For example, the generating step may include: - applying said virtual relighting effect based on a video image representing a real-world scene to obtain a relighted real-world scene; - obtaining at least one relit virtual object by applying the blended relighting effect based on 3D information representing the at least one virtual object; and aggregating the relit real-world scene and the at least one relit virtual object into an AR video image.

[0093] As described above, prior to aggregating, i.e., aggregating the relit real-world scene and the relit at least one virtual object, a virtual relighting effect and a blended relighting effect can be applied to the real-world scene and the at least one virtual object, respectively.

[0094] In a variant, when the real-world scene and the at least one virtual object have been aggregated into an aggregated video image, i.e., after aggregation, a virtual relighting effect and a mixed relighting effect are applied to the real-world scene and the at least one virtual object, respectively. In this variant, the virtual relighting effect is applied to the real-world scene in the aggregated video image (i.e., applied to a first portion of the aggregated video image representing the real-world scene) based on the virtual lighting information, and the mixed relighting effect is applied to the at least one virtual object in the aggregated video image (i.e., applied to a second portion of the aggregated video image representing the at least one virtual object) based on the physical lighting information and the virtual lighting information.

[0095] One aspect of the present application is to realize AR content rendering by obtaining physical lighting information and virtual lighting information. The physical lighting information can be obtained by performing a first lighting estimation based on a (source) video image of a real-world scene. The virtual lighting information can be obtained by analyzing virtual scene information (e.g., a scene description document or file that describes various elements of a virtual scene including at least one virtual object) or by calling a programming function (e.g., an API) that provides the virtual scene information. The virtual object can be, for example, a virtual light source that emits virtual light.

[0096] The physical lighting information and the virtual lighting information can be obtained in different ways by the AR device responsible for AR content rendering. In an embodiment, the AR device performs the first lighting estimation and / or the second lighting estimation by itself to determine the physical lighting information and / or the virtual lighting information. The AR device DV1 may, for example, obtain the physical lighting information and / or the virtual lighting information from a local memory, or may receive the physical lighting information and / or the virtual lighting information from outside the AR device (e.g., from an external processing device DV2). In an embodiment, the AR device DV1 executes a lighting estimation algorithm to determine the physical lighting information representing at least one physical light source in a real-world scene. The AR device DV1 can determine the virtual lighting information by obtaining and analyzing virtual scene information (e.g., in a scene description document or file).

[0097] In an embodiment, lighting estimation is performed by a processing device external to the AR device to determine physical lighting information, which can then be transmitted (together with the virtual lighting information) to the AR device in any suitable manner (e.g., in a scene description document, one or more files, or a bitstream of encoded video image data) to enable AR rendering.

[0098] In an embodiment, a method for realizing AR content rendering is provided, the method comprising: - obtaining physical lighting information representing at least one physical light source by performing lighting estimation based on a video image of a real-world scene; - obtaining virtual lighting information representing at least one virtual light source; - transmitting physical lighting information and virtual lighting information to realize AR rendering of the AR video image.

[0099] The present application can generate and render high-quality AR video images based on an aggregation of a relit real-world scene and at least one relit virtual object (i.e., one or more relit virtual objects) in three dimensions (3D). Thus, two different types of lighting information (i.e., physical lighting information representing at least one physical light source and virtual lighting information representing at least one virtual light source) can be used to realize an improved relighting process that takes into account physical light and virtual light.

[0100] That is, the present application can integrate two types of lighting information from a real-world scene (one or more physical objects) and a virtual scene (one or more virtual objects) to achieve realistic relighting. In particular, the present application can distinguish between physical light sources and virtual light sources and appropriately consider them when relighting AR content.

[0101] More specifically, based on the virtual lighting information, virtual light emitted from a virtual object (which functions as a virtual light source) is recognized and taken into consideration to relight a real-world portion (real-world scene) defined in the source video image and / or relight a virtual portion (at least one virtual object) defined in the 3D information. For example, light reflections and / or shadows due to the virtual light can be incorporated into (overlaid on) the AR video image. Such relighting effects can be applied to the real-world portion of the AR video image and / or the virtual content (i.e., at least one virtual object) in the AR video image.

[0102] It should be noted that, based on the physical lighting information, physical light emitted from a physical light source can be recognized and taken into account for relighting the virtual portion (at least one virtual object) of the AR video image. In an embodiment, it is not necessary to relight the real-world portion of the aggregated video based on the physical lighting information, because the portion of the real-world portion corresponds to a real-world scene that, by definition, includes physical light (captured as part of the source video image). For example, light reflections and / or shadows due to the physical light may be incorporated into (overlay on) the virtual content (i.e., at least one virtual object) of the aggregated video image.

[0103] However, variations are possible in which physical relighting effects are also applied to the real-world scene based on the physical lighting information to alter (e.g., cancel or reduce) the physical lighting emitted from one or more physical light sources in the real-world scene, which in this particular example allows for canceling, dimming, or hiding physical light in the real-world scene aggregated into the AR video image.

[0104] This allows virtual objects to be elegantly and smoothly integrated into video images of real-world scenes, providing users with a realistic and immersive AR experience.

[0105] As will be apparent from the examples below, a variety of relighting effects can be achieved by practicing the present invention.

[0106] The present application may further permit efficient relighting of AR content regardless of available resources at the AR device level. Therefore, optimal rendering quality of AR content can be achieved even by AR devices with limited processing resources, such as mobile devices (e.g., AR Google). In particular, realistic lighting can be effectively reconstructed in AR video images for realism. For example, accurate lighting estimation with no or limited latency can be achieved in real time. In particular, some processing requests can be performed externally to the AR device, i.e., only some processing operations required to render AR content, including lighting estimation for determining physical lighting information, can be requested from an external processing device. However, other processing operations (e.g., aggregating the relighted real-world scene and aggregating at least one relighted virtual object (as described above)) may be performed locally by the AR device, limiting communication latency between the AR device and the processing device and maximizing AR rendering efficiency.

[0107] Therefore, even when the bandwidth available for communication and the local resources within the AR device are limited, this mixed processing distribution between the AR device and the processing device provides an optimal trade-off, enabling efficient and realistic AR content rendering.

[0108] Other aspects and advantages of the present application will be described in specific embodiments with reference to the accompanying drawings.

[0109] FIG. 4 illustrates a method 400 for rendering AR content according to an embodiment of the present application.

[0110] In this embodiment, an AR video image PR3 is generated and rendered. The AR video image PR3 includes a real-world portion (a first portion) representing a real-world scene and a virtual portion (a second portion corresponding to the augmented portion) representing at least one virtual object OB2. The AR video image PR3 is generated based on a source video image (also referred to as a video image) PR1 of the real-world scene and 3D information DT1 representing virtual content (i.e., at least one virtual object OB2).

[0111] The 3D information DT1 may include volume information that defines at least one virtual object in space as a non-zero volume in space, although variations are possible in which the 3D information DT1 defines one or more virtual objects as 3D points (i.e., by definition, no volume and no surface) and / or two-dimensional (i.e., no volume but a surface, e.g., a plane).

[0112] In this example, the real-world scene includes a physical (real) object OB1. In an embodiment, one or more physical objects OB1 may be included. As shown in FIG. 4 , the real-world scene includes physical light generated by a physical light source SC1 (which may or may not be part of the real-world scene). A physical shadow SH1 exists in the real-world scene due to the interaction between the physical object OB1 and the physical light emitted from the physical light source SC1. That is, the physical shadow SH1 is generated by physically projecting the physical light emitted from the physical light source SC1 onto the physical object OB1.

[0113] The configuration of the physical and virtual objects and physical and virtual light sources integrated into the AR content, in terms of their properties, shapes, numbers, arrangements, etc., can be changed depending on the situation.

[0114] The steps of the method 400 may all be performed by the AR device DV1, or may be performed in cooperation between the AR device and an external processing device, as further described in the following examples.

[0115] More specifically, physical lighting information DT4 and virtual lighting information DT6 are obtained (402) in any suitable manner. The physical lighting information DT4 represents (defines) at least one physical light source SC1 of the real-world scene. That is, the physical lighting information DT4 defines physical light emitted from at least one physical light source SC1 in the real-world scene.

[0116] Similarly, the virtual lighting information DT6 represents (defines) at least one virtual light source SC2, i.e., virtual light emitted from at least one virtual light source SC2. The virtual light source SC2 is a virtual object that emits virtual light, and this virtual object can have various properties. The virtual light source SC2 may be, for example, a direct virtual light source (e.g., a virtual lamp) or an indirect virtual light source (reflection of physical light and / or virtual light on a virtual object).

[0117] As will be apparent from the examples below, the nature and configuration of the physical and virtual light sources (and the physical and virtual light they generate) can vary depending on the particular situation. Both the physical and virtual light sources may be of various types, such as punctual and / or ambient light sources that emit punctual and / or ambient light, respectively.

[0118] Physical lighting information DT4 can be obtained by performing lighting estimation (402) based on source video images PR1 of a real-world scene.

[0119] The virtual lighting information DT6 can be obtained in any suitable way, for example, by analytically defining data of a virtual scene including at least one virtual object OB2. Typically, a 3D virtual scene may be structured with a characteristic diagram including virtual lights or virtual light source(s). This data structure can be of various types and obtained in different ways. For example, the virtual lighting information DT6 can be determined by obtaining (e.g., retrieving from a local memory or receiving from an external processing device DV2) and analyzing a scene description document (or file) SD1 that defines a virtual scene including at least one virtual object OB2.

[0120] In an embodiment, the AR device DV1 can receive the above data structure from the external processing device DV2 via an API of an AR application running on the AR device DV1. For example, the AR device DV1 can call a programming function (e.g., an API) to receive the scene description document (or file) SD1.

[0121] In an embodiment, the AR device DV1 receives physical lighting information DT4 and virtual lighting information DT6 in a bitstream of video image data from an external processing device DV2.

[0122] In addition, the real-world scene (physical object OB1) defined by the source video image PR1 and at least one virtual object OB2 (i.e., virtual scene) defined by the 3D information DT1 are relighted (406), and the AR video image PR3 is generated (404) by aggregating (408) the relighted real-world scene and the relighted at least one virtual object OB2.

[0123] More specifically, virtual relighting effects can be obtained based on the virtual lighting information DT6, where the virtual relighting effects represent at least one virtual light source SC2. Mixed relighting effects can also be obtained based on the virtual lighting information DT6 and the physical lighting information DT4, which represents at least one physical light source SC1. The real-world scene relighted based on the virtual relighting effects and the at least one virtual object relighted based on the mixed relighting effects can then be aggregated to generate the AR video image PR3.

[0124] As previously described, virtual and blended relighting effects are applied to the real-world and virtual portions, respectively, before aggregating 408 (i.e., relighting the physical object OB1 and the at least one virtual object OB2 before aggregating, i.e., individually with the aggregate) or after aggregating 408 (i.e., aggregating the real-world scene defined by source video images PR1 and at least one virtual object OB2 defined by 3D information and relighting both the real-world scene and the at least one virtual object OB2 in the aggregated video images). In any situation, relighting 406 includes a first relighting 410 that applies virtual relighting effects to the real-world scene (or to the source video images PR1 that define the real-world scene) and a second relighting 412 that applies blended relighting effects to the at least one virtual object OB2 (or to the 3D information that defines the at least one virtual object OB2).

[0125] 4, relighting 410 the real-world scene of AR video image PR3 may include inserting (or generating) a virtual shadow SH2a associated with physical object OB1 into the real-world scene, where the virtual shadow SH2a represents the virtual exposure of physical object OB1 to virtual light emitted from a virtual light source (in this example, SC2).

[0126] For example, as shown in FIG. 4, the second relighting 412 of the virtual scene may include inserting (or generating) one or more virtual shadows and / or one or more light reflections associated with the virtual object OB2.

[0127] For example, one or two virtual shadows, namely, virtual shadow SH2b representing the virtual interaction between virtual object OB2 and the physical light emitted from physical light source SC1 and / or virtual shadow SH2c generated by the virtual interaction between virtual object OB2 and the virtual light emitted from virtual light source SC2, can be inserted into AR video image PR3. Virtual shadow SH2b represents the virtual exposure of virtual object OB2 to the physical light emitted from physical light source SC1, and virtual shadow SH2c represents the virtual exposure of virtual object OB2 to the virtual light emitted from virtual light source SC2.

[0128] In this example, in the second lighting 412, the virtual shadow SH2b may be generated based on the physical lighting information DT4, and the virtual shadow SH2c may be generated based on the virtual lighting information DT6.

[0129] For example, virtual object OB2 and physical light source S C1 One or two of the following light reflections may be inserted into the AR video image PR3: a light reflection RF1 on the virtual object OB2 representing a virtual interaction with physical light emitted from a virtual light source SC2, and / or a light reflection RF2 on the virtual object OB2 representing a virtual interaction between the virtual object OB2 and the virtual light of the virtual light source SC2. For example, a light reflection of a physical object (e.g., OB1) and / or a light reflection of another virtual object may be overlaid on the virtual object OB2 in the aggregated video image PR2.

[0130] In this application, various virtual and mixed relighting effects can be applied, as will become clear in the following examples. By definition, since the real-world scene in the aggregated video image PR2 already contains physical light (e.g., physical light from physical light source SC1 in the example of FIG. 4), virtual relighting effects are applied to the real-world portions of the aggregated video image PR2 separately from (without considering) the physical lighting information DT4.

[0131] As will be explained in more detail in the examples below, the relighting effects applied to the real-world scene and the virtual object(s), respectively, can be tailored accordingly.

[0132] In an embodiment, based on the physical lighting information DT4, a physical relighting effect is applied to the real-world scene (either before or after aggregating 408) to change the physical lighting emitted from one or more physical light sources in the real-world scene. Thus, physical light in the real-world scene can be canceled, darkened, reduced, or hidden. In this way, excess physical light in the real-world portion of the AR video image PR3 can be reduced or canceled.

[0133] It should be noted that the present application can be applied in the same manner to one or more source video images PR1 to obtain and transfer one or more AR video images PR3. It should be noted that the real-world scene represented by the source video images PR1 may have various properties and may include one or more physical objects OB1 that may or may not have real shadow(s) SH1, or may not include any such physical objects (e.g., only a background). Similarly, the AR video image PR3 can be obtained by aggregating the (source) video images and 3D information representing one or more virtual objects OB2 (also referred to as 3D objects OB2).

[0134] Specific examples of the present application will now be described with reference to Figures 5 to 11. The examples described below may be exemplary implementations of the specific example of Figure 4.

[0135] In the following examples, it is assumed that an AR video image PR3 is obtained based on a source video image PR1 (also referred to as video image PR1) of a real-world scene. The source video image PR1 represents a real-world scene and may include one or more physical (real) objects OB1 of this scene. However, the present application can apply multiple source video images PR1 in the same manner, such as video images PR1 of a video data stream. In particular, the following examples can be applied when rendering AR video images as a video sequence.

[0136] In the following embodiment, the AR device DV1 may implement an API, for example, an API of an AR application for performing AR content rendering. This API may be, for example, an ARCore API, an ARKit (Apple®) API, or a WebXR API.

[0137] FIG. 5 is a block diagram that schematically illustrates steps S2 to S20 of a method 500 for performing AR content rendering performed by an AR device DV1, and steps S40 to S48 of a method 540 for realizing AR content rendering performed by a processing device DV2.

[0138] The AR device DV1 and the processing device DV2 can take various forms, as will be further explained below. Exemplary embodiments of the AR device DV1 and the processing device DV2 will now be described with reference to Figures 10-11.

[0139] 5-9, the AR device DV1 and the processing device DV2 cooperate to enable AR content rendering. To this end, a communication channel CN1 can be established between the AR device DV1 and the processing device DV2. That is, the processing device DV2 may be configured as a tethered device with respect to the AR device DV1.

[0140] However, it should be noted that although the AR device DV1 and the processing device DV2 are separate devices in this example, a variant is possible in which both methods 500 and 540 are performed by the AR device DV1. That is, the AR device DV1 and the processing device DV2 can form (or be part of) the same device. In particular, as will be explained further below, a variant is possible in which the AR device DV1 determines the physical lighting information DT4 (S44) and the virtual lighting information DT6 (S46). As will be explained further below, because all processing can be performed by the AR device DV1, in this variant, some steps of the methods 500 and 540 (e.g., steps of encoding and decoding various data and information, and steps of transmitting / receiving between the AR device DV1 and the processing device DV2) can be omitted.

[0141] In an acquisition step S4 (FIG. 5), the AR device DV1 acquires a source video image PR1 representing a real-world scene (or real-world environment). The real-world scene may have different characteristics depending on the situation. The source video image PR1 may have any suitable format.

[0142] For example, the source video image PR1 can be obtained by capturing (S4) a real-world scene as the source video image PR1 with at least one capture sensor built into (or locally included in) the AR device DV1. For example, the video image PR1 can be captured using a camera. In a variant, the AR device DV1 may receive (S4) the source video image PR1 from a capture device DV3 that captures the real-world scene as the source video image PR1 (e.g., via a camera or any other capture sensor). In a variant, the AR device DV1 retrieves (S4) the source video image PR1 from a local memory (i.e., a memory included in the AR device DV1).

[0143] As already explained, source video image PR1 may represent one or more physical objects OB1, which may or may not be physical light sources. In the example of Figure 5, it is assumed that the real-world scene represented by source video image PR1 includes physical objects OB1a and SC2, where SC2 represents a physical light source. Variations are possible in which physical light emitted from physical light source SC1 is present in the real-world scene captured in source video image PR1, but physical light source SC1 itself is not part of the real-world scene.

[0144] In an encoding step S6, the AR device DV1 encodes the source video image PR1 of the real-world scene as encoded video image data DT2. To this end, the AR device DV1 may use any suitable encoding technique, such as the encoding method (method 100) described with reference to Figure 2. As mentioned above, variants in which there are no encoding and decoding steps are also possible.

[0145] In a transmission step S8, the AR device DV1 transmits the encoded video image data DT2 to the processing device DV2. In this example, the transmission S8 is in the form of an encoded bitstream BT1 and is performed by a communication channel (or communication link) CN1, which may be, for example, a 5G link or any other suitable channel, and may be performed, for example, using a communication network. 5G provides low-latency communication.

[0146] In an embodiment, in the encoding step S6, the AR device DV1 encodes spatial mapping information together with the encoded video image data DT2. In this case, the spatial mapping information is transmitted (S8) to the processing device DV2 together with or together with the encoded video image data DT2. The spatial mapping information defines a 3D coordinate system that can be used to ensure that the AR device DV1 and the processing device DV2 share the same 3D spatial coordinate system to achieve consistent lighting estimation.

[0147] In a receiving step S40, the processing device DV2 receives the encoded video image data DT2. The processing device DV2 then obtains (S42) the source video image PR1 by decoding the received encoded video image data DT2. The decoding can be performed in any suitable manner, such as the decoding method (method 200) described above with reference to Figure 3.

[0148] In an acquisition step S44, the processing device DV2 acquires physical lighting information DT4 by performing lighting estimation based on the source video image PR1. This lighting estimation S44 can be performed in any known manner, such as by implementing a lighting estimation algorithm based on ARCore® (Google Play Services for AR) or ARKit®. For example, a lighting estimator unit such as ARShadowGAN (Generative Adversarial Network for Shadow Generation), GLEAM (Lighting Estimation Framework for Real-Time Augmented Reality on Mobile Devices), or DeepLight (Light Source Estimation for Augmented Reality Using Deep Learning) can be used.

[0149] The physical lighting information DT4 acquired in S44 defines the characteristics of the physical light to which the real-world scene represented by the source video image PR1 is exposed. That is, the physical lighting information DT4 represents the physical light emitted from one or more physical light sources SC1 in the real-world scene. As shown in FIG. 5, in this embodiment, the physical lighting information DR4 defines the physical light source SC1 in the real-world scene or the physical light emitted from this physical light source SC1.

[0150] 6 is a schematic diagram illustrating a physical light source SC1 (i.e., a light source that exists in the real world) projecting light onto a physical object OB1 that exists in a real-world scene captured by a source video image PR1. By performing a first illumination estimation S44, one or more characteristics of the physical light emitted from the physical light source SC1 can be identified.

[0151] In an obtaining step S46 (FIG. 5), the processing device DV2 obtains the virtual lighting information DT6 in any suitable manner, for example by parsing a data structure (e.g., a scene description document SD1, a file, a function from an API, etc.) that describes the virtual scene (including the virtual scene information). As will be explained further below, the virtual scene may include at least one virtual object OB2.

[0152] In an embodiment, the physical lighting information DT4 defines at least one physical light source SC1 that emits physical light in the real-world scene.

[0153] In an embodiment, the virtual lighting information DT6 defines at least one virtual light source SC2 that emits virtual light in the real-world scene.

[0154] The physical lighting information DT4 and / or the virtual lighting information DT6 obtained in S44 and S46, respectively, may each include at least one (or any combination thereof) of the following light characteristics: lighting direction, lighting intensity, lighting position, color temperature, spectrum, and ambient spherical harmonics.

[0155] The physical lighting information DT4 and / or the virtual lighting information DT6 obtained in S44 and S46, respectively, may both be determined based on spatial mapping information (if present) received from the AR device DV1 as part of the encoded video image data DT2. For example, as shown in Figure 6, the lighting position 602 of the physical light source SC1 may be defined in 3D spatial coordinates (representing the real-world scene), for example, using the 3D spatial coordinates defined by the spatial mapping information (if present) received (S40) from the AR device DV1.

[0156] The ambient spherical harmonics may represent the spherical harmonic coefficients for ambient lighting.

[0157] For example, the physical lighting information DT4 may define a principal directional light representing the physical light source SC1 as the primary light source in the real-world scene and an environment spherical harmonic representing residual ambient light energy in the real-world scene. As described further below, determining the principal directional light may be used to cast shadows. Similarly, the virtual lighting information DT6 may define a principal directional light representing the virtual light source SC2 as the secondary primary light source corresponding to the virtual object OB2 and an environment spherical harmonic representing ambient light energy.

[0158] In an embodiment, one or more physical light sources SC1 and one or more virtual light sources SC2 are defined as ambient light sources and / or punctual light sources of any type, including directional light sources, point light sources, and spot light sources, by physical lighting information DT4 and virtual lighting information DT6, respectively.

[0159] In an embodiment, in the illumination estimation step S44, the processing device DV2 further determines environment map data DT5 based on the source video images PR1. These environment map data DT5 define a 3D view of the estimated physical light as an image. However, embodiments are also possible that do not use such environment map data.

[0160] More specifically, the environment map data DT5 defines an environment map, which is a mapping image that represents an omnidirectional view of the physical ambient lighting of a 3D scene (the 3D scene shown in the source video image PR1) as seen from a particular 3D position. Each pixel in the mapping image corresponds to a 3D direction, and the data stored in the pixel represents the amount of light arriving from that direction. In 3D rendering applications, the environment map can be used for image-based lighting techniques that approximate how an object is lit by its environment.

[0161] An environment map with sufficient dynamic range to represent the brightest light sources in the environment is sometimes called a "light detection image."

[0162] For example, the environment map data DT5 can represent the mapping image as a cubemap (also called an environment cubemap or HDR (High Dynamic Range) cubemap), which is a method of encoding 360-degree light information. In a cubemap, the environment lighting is projected onto the six faces of an axis-aligned cube. Each face of the cube is then aligned to form a 2D image.

[0163] For example, ARCore® (Google Play Services for AR) uses cube maps to represent environmental lighting. Using cube maps can achieve good performance for efficiently encoding 360 video sequences. However, any other 360 projection, such as iso-rectangular, fish-eye, or cylindrical, can be used. For example, an environment map can be represented as a "latitude and longitude map," also known as an isometric projection in the 360 ​​video field.

[0164] The environment map data DT5, which may take the form of a cube map or similar, can be advantageously used to reconstruct realistic lighting effects on virtual objects, and as will be explained further below, it can be particularly used to render reflections of shiny metallic objects in AR content.

[0165] In an embodiment, the physical lighting information DT4 determined in S44 includes a primary directional light component representing the physical light source SC1 as the primary light source in the real-world scene, ambient spherical harmonics representing the residual ambient light energy in the real-world scene, and the aforementioned environment map data DT5. Figure 7 illustrates the effect of these three components in a particular example.

[0166] In a variant, environment map data DT7 (e.g., taking the form of a cube map, similar to DT5) may be determined based on 3D information representing at least one virtual object OB2. These environment map data DT7 may define a 3D view of the estimated virtual light as an image. However, embodiments are also possible in which such environment map data DT7 are not used. These environment map data DT7 can then be advantageously used to reconstruct realistic lighting effects associated with the real-world scene and / or the at least one virtual object OB2.

[0167] In order to maximize the rendering quality of the AR content, the above three components may be refreshed (determined) by the processing device DV2 every time a new source video image PR1 (or video frame) is provided by the AR device DV1, so that the light direction information and the environment sphere harmonics become a timed sequence of metadata, and the environment map data DT5 becomes a timed sequence of textures, i.e., a video sequence.

[0168] In a transmission step S48, the processing device DV2 realizes AR content rendering by transmitting the physical lighting information DT4 and the virtual lighting information DT6 (environment map data DT5 and / or DT7 may also be present) to the AR device DV1. This transmission S48 may be performed, for example, via a communication channel CN1 established between the AR device DV1 and the processing device DV2 (e.g., via a 5G link).

[0169] In a receiving step S10, the AR device DV1 receives physical lighting information DT4 and virtual lighting information DT6, and possibly also environment map data DT5 and / or DT7 (if present).

[0170] The transmission S48 of the physical lighting information DT4 and the virtual lighting information DT6 to the AR device DV1 can be performed in various ways, and some examples are provided below.

[0171] In an embodiment, the physical lighting information DT4 and the virtual lighting information DT6 (environment map data DT5 and / or DT7 may also be present) are transmitted to the AR device DV1 in a scene description document (or one or more files) SD1 (S48) and received by the AR device DV1 (S14). For example, the physical lighting information DT4 and the virtual lighting information DT6 may be received via a content delivery network (CDN) using the HTTP protocol or any other suitable protocol. In a particular example, the scene description document SD1 including the physical lighting information DT4 and the virtual lighting information DT6 is encapsulated in a file format structure such as the ISO Base Media File Format structure and is retrieved by the AR device DV1 via the HTTP protocol.

[0172] Therefore, the AR device DV1 can analyze the received scene description document (or file(s)) SD1 to obtain the physical lighting information DT4 and the virtual lighting information DT6.

[0173] Although this is not required, the physical lighting information DT4 and the virtual lighting information DT6 may be transmitted in an encoded (or compressed) format (S48). In this case, the AR device AR1 can obtain the physical lighting information DT4 and the virtual lighting information DT6 by decoding (e.g., performing video decompression). The decoding may be performed in any suitable manner, such as the decoding method (method 200) described above with reference to FIG. 3.

[0174] In an embodiment, transmission S48 serves as a bitstream BT2 of coded video image data in coded form, this bitstream BT2 including physical lighting information DT4 and virtual lighting information DT6, and possibly also environment map data DT5 and / or DT7 (if present). The coding may be performed in any suitable manner, such as the coding method (method 100) described with reference to Figure 2.

[0175] In an embodiment, before transmitting S48 to the AR device DV1, the processing device DV2 inserts the physical lighting information DT4 and the virtual lighting information DT6 as metadata in a bitstream BT2 of the encoded video image data (S48). For example, the physical lighting information DT4 and the virtual lighting information DT6 are carried as metadata in one or more SEI messages (Supplementary Extension Information) of the bitstream BT2. A specific implementation of carrying the physical lighting information DT4 and the virtual lighting information DT6 in the bitstream BT2 will be further described below.

[0176] In an embodiment, the environment map data DT5 determined in S44 is also encoded (S46) as an encoded image in the encoded bitstream BT2 of encoded video image data before S48 is transmitted to the AR device DV1.

[0177] It does not matter how the transmission S48 is performed (e.g., via a scene description document, one or more files, or a bitstream of encoded video image data). In an embodiment, since the AR device DV1 already has the source video image PR1 (acquisition step S4), the processing device DV2 does not return the source video image PR1 to the AR device DV1 to avoid unnecessary data exchange and communication delays. While waiting to receive the physical lighting information DT4 and the virtual lighting information DT6, the AR device DV1 can store the source video image PR1 in its local memory (for later retrieval).

[0178] In an embodiment, the AR device DV1 can obtain the physical lighting information DT4 and the virtual lighting information DT6 (and possibly also the environment map data DT5 and / or DT7) by decoding a bitstream BT2 of coded video image data received by the processing device DV2 (e.g., via the communication channel CN1). The decoding can be performed in any suitable manner, such as the decoding method (method 200) described above with reference to FIG. 3.

[0179] 5 and 8A, the AR device DV1 relights (S14) the real-world scene (i.e., physical object OB1) defined by the source video image PR1 and relights (S16) at least one virtual object OB2 (virtual scene) defined by the 3D information DT1. As further described below, the AR device relights the real-world scene based on a virtual relighting effect RL1 and relights the at least one virtual object OB2 based on a blended relighting effect.

[0180] To this end, the AR device DV1 acquires (S2) 3D information DT1 in any suitable manner. The 3D information DT1 represents at least one virtual (or 3D) object OB2, which constitutes virtual (or augmented) content that is aggregated with a real-world scene to generate AR content. Each virtual object OB2 can be defined by the 3D information DT1 in any suitable 3D (or volumetric) format.

[0181] The one or more virtual objects OB2 integrated into the source video image PR1 can have various properties, such as being dynamic (or moving) objects, i.e. virtual objects that are defined to move virtually within the real-world scene by the 3D information DT1.

[0182] As already explained, the 3D information DT1 may include volumetric information that defines at least one virtual object in space as having a non-zero volume in space (volumetric object), although variants are possible in which the 3D information DT1 defines one or more virtual objects as 3D points (i.e. having no volume and no surface by definition) and / or two-dimensional (i.e. having no volume but having a surface, e.g. a plane).

[0183] As shown in the examples of FIGS. 5 and 8A, in this example, it is assumed that the 3D information DT1 defines two virtual objects OB2a and SC2 (collectively referred to as OB2). It is assumed that the virtual object SC2 is a virtual light source that emits virtual light in a real-world scene. Of course, it is possible to adjust the structure, such as the number, type, and arrangement, of one or more virtual objects OB2. For example, the 3D information DT1 can define a single virtual object OB2 as the virtual light source SC2. Based on the present disclosure, those skilled in the art will easily understand how to adjust the embodiment.

[0184] In an embodiment, the AR device DV1 receives (S2) the 3D information DT1 from a content server DV4 (FIG. 10) separately from (without) the physical lighting information DT4 and the virtual lighting information DT6 (e.g., separately from the scene description document (or file) SD1 received from the processing device DV2), for example via a second communication channel CN2 separate from the first communication channel CN1. The 3D information DT1 can, for example, be received (S2) from a content server DV4 different from the processing device DV2.

[0185] In an embodiment, the AR device DV1 receives (S2) the 3D information DT1 and the physical lighting information DT4 and virtual lighting information DT6 as part of the scene description document SD1 or a file or a bitstream of encoded video image data received at S10 as described above. In this particular case, the processing device DV2 transmits the 3D data DT1 to the AR device DV1.

[0186] In an embodiment, the content server DV4 transmits the 3D data DT1 to the processing device DV2, and the processing device DV2 encodes the received 3D data DT1, physical lighting information DT4, and virtual lighting information DT6 (environment map data DT5 may also be present) as a bitstream BT2 and transmits it to the AR device DV1. In this variant, the processing device DV2 acts as a repeater to transmit the 3D information DT1 to the AR device DV1.

[0187] In the embodiment, the AR device DV1 retrieves the 3D information DT1 from the local memory of the AR device DV1 (S2).

[0188] More specifically, as explained below, the relighting step S12 (shown in FIGS. 5 and 8A) performed by the AR device DV1 includes a first relighting step S14 and a second relighting step S16. In this example, a virtual relighting effect RL1 and a blended relighting effect RL2 are applied to the real-world scene and the at least one virtual object OB2, respectively, before aggregating them into the AR content, although, as already explained, variations are possible.

[0189] In a first relighting step S14, a virtual relighting effect RL1 is obtained (or determined) based on the virtual lighting information DT6 representing at least one virtual light source (e.g., SC2 in this example) obtained in S10. Then, based on the virtual relighting effect RL1, the real-world scene defined by the source video image PR1 is relighted (S14). For example, 14 The virtual relighting effect RL1 applied in is associated with the physical object OB1a by changing the lighting of the physical object OB1a itself and / or changing the lighting of other regions of the aggregated video images PR2 as a result of the physical object OB1a being virtually exposed to virtual light (e.g., the virtual light of virtual light source SC2). A first relighting step S14 may be performed based on environment map data DT7 (if present).

[0190] In the second re-lighting step S16, the virtual lighting information DT6 obtained in S10 and at least one physical light source (e.g., SC in this example) obtained in S10 are used. 1 The mixed relighting effect RL2 is obtained (or determined) based on the physical lighting information DT4 representing the 3D information DT1 and the virtual lighting information DT4 representing the virtual lighting. RL2 represents a mixed relighting effect because it takes into account both physical lighting and virtual lighting. Then, at least one virtual object OB2 defined by the 3D information DT1 is relighted (S16) based on the mixed relighting effect RL2.

[0191] The second relighting step S16 may be performed based on the environment map data DT5 and / or DT7 (if present). For example, the blended relighting effect RL2 applied in S16 is associated with the virtual object OB2a because it has changed the lighting of the virtual object OB2a itself, which is a result of the virtual object OB2a being virtually exposed to virtual light (e.g., the virtual light of virtual light source SC2).

[0192] Generally, the relighting effects RL1, RL2 applied in the relighting steps S14, S16 are configured to modify the lighting of the real-world and virtual parts to provide a more immersive and realistic AR experience to the user. Both relighting steps S14 and S16 incorporate or overlay relighting effects (or reconstructed lighting) into the real-world scene and the virtual content, respectively. Thus, the relighted real-world scene and the relighted at least one virtual object OB2 generated by the relighting steps S14, S16 include relighting effects RL1, RL2. These relighting effects RL1, RL2 can be of various types and can include at least one (or a combination) of shadow effects, ambient lighting, specular highlights, and light reflections, among others.

[0193] In an embodiment, the first relighting step S14 may include one or both of the steps of overlaying the reconstructed lighting onto the physical object OB1 and / or associating (or combining) the physical object OB1 with at least one reconstructed shadow effect representing the physical object OB1, where these reconstructed lighting and shadow effects are determined based on the virtual lighting information DT6.

[0194] In an embodiment, the second relighting step S16 may include either or both of overlaying the reconstructed lighting onto the virtual object OB2 and associating (or combining) the virtual object OB2 with at least one reconstructed shadow effect representing the virtual object OB2, where these reconstructed lighting and shadow effects are determined based on the physical lighting information DT4 or the virtual lighting information DT6 (or both).

[0195] 8, the virtual heavy lighting effect RL1 applied in the first relighting step S14 may include a virtual shadow SH2a in the relighted real-world scene, the virtual shadow SH2a representing the virtual interaction between the physical object OB1a and the virtual light emitted from the virtual light source SC2. That is, the virtual shadow SH2a formed by virtually exposing the physical object OB1a to the virtual light of the virtual light source SC2 is generated and incorporated into the relighted real-world scene.

[0196] In this embodiment, physical light (e.g., light emitted from SC1) is already present in the real-world scene defined by the source video image PR1 according to the definition, so the first re-lighting step S 14The virtual relighting effect RL1 applied in does not take into account the physical lighting information DT4. In this example, the physical shadow SH1 generated by physically exposing the physical object OB1 to the physical light of the physical light source SC1 is already shown in the source video image PR1. As already explained, a variation is possible, in which the physical relighting effect based on the physical relighting information DT4 is also applied to the real-world scene, for example to cancel, reduce, or darken the physical light present in the real-world scene (e.g., the physical light emitted from the physical light source SC1).

[0197] For example, as shown in FIG. 8, the mixed relighting effect RL2 applied in the second relighting step S16 may apply various relighting effects associated with the virtual object OB2a, such as (one or more) virtual shadows SH2 and / or (one or more) virtual light reflections RF.

[0198] More specifically, the mixed relighting effect RL2 can generate virtual shadows such as either (or both) of the following: generating a virtual shadow SH2b based on the physical lighting information DT4, the virtual shadow SH2b representing a virtual interaction between the virtual object OB2a and the physical light emitted from the physical light source SC1; - generating a virtual shadow SH2c based on the virtual lighting information DT6, the virtual shadow SH2c representing a virtual interaction between the virtual object OB2a and the virtual light emitted from the virtual light source SC2;

[0199] That is, virtual shadows SH2b and SH2c formed by virtually exposing the virtual object OB2a to the physical light of the physical source SC1 and the virtual light of the virtual light source SC2, respectively, can be integrated into the aggregated video image PR2.

[0200] Note that the mixed relighting effect RL2 applied in the second relighting S16 can generate a virtual light reflection RF, i.e., a virtual light reflection RF1 of physical light from a physical light source SC1 on a virtual object OB2a; - Virtual light from a virtual light source SC2, and a virtual light reflection RF2 in the virtual object OB2a can be generated.

[0201] In the aggregation step S18 shown in Figures 5 and 8B, the AR device DV1 generates an AR video image PR3 by aggregating the real-world scene relighted based on the virtual relighting effect RL1 in S14 and at least one virtual object OB2 relighted based on the mixed relighting effect RL2 in S16.

[0202] As a result of the two-step relighting S12, if the light source is virtual, virtual light emitted from the light source can be coupled (or rendered) to the real-world scene and to one or more virtual objects OB2 in the AR video image PR3. If the light source is real (i.e., a physical light source), physical light emitted from the light source is coupled (or rendered) only to one or more virtual objects OB2 in the AR video image PR3, for example.

[0203] In this embodiment, aggregation S18 overlays (or integrates) the re-illuminated virtual object OB2 onto (or into) the re-illuminated real-world scene, thereby generating an AR video image PR3. Thus, the AR video image PR3 includes two parts (or multiple parts): a real-world part (or first part) representing the re-illuminated real-world scene, and a virtual part (or second part) representing the re-illuminated virtual object OB2. The virtual object OB2 constitutes augmented content that is overlaid on the real-world environment perceived by the user as defined by the source video image PR1.

[0204] 9 illustrates the generation of an AR video image PR3 by inserting virtual objects OB2a (rocket) and SC2 (light bulb) into source video images PR1 according to an embodiment. In a second relighting step S18 (FIG. 5), a virtual light reflection RF2 of virtual light emitted from a virtual light source SC2 is superimposed on the virtual object OB2a in the aggregated video images PR2, the virtual light reflection RF2 being determined based on virtual lighting information DT6.

[0205] 5 , in a rendering step S20, the AR device DV1 may render the AR video image PR3 generated in S18 by any suitable means. The rendering may be performed by displaying the AR video image PR3 (or using any 3D projection, etc.). In an embodiment, the AR device DV1 displays the AR video image PR3 using a local display unit integrated into the AR device DV1. The display unit may take any suitable form, such as an AR Google, headphones for rendering AR content, a screen, etc.

[0206] The AR device DV1 may perform appropriate calculations to render the AR video image PR3 as AR content. In an embodiment, the AR device DV1 may transmit the AR video image PR3 to an external display unit for display.

[0207] As already explained, in the embodiment shown in FIG. 5, the AR device DV1 and the processing device DV2 are separate devices. In the embodiment, both methods 500 and 540 are executed by the AR device DV1. That is, the AR device DV1 and the processing device DV2 can form (or be part of) the same device. Thus, the AR device DV1 can obtain or determine (S46) the physical lighting information DT4 (S44) and the virtual lighting information DT6, as described above. In this case, steps S8 and S48 of transmitting / receiving data and / or information may be omitted. Similarly, although encoding and decoding may also occur in the AR device DV1, the encoding / decoding operations (S6, S42, S48, S14) may be omitted. For example, the source video image PR1 may be stored in an unencoded (or uncompressed) form in the local memory of the AR device DV1 and retrieved from this local storage in step S4 (FIG. 5), without the need for decoding (or decompression). For example, the AR device DV1 acquires 3D information DT1 (S2), acquires a source video image PR1 (S4), and performs relighting S 12 and Aggregate S 18 and proceed to the rendering step S20.

[0208] Various variations of the methods 500 and 540 may be included herein, some of which are described below by way of example.

[0209] As previously described, this application enables the generation and rendering of high-quality AR video images by performing complex relighting of real-world and virtual elements in aggregated video images, taking into account the physical and virtual light that affects the AR content. In particular, virtual lights can be reflected off of each other and virtual shadows can be generated to make virtual objects appear to visually interact. Physical elements can also be made to appear as if they are exposed to virtual light emitted from virtual light sources that are combined into the real-world scene as augmented content.

[0210] As previously mentioned, although variations are possible, some processing tasks can be delegated (or externalized) to a processing device DV2 external to the AR device DV1, particularly for determining physical and / or virtual lighting information, which may require significant processing and power resources. To avoid delays and ensure efficient AR content rendering, a trade-off can be achieved by performing as much of the necessary processing locally as possible at the AR device level (including aggregating real and virtual content and two types of relighting reconstruction). A solution would be for the external processing device DV2 to remotely perform all processing operations (including aggregating and relighting), which could be detrimental to the quality of the AR content rendering. In particular, a large amount of data (corresponding to the aggregated video image including the relighting effect) must be transmitted from the processing device DV2 to the AR device DV1, requiring significant bandwidth and communication resources. Furthermore, additional processing requires additional delays. In this application, limited data is transmitted from the processing device DV2 to the AR device DV1 because the processing device DV2 does not need to return data corresponding to the source video image PR1.

[0211] Exemplary embodiments of methods 500 and 504 are described below with reference to FIGS.

[0212] In an embodiment, the processing device DV2 also transmits 3D coordinates associated with the virtual lighting information DT6 and the physical lighting information DT4 to the AR device DV1 (S48, FIG. 5). For example, as described above, the processing device DV2 provides the 3D coordinates as part of the scene description document (or file) SD1. The AR device DV1 then transmits the virtual lighting information DT6 and the physical lighting information DT4 to the AR device DV1 based on the 3D coordinates (spatialization information). 6 and the spatial position where the physical lighting information DT4 is applied to the real-world scene can be determined. Therefore, the virtual relighting effect RL1 and the mixed relighting effect RL2 can be determined based on the spatial positions of the virtual lighting information DT6 and the physical lighting information DT4 (S14, S16).

[0213] It should be noted that, as mentioned above, the AR device DV1 can receive the physical lighting information DT4 and the virtual lighting information DT6 (and environment map data DT5 and / or DT7 may also be present) from the processing device DV2 in various ways in the acquisition step S10 (e.g., in a scene description document SD1, in a file, or as part of a bitstream of encoded video data). In the following examples, it is assumed that a scene description document SD1 is used to transmit the physical lighting information DT4 and / or the virtual lighting information DT6, although variations are possible in which these items are transmitted in a file or as part of the bitstream BT2 of encoded video data.

[0214] In an embodiment, the processing device DV2 generates a scene description document SD1 (S48, FIG. 5) that includes at least one syntax element representing the virtual lighting information DT6 and the physical lighting information DT4, and at least one indicator (or descriptor) IT1 indicating the presence of the virtual lighting information DT6 and / or the lighting information DT4 in the scene description document SD1. Examples of syntax elements are described in more detail below. The at least one indicator IT1 is used to signal (or define) the type of lighting information (physical and / or virtual) that is carried in the scene description document SD1.

[0215] Therefore, the AR device DV1 can detect, based on the indicator IT1 of the received scene description document SD1, that said scene description document SD1 includes physical lighting information and / or virtual lighting information, and can therefore analyze the scene description document SD1 and extract therefrom physical lighting information DT4 and / or virtual lighting information DT6 for subsequent processing in the re-lighting step S15.

[0216] More specifically, in an embodiment, the scene description document SD1 generated by the processing device DV2 includes first lighting information associated with the first indicator IT1 (environment map data DT5 may also be present). The scene description document may include second lighting information that may be associated with the second indicator IT2, although variations are possible in which the second lighting information DT6 is acquired by the AR device DV1 in other ways. Thus, based on the associated first indicator IT1 present in the received scene description document SD1, the AR device DV1 can detect that the first lighting information has a type of physical lighting information and use (or interpret) this first lighting information as physical lighting information. That is, based on the first indicator IT1 in the received bitstream BT2, the AR device DV1 can recognize (or interpret) the first lighting information as physical lighting information. Similarly, based on the second indicator IT2 present in the received bitstream BT2, the AR device DV1 can detect that the second lighting information has a type of virtual lighting information and use (or interpret) this second lighting information as virtual lighting information. That is, based on the second indicator IT2 in the received bitstream BT2, the second lighting information can be recognized (or interpreted) as virtual lighting information.

[0217] In this case, the AR device DV1 performs a relighting process S based solely on the information provided in the scene description document SD1. 12 This simplifies the process from the AR application's perspective and allows for the creation of an interface that is interoperable with different ecosystems based on the scene description document SD1. In this way, any service that provides distant lighting estimation to an AR application can be used to provide physical relighting information DT4 and virtual relighting information DT6 via the scene description document SD1 without the need to integrate with an API.

[0218] In an embodiment, the scene description document SD1 includes an indicator indicating additional information, such as the type of physical and / or virtual light source defined in the physical lighting information DT4 and the virtual lighting information DT6, respectively. For example, the AR device DV1 may detect, based on the indicator present in the received scene description document SD1, that the physical (or virtual) lighting information defines at least a predetermined type of light source (e.g., a punctual light source, an ambient light source, etc.).

[0219] In an embodiment, the processing device DV2 inserts the scene description document SD1 into the bitstream BT2 transmitted to the AR device DV1 in S48 (FIG. 5), as described above. The AR device DV1 obtains the scene description document SD1 by decoding the received bitstream BT2 in S14. Then, as already described, the AR device DV1 can process the lighting information included in the scene description document SD1.

[0220] In the embodiment, the virtual lighting information DT6 is not transmitted from the processing device DV2 in the scene description document SD1 and is instead retrieved by any other appropriate means. The AR device DV1 can acquire the virtual lighting information DT6 separately from the received scene description document SD1, for example, via an API of the AR device DV1. For example, the virtual lighting information DT6 can be received by the API of an AR content rendering application executed by the AR device DV1. For example, a second scene description document including the virtual lighting information DT6 may be distributed to the AR device DV1 via a content delivery network (CDN) using the HTTP protocol and received by the AR device DV1 (S10). This scene description document SD1 may also be encapsulated in a file format structure (for example, an ISO Basic Media File Format structure) and similarly retrieved via the HTTP protocol.

[0221] In an embodiment, SD1 received by AR device DV1 from processing device DV2 (S10) also includes descriptor IT3. For example, this scene description document is transmitted by processing device DV2 together with bitstream (or data stream) BT2 (S48, FIG. 5). As already described, this bitstream (or data stream) BT2 may include physical lighting information DT4 and virtual lighting information DT6 (and, if present, environment map data DT5). In this case, bitstream BT2 does not include the scene description document provided to AR device DV1. Then, AR device DV1 can access the physical lighting information DT4 and virtual lighting information DT6 included in the received bitstream BT2 based on descriptor IT3. That is, the scene description document provides AR device DV1 with only a means to access related data, but does not include the related data itself.

[0222] The bitstream BT2 can be obtained by using an accessor, which is a general way of describing structured data that can be accessed by an AR application of the AR device DV1. The descriptor IT3 can describe the lighting information contained in the bitstream BT2 and / or describe the means for accessing the lighting information.

[0223] In an embodiment, the AR device DV1 obtains environment map data DT5 and / or DT7 by decoding encoded video image data included in the received bitstream BT2 (S10, FIG. 5). Then, the AR device DV1 detects that the received bitstream BT2 carries virtual lighting information DT6 and physical lighting information DT4 as metadata based on at least one indicator IT1 included in the received bitstream BT2. The AR device DV1 extracts the virtual lighting information DT6 and physical lighting information DT4 from the received bitstream BT2 (obtaining step S14).

[0224] In this way, instead of declaring multiple buffers and buffer views for each parameter and light type (punctual light, ambient light, etc.), all relevant physical and virtual lighting information can be encapsulated and transmitted in a video bitstream BT2. This video bitstream BT2 can include environment map data DT5 in the video data, and physical lighting information DT4 and virtual lighting information DT6 as metadata (which may be included in an SEI message in the case of AVC, HEVC, VVC, EVC video codecs).

[0225] The following describes embodiments of the present application using an example of an extended standard representation of a scene description document and an example of a signaling syntax for signaling the type of associated data transmitted by the processing device DV2 to the AR device DV1 (S48, FIG. 5), each of which can be used to implement any of the embodiments described herein.

[0226] Because existing scene description documents (e.g., USD, glTF, see below) do not distinguish between real and virtual light sources in an AR scene, the present application can extend any of these to represent various items of lighting information (static and / or dynamic) and associated information (designators) that indicate the type or nature (i.e., real or virtual) of each item. Based on this information, the AR device DV1 (FIG. 5) can re-light the AR scene taking into account both the physical lighting information DT4 and the virtual lighting information DT6.

[0227] More specifically, a 3D scene can be a complex description of 3D objects with physical properties that move in a virtual environment. These scenes can be represented in several languages ​​as files called scene description documents. For example, the Khronos Group® has standardized the glTF (a derivative abbreviation for Graphic Language Transmission Format) specification. A glTF document can describe 3D objects and entire 3D scenes. Depending on the format defined in the glTF standard, this document can be saved as a JSON document (i.e., plain text) or a binary file.

[0228] The Khronos extension named "KHR_lights_punctual" is a Khronos groups extension to glTF 2.0. This extension allows you to describe light sources in a glTF scene. As indicated by the extension name and described in the extension description, the lights have a punctual nature. This extension defines three types of "punctual" lights: directional lights, point lights, and spot lights. Punctual lights are defined as parameterized infinitesimal points that emit light with a well-defined direction and intensity.

[0229] According to the glTF 2.0 specification, the Khronos registry for extensions also includes a multi-vendor extension (provided by Microsoft® and Adobe®) called "EXT_lights_image_based" that allows adding environment map data and spherical harmonics to scene description documents.

[0230] The Khronos glTF specification specifies extensions to the MPEG-I Scene Description (SD) published in ISO / IEC DIS 23090-14. These extensions allow MPEG media data (e.g., MPEG-DASH Media Rendering Description (MPD) for self-adaptive stream content) to be integrated into the glTF format for scene description. SD defines an architecture for requesting, accessing, synchronizing, and rendering media contained in a scene description document (i.e., a glTF scene).

[0231] Here, in the embodiment of the present application, an extension of the glTF scene description document specification and a change to the MPEG-I scene description standard (extended glTF specification standard) will be described.

[0232] In an embodiment, for EXT_lights_image_based, a new parameter (e.g. called "nature") can be introduced to indicate the nature of the described lighting information. This can correspond to physical lighting information, virtual lighting information, or a mix of both. Here is an example of an extended attribute table modified for this purpose (changes in yellow): [Table 1]

[0233] Here is an example code segment showing the new parameters (not the complete glTF file, just the relevant parts): In this example, the parameter "nature" indicates that the lighting information corresponds to physical lighting information. "extensions": { "EXT_lights_image_based" : { "lights": [ { "intensity": 1.0, "rotation": [0, 0, 0, 1], "irradianceCoefficients": [...3 x 9 array of floats...], "specularImageSize": 256, "specularImages": [ [... 6 cube faces for mip 0 ...], [... 6 cube faces for mip 1 ...], ... [... 6 cube faces for mip n ...] ], "nature":"physical" } ] } }

[0234] We can make a similar extension to KHR_lights_punctual and change the attribute table as follows: [Table 2]

[0235] Note that because the data structure describes a unique light source, a mix of physical and virtual sources may not occur frequently, so variations including a mix are possible, although omitted in the above example. However, unknown situations can be useful, and a signal can be issued when the source properties cannot be defined.

[0236] Here is an example code segment showing the new parameters (not the complete glTF file, just the relevant parts): In this example, the parameter "nature" indicates that the lighting information corresponds to virtual lighting information. "extensions": { "KHR_lights_punctual" : { "lights": [ { "color": [ 1.0, 1.0, 1.0 ], "type": "directional", "nature":"virtual" } ] } }

[0237] As shown in the above example, by modifying existing glTF extensions that describe lighting information, applications can differentiate the rendering process between virtual and physical lighting information, but for the rendering process associated with this glTF scene description document, this information remains static.

[0238] In an embodiment, the local lighting conditions in a user's room may change over time (e.g., as clouds pass by, the sun rises, etc.). In these cases, the scene description document can be updated when these conditions occur. However, changing the frequency or granularity of updates would result in multiple document updates (hundreds per second), which would exponentially increase the scene document generation and application analysis, resulting in excessive interruptions and power consumption. Instead, the time-related lighting information can be passed to the application as a data stream. For the application, the scene description document simply provides the application with a means to access this data stream, but (as already explained) is not the data itself.

[0239] Therefore, it is possible to define a new glTF extension that provides lighting information data as part of a data stream. This data stream can be obtained using accessors, which are a general way of describing structured data that applications can access. However, because our goal is to provide data that changes over time, we must use an existing MPEG extension called MPEG_buffer_circular, which provides general access to timing data. Therefore, we propose a new extension called "MPEG_lighting_information" as part of the MPEG-I SD standard.

[0240] Definition of the top-level object for the MPEG_lights_punctual extension: [Table 3]

[0241] Starting example glTF file: … "extensions":{ "MPEG_media": { "media": [ { "alternatives": [ { "mimeType":"application / octet-stream", "uri": "punctual_light_color_information.bin" } ], "loop": true, "name": "media_0" }, { "alternatives": [ { "mimeType":"application / octet-stream", "uri": "punctual_light_intensity_information.bin" } ], "loop": true, "name": "media_1" }, { "alternatives": [ { "mimeType":"application / octet-stream", "uri": "punctual_light_range_information.bin" } ], "loop": true, "name": "media_2" }, ] }, "MPEG_lights_punctual" : { "lights": [ { "type": "directional", "nature":"virtual", "color_accessor": 0, "intensity_accessor": 1 "range_accessor": 2 } ] } }, … "accessors": [ { / / Color information "componentType": 5126, / / This means floating "count": 1024, "extensions":{ "MPEG_accessor_timed": { "bufferView": 0, "immutable": true, "suggestedUpdateRate": 25 } }, "type": "VEC3" / / This means a vector of three components }, { / / Strength information "componentType": 5126, / / This means floating "count": 1024, "extensions":{ "MPEG_accessor_timed": { "bufferView": 1, "immutable": true, "suggestedUpdateRate": 25 } }, "type": "SCALAR" / / This means there is only one scalar (i.e. one value) }, { / / Range information "componentType": 5126, / / This means floating "count": 1024, "extensions":{ "MPEG_accessor_timed": { "bufferView": 2, "immutable": true, "suggestedUpdateRate": 25 } }, (…) } } } ], … End of sample glTF file

[0242] Optionally, MPEG_lights_punctual may refer to the current light given in KHR_punctual_light. In this case, only the light index and accessor are given. In this case, the value present in the KHR_punctual_light structure must be the value of the first sample in the timing data sequence. The extension is as follows: [Table 4]

[0243] Note that a limitation of the current AR framework and existing Khronos extensions is that the ambient lighting is global to the entire scene, and spatializing the ambient lighting information would be more realistic (as mentioned above).

[0244] Therefore, in an embodiment, the new MPEG_lights_image_based extension also provides spatial information: each data structure is associated with a 3D coordinate in the scene. At rendering time, the AR application can take into account the distance from each light information to render a relighting effect. For this effect, we have introduced the following "position" parameter:

[0245] glTF extensions for dynamic lighting information in video streams: [Table 5]

[0246] Instead of declaring multiple buffers and buffer views for each parameter and light type (punctual and ambient), all this lighting information can be packaged and stored in a video bitstream, which can include the video data, the environment map, and all the lighting information as metadata (potentially included in the SEI message for AVC, HEVC, VVC, and EVC video codecs).

[0247] In such cases, the application needs to be able to determine the existence of such a video stream that contains all this information, so MPEG_lights_video can be defined as follows: [Table 6]

[0248] Starting sample glTF file … "extensions": { "MPEG_media": { "media": [ { "alternatives": [ { "mimeType": "video / mp4", "uri": " lights_video.mp4" } ], "loop": true, "name": "media_0" }, ] }, "MPEG_lights_video" : { "media": 0 / / Indicates that item index 0 in the MPEG_media media list contains lighting information (punctual and ambient) } }, … "accessors": [ { / / Mirror image information "componentType": 5121, / / This means an unsigned byte (8 bits) "count": 1048576 / / Number of samples in one frame "extensions": { "MPEG_accessor_timed": { "bufferView": 3, "immutable": true, "suggestedUpdateRate": 25 } }, "type": "VEC3" / / This means a vector of three components } ] … "bufferViews": [ { "buffer": 0, "byteLength": 32 }, ] "buffers": [ { "byteLength": 15728732, "extensions": { "MPEG_buffer_circular": { "count": 5, "headerLength": 12, "media":0, } } } ], … Ending a glTF file

[0249] For all the above examples, the MPEG media buffer can also be referenced using a URI to the RTP stream, in which case all signaled information can be obtained by any transmission protocol or from an external sender rather than from a local file that may be downloaded before starting the AR rendering session.

[0250] In an embodiment, the physical lighting information DT4 and virtual lighting information DT6 (FIG. 5) may come from simple binary files, where the video bitstream BT2 may contain SEI messages, or an ISOBMFF file (e.g., an mp4 file) may come from an external source (an input RTP stream), where each type of data has a different track. In the case of an RTP stream, a new RTP payload can be defined to store the lighting information.

[0251] It should be noted that the present application also relates to an AR device configured to perform the method 400 for AR content rendering (FIG. 4) as described above. In particular, the AR device may include suitable means (or units, or modules) configured to perform each step of the method 400 of any one of the embodiments of the present application.

[0252] The present application also relates to an AR device DV1 and a processing device DV2 configured to perform the methods 500 and 540, respectively, as described above (FIGS. 5-9). In particular, the AR device DV1 and the processing device DV2 may include suitable means (or units, or modules) configured to perform the steps of the methods 500 and 540, respectively, of any one of the embodiments of the present application.

[0253] In the following, specific embodiments of the AR device DV1 and the processing device DV2 will be described.

[0254] 10 illustrates an exemplary embodiment of a system SY1 including an AR device DV1 and a processing device DV2 configured to cooperate with each other to enable AR content rendering as described above, with particular reference to different example methods 500 and 540 illustrated in FIGS.

[0255] The AR device DV1 may be or may include, for example, an AR Google, a smartphone, a computer, a tablet, headphones for rendering AR content, or any other suitable device or apparatus configured to perform steps of the method for rendering AR content according to any embodiment of the present application.

[0256] The AR device DV1 may include an image acquisition unit 1002 for acquiring a source video image PR1 (e.g., from a capture device DV3), and a video encoder 1004 for encoding the source video image PR1 as an encoded bitstream BT1.

[0257] The AR device DV1 may also include a 3D information acquisition unit 1006 for acquiring 3D information DT1 representing at least one virtual object OB2 (e.g., receiving and decoding said DT1 from a content server DV4), and an AR generation unit 1008 including a relighting unit 1010 and an aggregator unit 1012. The relighting unit 1010 is for relighting the real-world scene defined by the source video image PR1 and for relighting the at least one virtual object OB2 defined by the 3D information DT1, as described above. The aggregator unit 1012 is for generating an AR video image PR3 by aggregating the relighted real-world scene acquired by the relighting unit 1010 and the relighted at least one virtual object OB2, as described above.

[0258] The AR device DV1 may further include an AR rendering unit 1014 for rendering the AR video image PR3 generated by the AR generation unit 1008.

[0259] As shown in FIG. 10, the processing device DV2 may be a high-end workstation, a mobile device, an edge processing unit, a server, etc., or any other suitable device or apparatus configured to perform steps of the method for realizing AR content rendering according to any embodiment of the present application.

[0260] The processing device DV2 may include a video decoder 1040 for decoding the encoded bitstream BT1 received from the AR device DV1, and a lighting information unit 1042 for obtaining or determining physical lighting information DT4 and virtual lighting information DT6 as described above.

[0261] The capture device DV3 is configured to capture the source video image PR1, for example by means of any suitable capture sensor. In an embodiment, the AR device DV1 and the capture device DV3 form the same device.

[0262] The content server DV4 may include a storage means for acquiring / storing the 3D information DT1.

[0263] FIG. 11 illustrates an exemplary schematic block diagram for implementing a system 1100 according to aspects and embodiments.

[0264] System 1100 may be implemented as one or more devices and includes various components described below. In various embodiments, system 1100 may be configured to implement one or more aspects described herein. For example, system 1100 may be configured to execute a method for rendering AR content or a method for implementing AR content rendering of any one of the embodiments described above. Thus, system 1100 may constitute an AR device or a processing device within the meaning of the present application.

[0265] Examples of devices that may comprise all or part of system 1100 include a personal computer, a laptop computer, a smartphone, a tablet, a digital multimedia set-top box, a digital television receiver, a personal video recording system, a connected home appliance, a connected car and its associated processing system, a head-mounted display (HMD, see-through glasses), a projector, a "cave" (a system including multiple displays), a server, a video encoder, a video decoder, a post-processor that processes output from a video decoder, a pre-processor that provides input to a video encoder, a web server, a video server (such as a broadcast server, a video-on-demand server, or a network server), a static or video camera, an encoding or decoding chip, or any other communications device. Elements of system 1100 can be implemented singly or in combination on a single integrated circuit (IC), multiple ICs and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1100 can be distributed across multiple ICs and / or discrete components. In various embodiments, system 1100 can be communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or dedicated input and / or output ports.

[0266] The system 1100 includes at least one processor 1110, configured to execute instructions loaded therein to, for example, implement aspects described herein. The processor 1110 may include embedded memory, input / output interfaces, and various other circuits known in the art. The system 1100 may include at least one memory 1120 (e.g., a volatile memory device and / or a non-volatile memory device). The system 1100 may include a storage device 1140, including non-volatile and / or volatile memory, including, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, magnetic disk drives, and / or optical disk drives. By way of non-limiting example, the storage device 1140 may include an internal storage device, an additional storage device, and / or a network-accessible storage device.

[0267] The system 1100 may include an encoder / decoder module 1130 configured to, for example, process data to provide encoded / decoded video image data, and the encoder / decoder module 1130 may include its own processor and memory. The encoder / decoder module 1130 may represent a module(s) or unit(s) included in a device to perform encoding and / or decoding functions. As is known, a device may include either one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 1130 may be implemented as a separate element of the system 1100 or may be coupled to the processor 1110 as a combination of hardware and software known to those skilled in the art.

[0268] Program code loaded into the processor 1110 or the encoder / decoder 1130 to perform aspects described herein may be stored in the storage device 1140 and then loaded into the memory 1120 and executed by the processor 1110. According to various embodiments, during execution of the processes described herein, one or more of the processor 1110, the memory 1120, the storage device 1140, and the encoder / decoder module 1130 may store one or more of the following items: video images, information data for encoding / decoding video image data, bitstreams, matrices, variables, and equations, formulas, logic for operations, and intermediate or final results of operations.

[0269] In some embodiments, memory within the processor 1110 and / or encoder / decoder module 1130 can be used to store instructions and provide working memory for processes performed during encoding or decoding.

[0270] However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 1110 or the encoder / decoder module 1130) is used for one or more of these functions. The external memory may be memory 1120 and / or storage device 1140, such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, fast external dynamic volatile memory such as RAM may be utilized as working memory for video encoding / decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, and also referred to as MPEG-2 Video), AVC, HEVC, EVC, VVC, AVI, etc.

[0271] As shown in block 1190, inputs can be provided to the elements of system 1100 via various input devices, including, but not limited to, (i) an RF section capable of receiving RF signals transmitted wirelessly, such as by a broadcast station, (ii) a combined input terminal, (iii) a USB input terminal, (iv) an HDMI input terminal, and (v) if the invention is implemented in the automotive field, a bus, such as a CAN (Controller Area Network), CAN FD (Flexible Data Rate for Controller Area Network), FlexRay (ISO 17458), or Ethernet (ISO / IEC 802-3) bus.

[0272] In various embodiments, the input devices of block 1190 have associated corresponding input processing elements, as known in the art. For example, the RF section may be associated with each of the following required elements: (i) selecting a desired frequency (also called signal selection, or limiting the signal to a frequency band); (ii) downconverting the selected signal; (iii) selecting a signal frequency band by controlling the frequency band back to a narrower frequency band, (e.g., referred to as a channel in some embodiments); (iv) demodulating the downconverted and frequency-band-limited signals; (v) performing error correction; and (vi) demultiplexing to select a desired data packet flow. The RF section of various embodiments may include one or more elements that perform these functions, such as a frequency selector, a signal selector, a frequency band limiter, a channel selector, a filter, a downconverter, a demodulator, an error correction device, and a demultiplexer. The RF section may also include a tuner that performs each of these functions, including, for example, downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband.

[0273] In one set-top box embodiment, the RF section and its associated input processing elements can receive RF signals transmitted over a wired (e.g., cable) medium, after which the RF section can perform frequency selection by filtering, downconverting, and re-filtering to obtain a desired frequency band.

[0274] Various embodiments may rearrange the order of these (and other) elements, eliminate some of these elements, and / or add other elements that perform similar or different functions.

[0275] Adding elements can include inserting elements such as amplifiers and analog-to-digital converters between existing elements. In various embodiments, the RF section can include an antenna.

[0276] Additionally, the USB and / or HDMI terminals allow you to connect your system via USB and / or HDMI connections. 1100 11. The input signal may include a corresponding interface processor for connecting the input signal to other electronic devices. It should be noted that, when desired, aspects of the input processing (e.g., Reed-Solomon error correction) may be implemented, for example, in a separate input processing IC or within processor 1110. Thus, it should be understood that, when desired, aspects of the USB or HDMI interface processing may be implemented in a separate interface IC or within processor 1110. Upon demodulation, the error corrected and demultiplexed stream may be provided to various processing elements, including processor 1110 and encoder / decoder 1130 operating in conjunction with memory and storage elements, to process the data stream when desired for display on an output device.

[0277] The various elements of system 1100 may be provided within a unitary housing, within which a suitable connection layout 1190 (e.g., internal buses known in the art, including I2C buses, wires, and printed circuit boards) may be used to connect the elements to one another and transmit data therebetween.

[0278] System 1100 may include a communication interface 1150 such that it can communicate with other devices over a communication channel 1151. Communication interface 1150 includes, but is not limited to, a transceiver configured to transmit and receive data over communication channel 1151. Communication interface 1150 may include, but is not limited to, a modem or a network card, and communication channel 1151 may be implemented within a wired and / or wireless medium, for example.

[0279] In various embodiments, a Wi-Fi network, such as IEEE 802.11, can be used to stream data to system 1100. The Wi-Fi signal in these embodiments can be received via communication channel 1151 and communication interface 1150 suitable for Wi-Fi communication. Communication channel 1151 in these embodiments can typically be connected to an access point or router that provides access to external networks, including the Internet, allowing streaming applications and other over-the-top wireless communications.

[0280] Other embodiments may provide streaming data to system 1100 using a set-top box, which carries the data through an HDMI connection in input block 1190.

[0281] In some embodiments, the RF connection of input block 1190 is used to provide streaming data to system 1100 .

[0282] The streaming data can be used as a form of signaling information, such as component transform information DT1 (as described above), used by the system 1100. The signaling information can include information about the bitstream B and / or the number of pixels in the video image, and / or any encoding / decoding setting parameters.

[0283] It should be noted that signaling can be achieved in various manners, for example, in various embodiments, one or more syntax elements, flags, etc. can be used to send signaling information to corresponding decoders.

[0284] System 1100 can provide output signals to various output devices, including a display 1161, speakers 1171, and other peripherals 1181. In various example embodiments, other peripherals 1181 can include one or more of a separate DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 1100.

[0285] In various embodiments, control signals may be sent between the system 1100 and the display 1161, speakers, and other devices using AV.Link (Audio / Video Link), CEC (Consumer Electronic Control), or other communication protocol signaling that enables device-to-device control. 1171 or other peripheral devices 1181, with or without user intervention.

[0286] Output devices can be communicatively connected to the system 1100 via dedicated connections by corresponding interfaces 1160, 1170 and 1180.

[0287] Optionally, an output device can be connected to system 1100 using communication channel 1151 via communication interface 1150. Display 1161 and speakers 1171 can be integrated into a single unit along with other components of system 1100, such as an electronic device (e.g., a television).

[0288] In various embodiments, the display interface 1160 can include a display driver such as a timing controller (T Con) chip.

[0289] For example, if the RF portion of input 1190 is part of a separate set-top box, display 1161 and speakers 1171 are optionally separate from one or more of the other components. In various embodiments where display 1161 and speakers 1171 may be external components, the output signals may be provided via dedicated output connections (including, for example, an HDMI port, a USB port, or a COMP output terminal).

[0290] 2-9, various methods are described herein, each of which includes one or more steps or actions to achieve the described method. Unless a specific order of steps or actions is required for the precise operation of the method, the order and / or use of specific steps and / or actions can be modified or combined.

[0291] While some examples have been described with reference to block diagrams and / or operational flowcharts, each block represents a circuit element, module, or portion containing one or more executable instruction codes for implementing a specified logic function(s). It should be noted that in other embodiments, the function(s) shown in the blocks may not occur in the order shown. For example, depending on the functionality involved, two blocks shown in succession may in fact be executed essentially in parallel, or the blocks may be executed in the reverse order.

[0292] For example, the embodiments and aspects described herein may be implemented in a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if discussed only in the context of a single type of embodiment (e.g., discussed only as a method), embodiments of the discussed features may be implemented in other forms (e.g., an apparatus or a computer program).

[0293] The methods may be implemented in, for example, a processor, which generally refers to a processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device, etc. Processors further include communication devices.

[0294] Additionally, methods may be implemented with instructions executed by a processor, and such instructions (and / or data values ​​produced by embodiments) may be stored on a computer-readable storage medium, such as storage device 1140 (FIG. 11). A computer-readable storage medium may take the form of a computer-readable program product having computer-readable program code embodied in, and embodied therein, executable by, a computer. Given their inherent ability to store information thereon and retrieve information provided thereby, computer-readable storage media as used herein may be considered non-transitory storage media. A computer-readable storage medium may be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Although the following provides more specific examples of computer-readable storage media to which the present embodiment can be applied, it should be understood that these are merely illustrative and not an exhaustive list, as would be readily apparent to one skilled in the art: portable computer floppy disks, hard disks, read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0295] The instructions may create an application tangibly embodied on a processor-readable medium.

[0296] For example, instructions may reside in hardware, firmware, software, or a combination thereof. For example, instructions may be found in an operating system, a standalone application, or a combination of both. A processor may therefore be characterized as, for example, a device configured to perform a process or a device that includes a processor-readable medium (e.g., a storage device) having instructions for performing a process. Also, in addition to or in place of instructions, the processor-readable medium may store data values ​​produced by an embodiment.

[0297] The device may be implemented in, for example, appropriate hardware, software, and firmware. Examples of such devices include personal computers, laptop computers, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected consumer electronics products, head-mounted displays (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process output from video decoders, pre-processors that provide input to video encoders, web servers, set-top boxes, and any other device for processing video images or other communications devices. The device may also be mobile and mounted in a moving vehicle.

[0298] The computer software may be implemented in the processor 1110, in hardware, or in a combination of hardware and software. By way of non-limiting example, an embodiment may be implemented in one or more integrated circuits. The memory 1120 may be of any type compatible with the technology environment and may be implemented in any suitable data storage technology (by way of non-limiting example, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memories, and removable memories). By way of non-limiting example, the processor 1110 may be of any type compatible with the technology environment and may cover one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0299] As will be apparent to those skilled in the art based on this disclosure, embodiments can generate a variety of signals shaped to carry, for example, storable or transmittable information. Information can include, for example, instructions for performing a method or data generated by one of the described embodiments. For example, a signal can be shaped to carry a bit stream of the described embodiments. The signal can be shaped, for example, as an electromagnetic wave (e.g., a radio frequency portion of the frequency spectrum) or a baseband signal. Shaping can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is well known, the signal can be transmitted over different wired or wireless links. The signal can be stored on a processor-readable medium.

[0300] The terms used herein are used only to describe particular embodiments and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" as used herein also include the plural. Furthermore, as used herein, the terms "include / comprise" and / or "including / comprising" may indicate the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. Furthermore, when an element is referred to as being "responsive to," "connected to," or "associated with" another element, it may be directly responsive to, connected to, or related to the other element, or intermediate elements may be present. Conversely, when an element is referred to as being "directly responsive to," "directly connected to," or "directly related to" another element, intermediate elements are not present.

[0301] For example, in the case of "A / B," "A and / or B," and "at least one of A and B," the use of any one of the symbols / terms " / ," "and / or," and "at least one" is intended to cover the selection of the first listed alternative (A), or the selection of the second listed alternative (B), or the selection of two alternatives (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such language is intended to cover the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second listed alternatives (A and B), or the selection of the first and third listed alternatives (A and C), or the selection of only the second and third listed alternatives (B and C), or the selection of all three alternatives (A, B, and C). This can be extended to any number of listed items, as would be apparent to one skilled in the art.

[0302] Various values ​​may be used in this application, and the specific values ​​are exemplary and the described aspects are not limited to these specific values.

[0303] It should be noted that terms such as "first," "second," etc. may be used to describe various elements herein, but are not limited to these terms. These terms are used only to distinguish one element from another. For example, a first element can be referred to as a second element, and similarly, a second element can be referred to as a first element, without departing from the teachings of the present application. No ordering is implied between a first element and a second element.

[0304] References to "one example" or "example" or "one embodiment" or "embodiment" and other variations are often used to convey that a particular feature, structure, characteristic, etc. (as described in connection with the example / embodiment) is included in at least one example / embodiment. Thus, appearances of the terms "in one example" or "in an example" or "in some examples" or "in one embodiment" or "in an embodiment" and any other variations appearing in various places in this application do not necessarily refer to the same example.

[0305] Similarly, references herein to "according to an embodiment" or "in an embodiment" and other variations are often used to convey that a particular feature, structure, or characteristic (as described in connection with an embodiment) may be included in at least one embodiment. Thus, the appearances of "according to an embodiment" or "in an embodiment" in various places in this application do not necessarily refer to the same embodiment, nor are independent or alternative embodiments mutually exclusive of other embodiments.

[0306] The reference numerals of the drawings appearing in the claims are for illustrative purposes only and do not limit the scope of the claims. Although not explicitly stated, the present embodiments / examples and modifications can be used in any combination or partial combination.

[0307] It is understood that when a figure is shown as a flow chart, a block diagram of the corresponding apparatus is also provided. Similarly, it is understood that when a figure is shown as a block diagram, a flow chart of the corresponding method / process is also provided.

[0308] Some figures include arrows that indicate the primary direction of communication in a communication path, however, it should be understood that communication can also occur in the direction opposite to the illustrated arrow.

[0309] Various embodiments relate to decoding. As used herein, "decoding" may cover all or part of the processes performed, for example, on received video images (which may include a received bitstream encoding one or more video images) to generate a final output suitable for display or further processing in the reconstructed video domain. In various embodiments, such processes may include one or more of the processes typically performed by a decoder. In various embodiments, for example, such processes may optionally include the processes performed by the decoders of various embodiments described herein.

[0310] As a further example, in one embodiment, "decoding" may refer to only inverse quantization, in one embodiment, "decoding" may refer to entropy decoding, in another embodiment, "decoding" may refer to differential decoding, and in another embodiment, "decoding" may refer to a combination of inverse quantization, entropy decoding, and differential decoding. Depending on the context specifically described, it is obvious and understandable to one skilled in the art whether the term "decoding process" refers specifically to a subset of operations or to a more general decoding process.

[0311] Various embodiments relate to encoding. As at least in the manner discussed above regarding "decoding," "encoding" as used herein can cover all or part of the process of, for example, processing input video images to generate and output a bitstream. In various embodiments, this type of process includes one or more of the processes typically performed by an encoder. In various embodiments, such processes can include, or alternatively include, the processes performed by the encoder of each embodiment described herein.

[0312] As a further example, in one embodiment, "encoding" may refer only to quantization, in one embodiment, "encoding" may refer only to entropy coding, in another embodiment, "encoding" may refer only to differential coding, and in another embodiment, "encoding" may refer to a combination of quantization, differential coding, and entropy coding. Depending on the context in which a particular description is made, it may be clear and easy to understand for one skilled in the art whether the term "encoding process" refers specifically to a subset of operations or to a more general encoding process.

[0313] This application may also refer to "obtaining" various pieces of information, which may include one or more of estimating information, calculating information, predicting information, or examining information from memory, processing information, moving information, copying information, deleting information, calculating information, determining information, predicting information, or estimating information.

[0314] Additionally, this application may refer to "receiving" various information. Receiving information may include, for example, one or more of accessing the information or receiving the information from a communications network.

[0315] Also, as used herein, the term "signal" refers to instructing a corresponding decoder to do something specific. For example, in some embodiments, an encoder sends a signal to notify specific information, such as encoding parameters or encoded video image data. In this manner, in some embodiments, the same parameters can be used at the encoder and decoder sides. Thus, for example, the encoder can send specific parameters to the decoder (explicit signaling), allowing the decoder to use the same specific parameters. Conversely, if the decoder has specific parameters and other parameters, signaling that does not require transmission (indirect signaling) can be used to inform the decoder and facilitate selection of the specific parameters. Various embodiments achieve bit savings by avoiding transmission of any actual functions. It should be appreciated that signaling can be accomplished in various ways. For example, in various embodiments, one or more grammatical elements, flags, etc., are used to transmit information to a corresponding decoder. While the above relates to the verb form of the word "signal," the word "signal" may also be used as a noun in this specification.

[0316] While several embodiments have been described above, it should be understood that various modifications may be made. For example, elements of different embodiments may be combined, supplemented, modified, or deleted to produce other embodiments. Moreover, as will be appreciated by those skilled in the art, other structures and processes may be substituted for the disclosed structures and processes, thereby producing embodiments that perform essentially the same function(s) in essentially the same way(s) to achieve at least essentially the same result(s) as the disclosed embodiments. Accordingly, these and other embodiments are contemplated herein.

Claims

1. A method (400, 500) for performing AR content rendering, comprising: - obtaining (410, S14) a virtual relighting effect (RL1) based on virtual lighting information (DT6) representing at least one virtual light source (SC2); - obtaining (412, S16) a mixed relighting effect (RL2) based on said virtual lighting information (DT6) and on physical lighting information (DT4) representing at least one physical light source (SC1); - generating (408, S18) an AR video image (PR3) by aggregating a real-world scene relit according to said virtual relighting effect (RL1) and at least one virtual object relit according to said mixed relighting effect (RL2); - rendering (S20) said AR video image (PR3), A method for performing AR content rendering.

2. generating the AR video image includes: - obtaining (410, S14) a relit real-world scene by applying said virtual relighting effect (RL1) based on a video image representative of said real-world scene; - obtaining (412, S16) at least one re-illuminated virtual object (OB2) by applying said blended re-lighting effect (RL2) based on three-dimensional (3D) information representative of said at least one virtual object; aggregating the re-illuminated real-world scene and the at least one re-illuminated virtual object into the AR video image; The method for performing AR content rendering according to claim 1 .

3. The at least one physical light source and the at least one virtual light source are respectively determined by the physical lighting information (DT4) and the virtual lighting information (DT6): - a directional light source, - a point light source, a spot light source; defined as any type of ambient light source and / or punctual light source among A method for performing AR content rendering according to claim 1 or 2.

4. - obtaining a scene description document comprising first lighting information associated with a first indicator (IT1); - recognizing said first lighting information as said physical lighting information (DT4) based on said first indicator; - acquiring the virtual lighting information (DT6) via an API of an AR content rendering application; A method for performing AR content rendering according to claim 1 or 2.

5. - further comprising a step of obtaining a scene description document comprising at least one of first lighting information associated with the first indicator (IT1) and second lighting information associated with the second indicator (IT2), The method comprises: - determining, based on said first indicator (IT1), that said first lighting information is to be used as said physical lighting information (DT4); - determining, based on said second indicator (IT2), that said second lighting information is to be used as said virtual lighting information (DT6), A method for performing AR content rendering according to claim 1 or 2.

6. - obtaining a scene description document containing the descriptors; - accessing the virtual lighting information (DT6) and the physical lighting information (DT4) contained in the data stream based on the descriptor, A method for performing AR content rendering according to claim 1 or 2.

7. - detecting, based on at least one indicator (IT1, IT2) contained in the received bitstream (BT2), that said received bitstream (BT2) carries said virtual lighting information and said physical lighting information as metadata; - extracting said virtual lighting information (DT6) and said physical lighting information (DT4) from said received bitstream (BT2), A method for performing AR content rendering according to claim 1 or 2.

8. - obtaining 3D coordinates associated with said virtual lighting information (DT6) and said physical lighting information (DT4), respectively; - determining, based on the 3D coordinates, a spatial position at which the virtual lighting information and the physical lighting information are applied relative to the real-world scene; the virtual relighting effect (RL1) and the mixed relighting effect (RL2) are determined based on a spatial position of the virtual lighting information and a spatial position of the physical lighting information. A method for performing AR content rendering according to claim 1 or 2.

9. A method (540) for realizing AR content rendering, comprising: - obtaining (S44) physical lighting information (DT4) representative of at least one physical light source (SC1) by performing a first lighting estimation based on a video image (PR1) of a real-world scene; - obtaining (S46) virtual lighting information (DT6) representing at least one virtual light source (SC2) emitting virtual light in said real-world scene; - transmitting (S48) said physical lighting information and said virtual lighting information to realize AR rendering of an AR video image (PR3), transmitting the physical lighting information (DT4) and the virtual lighting information (DT6) in a scene description document to realize AR rendering of the AR video image (PR3); the scene description document comprises at least one syntax element representing the virtual lighting information and the physical lighting information and at least one indicator (IT1, IT2) indicating the presence of the virtual lighting information and / or the physical lighting information in the scene description document; A method for realizing AR content rendering.

10. 3. A method for performing a method according to claim 1, AR device (DV1) for AR content rendering.

11. means for carrying out the method of claim 9, A processing device (DV2) for realizing AR content rendering by the AR device (DV2).

12. A computer program comprising: The computer program, when executed by one or more processors, causes the one or more processors to perform the method of claim 1 or 2 or claim 9. Computer program.

13. comprising program code instructions for carrying out the method of claim 1 or 2 or claim 9, Non-transitory storage media.

Citation Information

Patent Citations

  • Display control device, display control method, and recording medium

    WO2021131781A1