System and method for real-time ray tracing in a 3D environment - Patents.com

JP2024523729A5Active Publication Date: 2025-07-08INTERDIGITALCE PATENT HLDG SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024501143
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-16
Filing Date
2022-06-30
Publication Date
2025-07-08
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

Ray tracing techniques are computationally expensive and memory-intensive, making them difficult to implement in real-time applications on devices like mobile phones, tablets, AR glasses, and embedded cameras for video capture.

Method used

A method and device for rendering 3D scenes using a hybrid acceleration structure that combines intermediate and classical acceleration structures, optimizing rendering quality and performance by classifying objects based on parameters such as shape complexity and environmental importance, and employing ray tracing techniques to integrate virtual objects into real-world videos.

Benefits of technology

Enables high-quality, real-time rendering of complex-shaped objects with accurate light interactions, allowing seamless integration of virtual objects into real-world videos, enhancing the immersive experience without significant computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method and system for rendering a 3D scene is disclosed. A set of parameters is identified for one or more objects in a 3D scene of a video. An intermediate structure and a corresponding object are determined for each of the one or more objects in the 3D scene based on the identified set of parameters. A hybrid acceleration structure is determined based on the determined intermediate structure and a classical acceleration structure. A color contribution is determined for the hybrid acceleration structure. The 3D scene is then rendered based on the determined color contribution of the hybrid acceleration structure.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] FIELD OF THE DISCLOSURE The present disclosure relates generally to augmented reality (AR) applications. At least one embodiment relates to the placement of virtual objects within a video, such as a live video feed of a 3D environment. [Background technology]

[0002] Traditionally, ray tracing is a technique used for high-quality non-real-time graphics rendering tasks, such as creating animated movies or creating 2D images that more faithfully model the behavior of light in various materials. As an example, ray tracing is particularly suited to introducing lighting effects into rendered images. For a scene, light sources can be defined that cast light onto objects in the scene. Some objects may cause shadows in the scene by occluding other objects from the light source. Rendering using ray tracing techniques allows the effects of light sources to be rendered accurately, since ray tracing is adapted to model the behavior of light in the scene.

[0003] Ray tracing rendering techniques are often relatively computationally expensive and memory intensive to implement, especially when rendering is desired to be performed in real time. Thus, ray tracing techniques are difficult to implement on devices such as mobile phones, tablets, AR glasses for displays, and built-in cameras for video capture. The embodiments herein have been devised with the above in mind. Summary of the Invention

[0004] The present disclosure is directed to a method for rendering a 3D scene of a video, e.g., a live video feed, that can be considered for implementation on devices such as, for example, mobile phones, tablets, AR glasses for displays, and built-in cameras for video capture.

[0005] According to a first aspect of the present disclosure, there is provided a method for rendering a 3D scene in a video, comprising: identifying a set of parameters for one or more objects in a 3D scene of the video; determining an intermediate structure and a corresponding substitute object for each of the one or more objects in the 3D scene based on the identified set of parameters; determining a hybrid acceleration structure based on the determined intermediate structure and the classical acceleration structure; determining a color contribution to the determined hybrid accelerating structure; and providing a video of the 3D scene based on the determined color contribution of a hybrid acceleration structure.

[0006] The general principle of the proposed solution concerns the rendering of non-planar, glossy and / or refractive objects with complex shapes in a video 3D scene. High quality reflections for objects with complex shapes are achieved by using a hybrid acceleration structure based on a rendering quality level defined for each object in the video 3D scene.

[0007] In an embodiment, the set of parameters includes at least one of: simple shapes, distant objects, complex shapes, shapes requiring internal reflection, and objects near reflective objects.

[0008] In an embodiment, the intermediate structure includes an identifier.

[0009] In an embodiment, the identifier includes color and depth information.

[0010] In an embodiment, the intermediate structure is determined using the array camera projection matrix.

[0011] In an embodiment, the corresponding substitute object includes geometry data.

[0012] In an embodiment, the geometry data is at least one of a center position and x, y and z extension values ​​for the primitive shape.

[0013] In an embodiment, the classical acceleration structure is one of a grid structure, a bounding volume hierarchy (BVH) structure, a k-dimensional tree structure, and a binary space partitioning data structure.

[0014] According to a second aspect of the present disclosure, there is provided a device for rendering a 3D scene of a video, comprising: identifying a set of parameters for one or more objects in a 3D scene of the video; determining an intermediate structure and a corresponding substitute object for each of the one or more objects in the 3D scene based on the identified set of parameters; determining a hybrid acceleration structure based on the determined intermediate structure and the classical acceleration structure; A device is provided, the device comprising at least one processor configured to: provide a rendering of the 3D scene of a video based on the determined color contribution of a hybrid acceleration structure.

[0015] In an embodiment, the set of parameters includes at least one of: simple shapes, distant objects, complex shapes, shapes requiring internal reflection, and objects near reflective objects.

[0016] In an embodiment, the intermediate structure includes an identifier.

[0017] In an embodiment, the identifier includes color and depth information.

[0018] In an embodiment, the intermediate structure is determined using the array camera projection matrix.

[0019] In an embodiment, the corresponding substitute object includes geometry data.

[0020] In an embodiment, the geometry data is at least one of a center position and x, y and z extension values ​​for the primitive shape.

[0021] In an embodiment, the classical acceleration structure is one of a grid structure, a bounding volume hierarchy (BVH) structure, a k-dimensional tree structure, and a binary space partitioning data structure.

[0022] Some processes performed by elements of the present disclosure may be computer-implemented. As such, such elements may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be referred to generally herein as a "circuit," "module," or "system." Furthermore, such elements may take the form of a computer program product embodied in any tangible medium of expression having computer usable code embodied in the medium.

[0023] Since elements of the present disclosure can be implemented in software, the present disclosure can be embodied in any suitable carrier medium as computer readable code for provision to a programmable device. Tangible carrier media can include storage media such as floppy disks, CD-ROMs, hard disk drives, magnetic tape devices, or solid-state memory devices. Transient carrier media can include signals such as electrical signals, optical signals, acoustic signals, magnetic signals, or electromagnetic signals, e.g., microwave or RF signals. [Brief description of the drawings]

[0024] Other characteristics and advantages of the embodiments will become apparent from the following description, given by way of illustrative and non-exhaustive example, and from the accompanying drawings, in which: [Figure 1] 1 illustrates an exemplary system for rendering a 3D scene in a video, according to an embodiment of the present disclosure. [Diagram 2] 1 is a flowchart of a particular embodiment of the proposed method for rendering virtual objects of an AR (Augmented Reality) or MR (Mixed Reality) application in a real-world 3D environment. [Diagram 3] FIG. 1 illustrates the positions of a virtual camera and a real camera capturing a typical AR scene. [Figure 4] 1 is a flowchart for constructing a hybrid acceleration structure by adding an intermediate structure to a classical acceleration structure. [Diagram 5] FIG. 1 is a diagram of an exemplary classical acceleration structure including sorted primitives into spatial partitioning nodes. [Figure 6] 1 is a diagram of the resulting hybrid acceleration structure in which only blue objects are classified as low quality rendering objects for a frame of a 3D scene. [Figure 7] 1 is a diagram of the resulting hybrid acceleration structure in which both blue and green objects are classified as low quality rendering objects for a frame of a 3D scene. [Figure 8] 13 is a flowchart of steps added to a node of a classical acceleration structure to support an intermediate acceleration structure of a hybrid acceleration structure. [Figure 9] 9 is a detailed flowchart of step 830 of the flowchart shown in FIG. 8. [Figure 10] FIG. 4 shows a comparison of target proxy intersection depth and read geometry depth for the bird shown in FIG. 3. [Figure 11] FIG. 7 depicts a ray traverse of a portion of the hybrid accelerating structure of FIG. 6. [Figure 12] 7 is a diagram depicting ray trajectories of another portion of the hybrid accelerating structure of FIG. 6. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] FIG. 1 illustrates an exemplary apparatus for rendering a 3D scene in a video, according to one embodiment of the disclosure. FIG. 1 illustrates a block diagram of an exemplary system 100 in which various aspects of the exemplary embodiment may be implemented. System 100 may be incorporated as a device including various components described below and configured to execute corresponding processes. Examples of such devices include, but are not limited to, mobile devices, smartphones, tablet computers, augmented reality glasses for displays, and built-in cameras for video capture. System 100 may be communicatively coupled to other similar systems and displays via communication channels.

[0026] Various embodiments of the system 100 include at least one processor 110 configured to execute instructions loaded therein to perform various processes as described below. The processor 110 may include internal memory, input / output interfaces, and various other circuits as known in the art. The system 100 may also include at least one memory 120 (e.g., volatile memory device, non-volatile memory device). The system 100 may further include a storage device 140. The storage device may include non-volatile memory including, but not limited to, EEPROM, ROM, PROM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 140 may include, by way of non-limiting example, an internal storage device, an attached storage device, and / or a network-accessible storage device.

[0027] Program code that is loaded onto one or more processors 110 to perform the various processes described herein below may be stored in storage device 140 and then loaded onto memory 120 for execution by processor 110. According to an example embodiment, one or more of processor(s) 110, memory 120, and storage device 140 may store one or more of a variety of items, including, but not limited to, ambient images, captured input images, texture maps, texture free maps, cast shadow maps, 3D scene geometry, 3D poses of viewpoints, lighting parameters, variables, operations, and operation logic, during execution of the processes described herein below.

[0028] System 100 may also include a communication interface 150, which allows communication with other devices over a communication channel. Communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data from a communication channel. Communication interface 150 may include, but is not limited to, a modem or a network card, and the communication interface may be implemented in a wired and / or wireless medium. The various components of communication interface 150 may be connected or communicatively coupled to each other using various suitable connections, including but not limited to an internal bus, wires, and printed circuit boards (not shown).

[0029] The system 100 also includes a video capture device 160, such as a camera, coupled to the processor for capturing video images.

[0030] The system 100 also includes a video rendering device 170, such as a projector or a screen, coupled to the processor for rendering the 3D scene.

[0031] Exemplary embodiments may be implemented by the processor 110 or by computer software implemented by hardware or by a combination of hardware and software. As a non-limiting example, exemplary embodiments may be implemented by one or more integrated circuits. The memory 120 may be of any type appropriate to the technological environment and may be implemented using any suitable data storage technology, such as, as non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memories and removable memories. The processor 110 may be of any type appropriate to the technological environment and may include, as non-limiting examples, one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0032] The implementations described herein may be implemented, for example, in a method or process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single type of implementation (e.g., only discussed as a method), the implementation of the features discussed may also be embodied in other forms (e.g., an apparatus or a program). An apparatus may be implemented, for example, in appropriate hardware, software, and firmware. The method may be implemented in an apparatus, such as, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include, for example, communication devices, such as computers, mobile phones, portable / personal digital assistants ("PDAs"), tablets, head-mounted devices, and other devices facilitating virtual reality applications.

[0033] The present disclosure is applicable to Augmented Reality (AR) applications, in which virtual objects are inserted (composite) into a live video feed of a real environment using see-through devices, such as a mobile phone, tablet, AR glasses for display, and a built-in camera for video capture.

[0034] The goal is that a viewer watching such a composited video feed should be unable to distinguish between real and virtual objects. Such applications require integrating 3D virtual objects into 2D video images with consistent real-time rendering of light interactions. The insertion of virtual objects into real-world video should be as realistic as possible, taking into account several technical aspects: virtual object position, virtual object orientation (dynamic precision as if the object was hooked / fixed in real 3D space when the user moves the camera), and lighting conditions (reflection / refraction of different light sources according to the virtual material properties of the virtual object).

[0035] Interaction with both the real and virtual worlds, which are composited together to render a composite plausible video feed (mixed reality with synthetic 3D objects inserted into a 2D live video feed, frame by frame in real-time conditions), is also very important. Such interactions may include, for example, as non-limiting examples, light interactions (global illumination with real and virtual lights) and physics interactions (interactions between both real and virtual objects).

[0036] The speed of the required computation is also important: the latency between the acquisition of a single frame of a real-world scene by a video camera and the presentation of the corresponding "augmented" frame (with 3D objects) on a display device needs to be close to 0 ms, so that the viewer (who is also the person using the camera) feels an immersive experience.

[0037] 2 is a flowchart 200 of a particular embodiment of the proposed method for placing virtual objects in an AR (Augmented Reality) or MR (Mixed Reality) application in a real-world 3D environment. In this particular embodiment, the method comprises five successive steps 210-250.

[0038] In an exemplary implementation described below, the method is performed by a rendering device 170 (e.g., a smartphone, tablet, or head-mounted display). In an alternative exemplary implementation, the method is performed by a processor 110 external to the rendering device 170. In the latter case, results from the processor 110 are provided to the rendering device 170.

[0039] A diagram of a typical AR scene is shown in Figure 3. The scene shown in Figure 3 is composed of tangible objects, including real objects and virtual objects. Exemplary tangible objects that are real objects include a refractive panel with a transparent portion (glass) 310. Some exemplary tangible objects that are virtual objects include a car 320, a bird 330, and a tree 340 that have metallic / shiny non-flat surfaces.

[0040] Additional virtual objects (not shown) can be added to the scene to increase realism and aid in rendering: an invisible virtual refractive panel that represents a real transparent panel with the same size and position as refractive panel 310 is a non-limiting example.

[0041] In the following sections, a specific embodiment of the proposed method 200 of Fig. 2 for placing virtual objects of an AR (Augmented Reality) or MR (Mixed Reality) application in a real-world 3D environment is described, which provides high-quality rendering for each object in an AR scene that displays different rendering aspects.

[0042] In step 210, a set of parameters is identified for objects in the 3D scene of the video. This step is necessary for objects other than the light and the camera in the scene and is based on prior knowledge of the AR scene. Each of the objects in the scene is identified as either a real object or a virtual object in the context of AR based on the set of parameters.

[0043] The set of parameters may include, for example, a rendering quality level, an environment importance value, and an array of N integer values ​​based on the number of reflective / refractive objects in a scene of the video.

[0044] A rendering quality level is defined for each object in a scene of the video by assigning it an integer value between 1 and 3. A rendering quality level with integer value 1 indicates low quality rendering and is assigned to objects with simple shapes or distant objects. A rendering quality level with integer value 2 indicates high quality rendering and is assigned to objects with complex shapes and / or objects that require internal reflections. A rendering quality level with integer value 3 indicates both low and high quality rendering and is assigned to objects with complex shapes whose rendering quality may change during scene rendering. Non-limiting examples that may be assigned a rendering quality level with integer value 3 are objects that are closer to the main camera or reflective objects.

[0045] Environment importance relates to the color contribution of an object from the environment of the 3D scene. An integer value between 0 and 2 may be used for the environment importance value. An environment importance value of 0 indicates that the color contribution from the environment is zero or negligible. An environment importance value of 1 indicates that the color contribution from the environment is carried by reflected rays. An environment importance value of 2 indicates that the color contribution from the environment is carried by reflected and refracted rays. Identifying an appropriate object's environment importance value is straightforward.

[0046] The size of the array of N integer values ​​corresponds to the number of reflective / refractive objects (e.g., objects with environmental importance different from 0) in the scene. The larger the integer value of N, the better the rendering of the selected object with respect to the associated reflective / refractive objects. The integer value of N corresponds to the number of intermediate (middle field) images to be generated for the selected object with respect to the associated reflective / refractive objects.

[0047] The array of N integer values ​​is set based on the selected object's geometric complexity and its position and distance relative to the associated reflective / refractive object. If these parameters are changed, the N integer values ​​may change. The array of N integer values ​​is used to optimize the tradeoff between rendering quality and performance (e.g., memory cost, frame rate).

[0048] A set of parameters for 3D scene objects is described below with reference to Fig. 3. The scene depicted in Fig. 3 has two reflective / refractive objects: a car 320 and a refractive (glass) panel 310. Thus, the size of the array of N integer values ​​is equal to 2. The first element of the array corresponds to the array value of the car 320. The second element of the array corresponds to the array value of the refractive (glass) panel 310.

[0049] Car 320 is assigned a rendering quality level with integer value 2. Car 320 has high quality rendering due to internal reflections. Car 320 has environmental importance value 1 because it has color contribution from reflected rays. The array of N integer values ​​for car 320 is [-1,1]. The first element of the array is set to -1 because the car cannot contribute color to itself. Setting this value to -1 makes it irrelevant. The second element of the array indicates that a single intermediate (middle field) image is generated to collect rays from car 320 to refractive (glass) panel 310.

[0050] Bird 330 is assigned a rendering quality level having an integer value of 3. Bird 330 has a complex shape whose rendering quality may be altered during scene rendering. Bird 330 has an environment importance value of 0 since there is no color contribution from the environment. The array of N integer values ​​for Bird 330 is [2,1]. The first element of the array indicates that two intermediate (middle field) images are used to collect light rays from the bird to the reflective / glossy car 320. The second element of the array indicates that a single intermediate (middle field) image is generated to collect light rays from the bird 330 to the refractive (glass) panel 310.

[0051] The refractive (glass) panel 310 is assigned a rendering quality level with an integer value of 2. The refractive (glass) panel 310 has high quality rendering due to internal reflections. The refractive (glass) panel 310 has an environmental importance value of 2 because there is color contribution from both reflected and refracted rays. The array of N integer values ​​for the refractive (glass) panel 310 is [1,-1]. The first element of the array indicates that a single intermediate (middle field) image is generated to collect rays from the refractive (glass) panel 310 to the car 320. The second element of the array is set to -1 because the refractive (glass) panel 310 cannot contribute color to itself. Setting this array element to -1 makes this value irrelevant.

[0052] The tree 340 is assigned a rendering quality level having an integer value of 1. The tree 340 is a distant object and therefore has a low quality rendering. The tree 340 has an environment importance value of 0 as there is no color contribution from the environment. The array of N integer values ​​for the tree 340 is [2,3]. The first element of the array indicates that two intermediate (middle field) images are used to collect the light rays from the tree 340 to the reflective / glossy car 320. The second element of the array indicates that three intermediate (middle field) images are generated to collect the light rays from the tree 340 to the reflective (glass) panel 310.

[0053] Referring to step 220 of Figure 2, for one particular embodiment, an intermediate structure is determined for the selected objects in the 3D scene and for each selected object's corresponding substitute (proxy) object based on the identified set of parameters from step 210. Step 220 is applicable to objects having a rendering quality level of 1 or 3 and having a non-zero environment importance. Such objects will have color contributions from the environment (e.g., objects in the surrounding scene), as described below with respect to the flowcharts of Figures 8-9.

[0054] In an exemplary embodiment, a corresponding substitute (proxy) object is used to represent the selected object. The purpose of this step is to generate middle field images and proxies that represent the low quality rendering object for all potential incident light directions. Each generated middle field image has a unique identifier (ID) and includes color and depth information. As a non-limiting example, the unique identifier (ID) can be stored as a vector of four floating point components (red, green and blue color information, and depth information values).

[0055] For objects with rendering quality levels of 1 or 3, the corresponding (proxy) objects preferably have a primitive shape, e.g. a bounding box structure, that defines an area for objects other than lights and cameras in the 3D scene.

[0056] One technique known as an axis-aligned bounding box (AABB) may be used to define primitive shapes, which has the advantage that it only requires coordinate comparisons, thus quickly eliminating coordinates that are far apart.

[0057] The AABB for a given set of points (S) is usually the smallest area subject to the constraint that the edges of the box are parallel to the coordinate (Cartesian) axes. It is a Cartesian product of n intervals, each defined by the minimum and maximum of the corresponding coordinates for the points in S.

[0058] The selected object for which the intermediate structure is determined is the source object. The source object has a source proxy (corresponding stand-in) object. For the selected object, the objects in the surrounding scene are target objects. Each target object also has a target proxy (corresponding stand-in) object.

[0059] At the end of step 220, each object having a low rendering quality level stores information such as an array of middle field image unique identifiers (IDs), an array of associated camera projection matrices used to render these middle field images, and proxy (corresponding substitute) geometry data (e.g., center position, x, y and z extension values ​​for an axis-aligned bounding box (AABB)).

[0060] Referring to step 230 of Figure 2, a hybrid acceleration structure is determined based on the intermediate structure and the classical acceleration structure determined in step 220. The hybrid acceleration structure is generated according to the flowchart of Figure 4, as described below. The classical acceleration structure is added to support the intermediate structure.

[0061] 5 is a diagram of an exemplary classical bounding volume hierarchy (BVH) acceleration structure including red 505, green 510, and blue 515 objects with triangular rendering primitives for a 3D scene and their corresponding binary tree 520. This classical acceleration structure sorts primitives, such as triangles, into space-partitioned nodes and uses simple shapes as boundaries to represent such nodes. This arrangement allows early ray miss detection instead of testing the intersection of each primitive. Such an approach generally speeds up ray-primitive intersection testing.

[0062] Referring to FIG. 5, assume that three objects 505, 510 and 515 are separated as follows: Blue object 515: Low quality rendering (i.e., rendering quality level 1), Red object 505: high quality rendering (i.e., rendering quality level 2); Green object 510: both low and high quality rendering (i.e., rendering quality level 3).

[0063] In step 405 of FIG. 4, a determination is made as to whether a leaf node exists. If the leaf node contains one or more low-quality rendering objects (step 415), a unique identifier (ID), a projection matrix and proxy data of the object middle field image are assigned to the leaf node (step 420). If the object is a low-quality rendering object (steps 405 and 410), this node is set as a leaf node (step 425) and a unique identifier (ID), a projection matrix and proxy data of the object middle field image are assigned to such leaf node (step 430). If the leaf node is a high-quality rendering object (step 415) or if the object is a high-quality rendering object (step 410), a classical accelerated node construction is performed (step 435).

[0064] 6 is a diagram of an exemplary hybrid acceleration structure in which only a blue object 615 is classified as a low-quality rendering object for a frame of a 3D scene. Red and green objects 605, 610 are classified as high-quality rendering objects for this frame of the 3D scene. The corresponding binary tree 620 of the resulting hybrid acceleration structure is also shown.

[0065] 7 is a diagram of another exemplary hybrid acceleration structure in which both a blue object 715 and a green object 710 are classified as low-quality rendering objects for a frame of a 3D scene. In this exemplary hybrid acceleration structure, only a red object 705 is classified as a high-quality rendering object for this frame of the 3D scene. A corresponding binary tree 720 of the resulting hybrid acceleration structure is also shown.

[0066] 2, a color contribution for each selected object is determined based on the hybrid acceleration structure configured in step 230. In step 250, the 3D scene is rendered based on the hybrid acceleration structure for the selected objects of the scene. The rendering of the 3D scene is performed by camera 350 (FIG. 3) using ray tracing.

[0067] The available ray types commonly used for ray tracing are: -Camera ray: The first ray coming out of the camera. - Secondary rays: (secondary rays) generated upon interaction with a material (e.g. reflection, refraction).

[0068] An exemplary mathematical definition of a ray set is as follows:

[0069]

number

[0070] The larger the lobe(W), the more smeared / anti-aliased the ray group will be if the ray is a camera ray. The larger the lobe(W), the more coarse the ray group will be if the ray is a secondary ray (reflection / refraction) and the more soft the ray group will be if the ray is a shadow ray. The minimum lobe size is equal to divergence 0. In such case, n=1 for middlefield images, since one ray is enough to sample the domain (camera ray, reflection ray, refraction ray, shadow ray).

[0071] For the camera chief ray, the rendering of the camera chief ray is performed by activating a middle field structure attached to the camera. The generated image corresponds to a rasterized image of the scene seen by the camera. The content of each image pixel is a vector of three floating-point components that store the color information (e.g., red, green, and blue color values) of the nearest hit object.

[0072] For secondary rays, rendering is performed for each visible object from the current camera. One exemplary embodiment for determining this visibility is to check if the bounding box of the object intersects with the camera frustum (i.e., for example, using a frustum culling technique). First, the middle field structure(s) attached to the object(s) are activated. Then, color information is retrieved by an appropriate lookup in the generated image.

[0073] When traversing each node of the hybrid acceleration structure (FIG. 5), steps 825-840 are added to the classical acceleration structure described in the flowchart shown in FIG. 8. Steps 825-840 are added to support the middle field acceleration structure of the hybrid acceleration structure. Referring to the flowchart of FIG. 8, in step 810, a ray-node intersection check is performed.

[0074] If there is no ray-node intersection in step 815 of Figure 8, then a determination is made as to whether a leaf node exists in step 820. If a leaf node exists, the flow chart proceeds to steps 825-840 to add midfield structure information to the classical acceleration structure. When a leaf node contains one or more low quality rendering objects (step 825), a check for midfield-ray intersections (for both the proxy and middlefield images) is performed (step 830).

[0075] Step 830 of Fig. 8 is further described below with reference to Fig. 9. For each low quality rendering object (step 905), object proxy geometry data is obtained (step 910). If the traversed leaf node contains one or more low quality rendering objects, a ray intersection check with the middle field representation (proxy image and middle field image) is performed (step 915). With reference to steps 920 and 925 of Fig. 9, the ray intersection check with the middle field representation (step 915) is performed using the object middle field image unique identifier (ID), associated camera projection matrix and proxy geometry data (step 430 of Fig. 4) stored in the corresponding leaf node(s).

[0076] 9, a determination is made as to whether the ray-Middlefield intersection point is valid. The criteria for determining a valid intersection point are as follows: - The intersection point lies inside the frustum of the virtual camera that corresponds to the current middle-field image. -The geometry hits inside the proxy (i.e. the read depth value must not be equal to the virtual camera's far plane distance). -The proxy intersection depth must be lower than the read geometry depth to avoid visual ghosting artifacts.

[0077] 10, a non-limiting example embodiment of analyzing a ray-Middlefield intersection point to determine whether it is valid is shown. A ray 1005 hits a target proxy at intersection point 1010, allowing for a valid target geometry depth reading 1015. However, the ray 1005 does not intersect with the target geometry (the bird). In addition, the target proxy intersection point 1010 is deeper than the read geometry depth 1015, so intersection point 1010 is not a valid intersection point.

[0078] Similar to steps 855-860 in FIG. 8, in steps 935-940 in FIG. 9, color contributions are taken for valid intersections. All color contributions are accumulated per ray energy. Accumulation of color contributions of all virtual cameras for hit objects requires calculation of blending weights. One example of accumulating color contributions uses the local cubemap blending weight calculation proposed by S. Lagarde in "Image-based Lighting Approaches and Parallax-corrected Cubemap", Siggraph 2012, Los Angeles, CA, USA, August 2012.

[0079] Figure 11 is a diagram depicting a ray traversal of a portion of a hybrid acceleration structure for the 3D scene depicted in Figure 6. In Figure 6, only the blue object 615 is classified as a low quality rendering object for this frame. Referring to Figure 11, a ray 1105 traverses the A 1110, C 1115, and G 1120 leaf nodes. At the G 1120 leaf node, the hit point 1125 of the low quality rendering blue object is closer to the hit point 1130 of the high rendering quality green object than the hit point 1135. Therefore, the hit point 1125 is considered the final intersection point.

[0080] Figure 12 is a diagram depicting ray traversals of another portion of the hybrid acceleration structure for the 3D scene depicted in Figure 6. Referring to Figure 12, a ray 1205 traverses the A 1210, B 1215, and D 1220 leaf nodes. At the D 1220 leaf node, a hit point 1225 of a red object with high rendering quality is closer than a hit point 1230 of a blue object with low rendering quality. Therefore, the hit point 1225 is considered the final intersection point.

[0081] Although the present embodiments have been described above with reference to specific embodiments, the present disclosure is not limited to the specific embodiments, and modifications that fall within the scope of the claims will be apparent to those skilled in the art.

[0082] Many further modifications and variations will be suggested to those skilled in the art upon reference to the specific embodiments described above, which are given by way of example only and are not intended to limit the scope of the present disclosure, which is determined solely by the appended claims. In particular, different features from different embodiments can be interchanged as necessary.

Claims

1. identifying a set of parameters for one or more objects within a 3D scene of a video; determining, based on the identified set of parameters, an intermediate structure and a corresponding alternative object for each of the one or more objects within the 3D scene, wherein the alternative object is a middle field image or a proxy object; determining a hybrid acceleration structure based on the determined intermediate structure and a classical acceleration structure, wherein the hybrid acceleration structure indicates whether color contribution for a reflective object of the 3D scene is ray-traced or derived from an alternative object; determining a color contribution for the determined hybrid acceleration structure; providing for rendering the 3D scene of the video based on the determined color contribution of the hybrid acceleration structure. A method comprising the above steps.

2. The method according to claim 1, wherein the set of parameters includes at least one of a rendering quality level, an environmental importance value, and information regarding the reflective and refractive objects within the 3D scene.

3. The method according to claim 1, wherein the intermediate structure includes an identifier including color and depth information.

4. The method according to claim 1, wherein the intermediate structure is determined using an array camera projection matrix.

5. The method according to claim 1, wherein the corresponding alternative object includes geometry data.

6. The method according to claim 5, wherein the geometry data is at least one of a center position for a primitive shape and x, y, and z extension values.

7. The method according to claim 1, wherein the classical acceleration structure is one of a grid structure, a bounding volume hierarchy (BVH) structure, a k-dimensional tree structure, and a binary space partitioning data structure.

8. A device for rendering a 3D scene of a video, comprising: identifying a set of parameters for one or more objects within the 3D scene of the video; Based on the set of the identified parameters, determining, for each of the one or more objects in the 3D scene, an intermediate structure and a corresponding alternative object, wherein the alternative object is a middle field image or a proxy object; Based on the determined intermediate structure and the classical acceleration structure, determining a hybrid acceleration structure, wherein the hybrid acceleration structure indicates whether the color contribution for the reflective objects in the 3D scene is ray-traced or derived from the alternative object; Determining the color contribution for the determined hybrid acceleration structure; Providing to render the 3D scene of the video based on the determined color contribution of the hybrid acceleration structure, a device comprising at least one processor configured to perform the above.

9. The device according to claim 8, wherein the set of parameters includes at least one of a rendering quality level, an environmental importance value, and information regarding the reflective objects and refractive objects in the 3D scene.

10. The device according to claim 8 or 9, wherein the intermediate structure includes an identifier including color and depth information.

11. The device according to claim 8 or 9, wherein the intermediate structure is determined using an array camera projection matrix.

12. The device according to claim 8 or 9, wherein the corresponding alternative object includes geometry data.

13. The device according to claim 12, wherein the geometry data is at least one of a center position for a primitive shape and x, y, and z extension values.

14. The device according to claim 8 or 9, wherein the classical acceleration structure is one of a grid structure, a bounding volume hierarchy (BVH) structure, a k-dimensional tree structure, and a binary space partitioning data structure.