Processing media assets for rendering objects
By encoding textures as video frames in a single file and using adaptive streaming, the method addresses inefficiencies in handling large 3D rendering assets, enhancing bandwidth efficiency and reducing HTTP requests for complex scenes.
Patent Information
- Application Number
- PCT/EP2024/088662
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-31
- Filing Date
- 2024-12-30
- Publication Date
- 2025-07-03
AI Technical Summary
Current 3D rendering technologies face inefficiencies in handling large asset sizes for realistic scenes, particularly with physically-based rendering (PBR), due to the need for multiple HTTP requests and inefficient bandwidth usage of asset bundles, which can exceed 64GB in size.
The use of video-based asset packs, where multiple textures are encoded as video frames in a single file, optimized for compression and streamed using adaptive protocols like DASH, allowing efficient retrieval and decoding on GPUs.
This approach reduces the number of HTTP requests and optimizes bandwidth usage, enabling efficient rendering of large and complex 3D scenes with realistic textures, particularly beneficial for PBR methods.
Smart Images

Figure EP2024088662_03072025_PF_FP_ABST
Abstract
Description
[0001] Processing media assets for rendering objects
[0002] Technical field
[0003] The embodiments relate to processing media assets for rendering objects, and, in particular, though not exclusively, to methods and systems for processing media assets for rendering objects, a video data structure and a computer program product for executing such methods.
[0004] Background
[0005] State of the art 3D rendering engines support real-time physically-based rendering (PBR) methods, where a plurality of textures may be used in the rendering of 3D objects to simulate more advanced light interactions such as scattering, colors, roughness, reflectance and transmittance. The use of PBR methods result in much more realistic display of the (micro)surface of 3D objects without requiring artists to manually describe all the object’s details using geometry. PBR methods have not been standardized across engines, however, gITF specifies and supports the use of PBR materials. However, when rendering realistic scenes using techniques as real-time physically-based rendering, the assets required to define the objects, materials, images and textures of a scene may grow to very large sizes (e.g., recent 3D computer games may exceed 64GB in total size).
[0006] Currently MPEG SD allows the use of textures but is limited in the sense that requests for textures are made individually, since gITF 2.0 only supports the use of simple single texture images, resulting in multiple HTTP GET requests. MPEG SD allows the use of asset bundles as known for Unity. An asset bundle is a file containing only images, which can be downloaded. The use of asset bundles however is not bandwidth efficient. Moreover, both texture requests and asset bundles require additional copying from RAM to the GPU before rendering.
[0007] Hence, from the above it follows there is a need in the art for improved methods and systems processing media assets, such as textures, for rendering objects
[0008] Summary of the invention
[0009] As will be appreciated by one skilled in the art, aspects of the present invention may be embodied as a system, method or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects that may all generally be referred to herein as a "circuit," "module" or "system." Functions described in this disclosure may be implemented as an algorithm executed by a microprocessor of a computer. Furthermore, aspects of the present invention may take the form of a computer program product embodied in one or more computer readable medium(s) having computer readable program code embodied, e.g., stored, thereon.
[0010] Any combination of one or more computer readable medium(s) may be utilized. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0011] A computer readable signal medium may include a propagated data signal with computer readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium may be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0012] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber, cable, RF, etc., or any suitable combination of the foregoing. Computer program code for carrying out operations for aspects of the present invention may be written in any combination of one or more programming languages, including an object-oriented programming language such as Java(TM), Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0013] Aspects of the present invention are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor, in particular a microprocessor or central processing unit (CPU), of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer, other programmable data processing apparatus, or other devices create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0014] These computer program instructions may also be stored in a computer readable medium that can direct a computer, other programmable data processing apparatus, or other devices to function in a particular manner, such that the instructions stored in the computer readable medium produce an article of manufacture including instructions which implement the function / act specified in the flowchart and / or block diagram block or blocks.
[0015] The computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. Additionally, the Instructions may be executed by any type of processors, including but not limited to one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FP- GAs), or other equivalent integrated or discrete logic circuitry.
[0016] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions. In an aspect, the embodiments may relate to a method of processing media assets wherein the method may comprise a step of receiving a scene description file comprising information about at least one 3D object and one or more source locations for obtaining media assets for rendering the 3D object, the assets including mesh data associated with the 3D object and a video file comprising a sequence of video frames, each video frame comprising encoded texture data for adding a texture to a surface of the 3D object. The method may also comprise a step of retrieving the mesh data and the video frames based on the information in the scene description. The method may also comprise the step of rendering by a rendering device the 3D object on a display based on the mesh data and the video frames, wherein the rending of the 3D object may include allocating one or more buffers, each buffer being configured to store a decoded video frame and each of the one or more allocated buffers being accessible by the rendering device; decoding the video frames into decoded video frames; and, storing at least part of the decoded video frames in the one or more allocated buffers. The method may further comprise adding a plurality of textures to the surface of the 3D object on the basis of the texture data from the decoded video frames.
[0017] Hence, the embodiments suggest using a video file comprising video frames comprising encoded texture data for adding a texture to a surface of the 3D object in the rendering process of a 3D object. Video coding is used for efficient compression of media assets such as textures that are used for rendering 3D objects in a scene based on a scene description. This approach is especially advantageous when using a large amount of texture images, for example in the case of physically-based rendering (PBR) methods. Further, adaptive streaming protocols such as DASH may be used for efficient retrieval of videobased asset packs for bandwidth-constrained client devices. The use of a single video file for retrieving assets allows optimizing the total amount of requests, e.g. HTTP requests, that are required to retrieve the media assets that are needed to render a scene. Existing HTTP adaptive streaming (HAS) CDN architecture may be used to efficiently distribute video-based asset packs to client devices. Further, in some embodiments, the GPU may be used for video decoding video so that in certain scenario’s the decoded output will be reused by the GPU.
[0018] In an embodiment, the scene description file may comprise information about the one or more buffers that should be allocated for rendering the 3D object and information about which decoded video frame from the decoded video frames should be stored in each of the one or more allocated buffers.
[0019] In an embodiment, the one or more buffers that should be allocated may be identified by a buffer identifier or a buffer index.
[0020] In an embodiment, the scene description may identify each video frame in the video file. In an embodiment, each the video frame may be identified by a frame identifier or a frame index.
[0021] In an embodiment, the presentation engine (PE) may be configured to control the rendering of the 3D object based on the information in the scene description file.
[0022] In an embodiment, the presentation engine may be further configured to instruct a media access function (MAF) to execute at least one of: the retrieval of the video frames, the allocation of the one or more buffers, the decoding of the video frames and the storage of at least part of the decoded video frames in the one or more allocated buffers.
[0023] In an embodiment, the instruction of the MAF to the PE may include one of more of the following information items: information about one or more source locations for obtaining media assets for rendering the 3D object; information about one or more buffer identifiers or buffer indices for identifying the one or more buffers that should be allocated for rendering the 3D object; and / or, information about which decoded video frame from the decoded video frames should be stored in each of the one or more allocated buffers.
[0024] In an embodiment, the rendering may further include: informing the rendering device that each of the one or more allocated buffers comprises a decoded video frame that comprises texture data.
[0025] In an embodiment, the texture data of each video frame in the sequence of video frames may be configured to add a different texture to the surface of the 3D object.
[0026] In an embodiment, the video frames may be retrieved based on a streaming protocol, preferably an HTTP adaptive streaming protocol, such as MPEG-DASH.
[0027] In an embodiment, the one or more buffers may be GPU-based buffers.
[0028] In a further embodiment, the embodiments in this disclosure may relate to a apparatus, preferably a client device, for processing media assets comprising: a computer readable storage medium having at least part of a program embodied therewith; and, a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, wherein the processor may be configured to perform one or more of the following executable operations: receiving a scene description file comprising information about at least one 3D object and one or more source locations for obtaining media assets for rendering the 3D object, the assets including mesh data associated with the 3D object and a video file comprising a sequence of video frames, each video frame comprising encoded texture data for adding a texture to a surface of the 3D object; retrieving the mesh data and the video frames based on the information in the scene description; and, rendering, by a rendering device, the 3D object on a display based on the mesh data and the video frames, the rending of the 3D object including: allocating one or more buffers, each buffer being configured to store a decoded video frame and each of the one or more allocated buffers being accessible by the rendering device; decoding the video frames into decoded video frames; and, storing at least part of the decoded video frames in the one or more allocated buffers.
[0029] In further embodiments, the processor of the apparatus maybe further configured to perform any of the method steps as described with reference to the above described embodiments.
[0030] The invention may also relate to a computer program product comprising software code portions configured for, when run in the memory of a computer, executing the method steps according to any of process steps described above.
[0031] The invention will be further illustrated with reference to the attached drawings, which schematically will show embodiments according to the invention. It will be understood that the invention is not in any way restricted to these specific embodiments.
[0032] Brief description of the drawings
[0033] Fig. 1A-1B illustrates a mesh representing a 3D object and textures which can be mapped onto the mesh;
[0034] Fig. 2 illustrate an example of a physically realistic rendered 3D object;
[0035] Fig. 3 depicts an overview of the hierarchical model of gITF entities and their relationships;
[0036] Fig. 4 depicts a media system for rendering assets based on a scene description;
[0037] Fig. 5 illustrates pipelines of a media system media system for rendering assets based on a scene description;
[0038] Fig. 6 illustrates a logical architecture of a rendering system that may be used by the embodiments in this disclosure;
[0039] Fig. 7A and 7B depict an example of a video-based asset pack and the processing of such video-based according to an embodiment;
[0040] Fig. 8A and 8B depict a method of rendering part of a scene based on a video-based asset pack according to an embodiment;
[0041] Fig. 9 depicts a flow diagram of a method of processing a video-based asset pack for rendering assets according to an embodiment; Fig. 10 depicts a schematic of an asset rendering system according to an embodiment;
[0042] Fig. 11 depicts a block diagram illustrating an exemplary data processing system that may be used with embodiments described in this disclosure.
[0043] Description of the embodiments
[0044] 3D assets are digital files that represent objects or elements in a three- dimensional space. These assets consist of data that defines the shape, texture, and appearance of these objects, allowing them to be rendered and animated in various software applications. One of the most common types of 3D assets are 3D models. These are digital representations of physical objects such as characters, vehicles, buildings, or props. 3D modeling is the process of developing a mathematical coordinate-based representation of a surface of an object (inanimate or living) in three dimensions.
[0045] A 3D model may represent the geometry of an object in the form of a mesh. Typically, a mesh may define collection of points in 3D space wherein the points define vertices which are connected by various geometric entities such as triangles, lines, curved surfaces, etc. to form a surface of a 3D object. 3D models can be created manually, algorithmically (procedural modeling), or by scanning. The surface of a 3D mesh may be further defined with texture mapping. A scene may be created using a plurality of 3D models. A 3D model can be displayed as a two-dimensional image through a process called 3D rendering.
[0046] Real-time rendering algorithms are typically structured using pipelines. Commonly used rendering pipelines are defined by platforms like DirectX and OpenGL. Graphical Processing Units (GPUs) have been designed and evolved with and as part of these rendering pipelines, and offer processing capacity for highly parallelized algorithms. A rendering pipeline can usually be divided into different stages, including an application stage, wherein the definition of the 3D environment (or scene) may be stored and sections of the scene are selected, pre-processed, and sent to the following stages, namely the geometry stage, at which where the relevant parts of the 3D objects in the scene are projected onto a 2D plane and the rasterization stage, wherein the color of each mapped object is determined.
[0047] As described a mesh is a particular type of 3D object, which is commonly used to describe a geometry for rendering pipelines. An example of a mesh representing an 3D object, in this example a torus, is provided in Fig. 1 A. Meshes consist of a set of 3D points or vertices, which may be connected using edges. Gaps between a group of vertices may be filled to form a ‘face’, wherein faces are typically defined using groups of three (triangles) or four (quadrilaterals) vertices. 3D meshes only describe the geometrical shape of an object. To introduce color and / or texture to 3D objects, one or more texture images (in gITF referred to as images) comprising texture data may be mapped onto the object using so-called UV maps, wherein each vertex of the mesh is mapped to a 2D point (u,v) on said image(s). Fig. 1B depicts a number of texture images 104I-6, which can be mapped onto a 3D object. In order to provide a 3D object a realistic appearance multiple texture images associated with different aspects (structure, lighting, color, etc..) may be required for rendering a 3D object. Fig. 2 depicts a rendered 3D torus which has a surface that has the appearance of bricks, wherein effects like shading and microstructure of the bricks are also include. This 3D object may be rendered based on the mesh and some or all texture images as illustrated in Fig. 1A and 1B.
[0048] Multiple 3D objects can be used to create 3D scenes for rendering. In a digital description of such scenes, additional data may be included such as: the position and orientation of each object, the configuration of the (virtual) camera, the position, configuration and type of lights, and references to external files and media. Typically, scene descriptions are based on a hierarchical model, wherein entities may be defined with respect to each other. Such hierarchical model simplifies the specification of complex and / or detailed relationships such as the location and relationships of bones in a skeleton. Different formats for scene descriptions include gITF, X3D and MPEG SD.
[0049] The gITF format is an interoperable format for the exchange of 3D and scene data developed by the Khronos Group. gITF documents are serialized using either the JSON format or a binary gITF-specific format (.gib). Fig. 3 depicts an overview of the hierarchical model of gITF entities and their relationships. As can be seen from this model, scenes may be defined as arrays of nodes, where the nodes may have sets of children. Vertex and texture data can be referred to using buffers and bufferviews respectively. Sets of textures may be mapped onto meshes using materials.
[0050] The ISO X3D standard describes an XML-based format for the specification of 3D scenes, which is maintained by the Web3D Consortium. Main implementations of the X3D standard are X3DOM and XJTE. In X3D, scenes are described using a hierarchical node structure (powered by XML), that allows for the definition of assets and the scene with similar features as gITF (in fact X3D supports including gITF documents).
[0051] The MPEG-I group is developing a standard for enabling description and execution of interactive media scenarios. The MPEG SD standard extends the gITF scene specification standard by adding support for media. Specifically, MPEG SD introduces an explicit decomposition between scene description, presentation and media operations. Fig. 4 depicts a media system for rendering assets based on a scene description as known from MPEG SD. The system may include a presentation engine PE 404 and a media access function MAF 402. The presentation engine may process assets, e.g. 3D objects in a 3D scene, 2D scenes and media content, e.g. video, and prepare the assets for rendering by a rendering device (not shown). The PE may process assets based on a scene description 403. The MAF and the PE may communicate with each other via a MAF application programming interface (MAF API) 414. This way, the PA may instruct the MAF to retrieve and prepare assets, e.g. 3D objects and / or media content, that are needed for rendering a scene as defined in the scene description. To that end, the scene description may include information about locations, e.g. resource locators such as URLs or URIs, where assets that are needed from a scene can be retrieved. This information may be provided by the PA to the MAF in one or more instructions, by the PA to the MAF (e.g. one or more MAF API calls). The MAF may retrieve the assets from a local storage 420 via a media access connection 422 or from the cloud or a server network 416 via one or more media requests. The PE may be responsible for rendering the scene provided by the scene description document, wherein the PE may delegate retrieval, parsing and decoding of media to the MAF. The MAF may then provide the required media in the requested format using buffers as an interface to the PE. Typically, the buffer is a circular buffer. The buffers may be managed by a buffer management module 410 wherein a buffer API 412 provides an interface for the MAF and the PE to the buffer management module. If assets, e.g. media objects, are received by the MAF, it may initiate and allocate one media pipeline 406i.nfor each media object, wherein each media pipeline is associated with a single buffer, which can be accessed by the PE.
[0052] The media pipelines and the buffers allow decoupling of the rendering by the PE and the media retrieval by the MAF. The PE can use information in the scene description to instruct the MAF to retrieve assets of a scene, e.g. media objects and to initiate pipelines with associated buffers to process media objects so that the PE can retrieve each processed objects via the buffer.
[0053] Fig. 5 illustrates the pipelines of the media system as depicted in Fig. 4 in somewhat greater detail. Each of the pipelines 502i-nmay be associated with encoded media data representing a media object which may be stored as a track 506 as known from the ISO Base Media File format. Encoded media data may be formatted in tracks and each track may include media data associated with a specific asset, e.g. video data, point cloud, a mesh (e.g. vertex positions), a texture image, etc. The MAF may retrieve media data of a track and provide the data to a decoder 508 for decoding the media data into decoded media data, which can be subsequently processed, e.g. formatted, by a media processing unit 510 before stored into a buffer 512.
[0054] Fig. 6 illustrates a logical architecture of a rendering system 600 that may be used by the embodiments in this disclosure. The system may include application 602 running on a processor, which may be configured to send logical instructions to one or more Graphical Processing Units (GPUs) 616 for rendering images onto a display 620, e.g. a display of a head-mounted device or the like. A GPU may comprise one or more frame buffers 618 which are used to collect and store a digital representation of the next image to be displayed. The application may be configured to display an interactive scene (either 2D or 3D) which is defined in a scene description 608 as described with reference to the embodiments in the application. One or more rendering pipeline abstractions 614 may be used to specify how to render the scene according to the application-specific view. In this way, the application does not need worry about sending instructions to the GPU, but defers this responsibility to the rendering (pipeline) library. The application may further include a Presentation Engine (PE) 604 and Media Access Function (MAF) 606 as described with reference to Fig. 4. The PE is responsible for setting up the presentation of the scene as indicated by a scene description document 608. The PE achieves this by instructing the MAF to retrieve and prepare assets (e.g. 3D meshes 612 and images 610 comprising specific assets information, e.g. texture, color, materials, etc.) as specified by the scene description.
[0055] It is noted that Fig. 6 is a non-limiting example of a logical architecture of a rendering system. Many variations are possible. For example, the frame buffer may also be used to write immediately to the display. Further, the application (or one or more parts thereof) may be implemented as software running on a Central Processing Unit (CPU), however it is also possible that the application (or one or more parts thereof) may be executed on a different system (e.g., the GPU itself or a System-on-a-chip (SoC)). Further, some parts of the application may be implemented in hardware. In further embodiment, the role of the GPU may be fulfilled by the CPU or other system, e.g., using software-based rendering or as an embedded specialized chip on the CPU. In another embodiment, the rendering pipeline may be part of the application itself.
[0056] Further, due to implementation details chosen by MPEG SD, the division between which assets are retrieved by the MAF or PE is not strict. The PE may retrieve assets specified in gITF, and the MAF may retrieve assets as specified by MPEG so there will be overlap in the assets that MAF and PE may retrieve.
[0057] Recent state of the art 3D rendering engines support real-time physically- based rendering (PBR) methods, where a plurality of textures may be used to simulate more advanced light interactions such as scattering, colors, roughness, reflectance and transmittance. An example of such texture images and associated rendered objects is shown in Fig. 1 and 2. The use of PBR methods result in much more realistic display of the (micro)surface of 3D objects without requiring artists to manually describe all the object’s details using geometry. PBR methods have not been standardized across engines, however, gITF specifies and supports the use of BRDF-, BSDF- and BTDF-based PBR materials. For example, rendering realistic scenes using techniques as real-time physically-based rendering, the assets required to define the objects, materials, images and textures of a scene may grow to very large sizes (e.g., recent 3D computer games may exceed 64GB in total size).
[0058] Currently MPEG SD allows the use of textures but is limited in the sense that requests for textures are made individually, since gITF 2.0 only supports the use of simple single texture images in jpg or png format, resulting in multiple HTTP GET requests. MPEG SD allows the use of asset bundles as known for Unity. An asset bundle is a file containing only images, which can be downloaded. The use of asset bundles however is not bandwidth efficient. Moreover, both texture requests and asset bundles require additional copying from RAM to the GPU before rendering.
[0059] The embodiments in this disclosure address this problem by defining a so- called video-based asset pack wherein multiple textures associated with 3D objects are encoded video frames of video file, which can be streamed to the MAF.
[0060] Fig. 7A and 7B depict an example of a video-based asset pack and the processing of such video-based according to an embodiment. As shown in Fig. 7A an asset pack may be created by encoding (compatible) asset, e.g. a number of texture images comprising texture data for one or more 3D objects, into a sequence of video frames and packing the encoded assets into a video file. As many similarities between available texture images may exist, encoding these images into a video file may allow substantial compression of the data, resulting in a video file of substantial smaller size compared to files that are conventionally used to handle texture images. The video file representing the asset pack may be stored as a track which can be retrieved based on a known streaming protocol such as DASH.
[0061] It is noted that that the timeline of the sequence of video frames of an asset pack 702 is not the timeline relating to the timeline of the rendering of the scene. When defining a scene description, the video frames packed within the video file need to be addressable. For example, in the case of textures, such addressing is needed to specify which frame(s) in the sequence of video frames in the video file need(s) to be used for a specific texture of a 3D object. Hence, the location of a particular video frame representing a specific texture in the sequence of video frames of the asset pack may be used to address this particular video frame. The MAF may request a video-based asset pack and use the location of the video frames in the video-based asset pack to store the frames into separate (memory) buffers for each defined asset in the scene so that the buffers can be accessed by the presentation engine.
[0062] Fig. 7B depicts the processing of a video-based asset pack according to an embodiment. The MAF may retrieve a video-based asset pack 706, a video file, based on instructions of the PE. The video-based asset pack may comprise assets, e.g. textures, associated with one or more 3D objects in a scene as defined in a scene description. The video file may be requested based on a streaming protocol such as DASH. The MAF may initiate and allocate at least one media pipeline for processing the sequence of frames in the video file and initiate and allocate one or more buffers 714i -4 for different defined assets, e.g. 3D objects, in the scene. For example, as shown in the figure, buffer 714i is configured to store a video frame representing a texture that is linked to a first mesh 7161 of a first 3D object in the scene so that the buffers can be accessed by the presentation engine. In a similar way, buffer 7142 is configured to store a video frame representing a texture that is linked to a second mesh 7161 of a second 3D object in the scene. This way, video frames comprising texture data stored in different buffers can be linked to different 3D objects (different meshes) in the scene.
[0063] In an embodiment, the PE may initiate and allocate the media pipelines and / or buffers. In yet another embodiment, the MAF may initiate and allocate a first part of the pipelines and / or buffers and the PE may initiate and allocate a second part of the pipelines and / or buffers. In an embodiment, the MAF may initiate and allocate the media pipelines and the buffers (or a part thereof) based on information in the instructions of the PE.
[0064] Once the buffers are set up and allocated, the sequence of video frames may be extracted from the video file and presented to the input of the decoder, which may decode the encoded video frames into a sequence of decoded video frames 710 (a stream) wherein each video frame may comprise different types of texture data. A fan-out buffer may be controlled to distribute the video frames over the different buffers based on their position in the sequence and based on a certain distribution scheme. For example, this distribution scheme may include copying the first video frame 0 and storing video frame 0 in the first buffer 714i and the second buffer 7142. Further, it may store video frame 1 into the third buffer 714s and video frame 3 in the fourth buffer 7144, while discarding video frame 2. Each of the buffers may be associated with an asset in a scene.
[0065] Fig. 8A and 8B depict a method of rendering part of a scene based on a video-based asset pack according to an embodiment. In particular, the figure illustrates a sequence diagram of a method of rendering a scene based on a video-based asset pack as described with reference to the embodiments in this disclosure. Here, the process is based on an implementation of the PE that operates similar to how Unity and other game engines operate.
[0066] The process may be executed by a rendering system comprising a presentation engine 810 (PE), a media access function 812 (MAF) and a render 814. As shown in the figure, the PE may receive an instruction to render a scene as described in a scene description 804. In a first step 816 the PE may load the scene description, e.g. a fltf file. The PE may parse the scene description (step 818) and initiate the Tenderer based on information in the scene description file, e.g. camera parameters defining e.g. the field of view and information about the format of the vertices of 3D mesh models that needs to be rendered. Further, the PE may determine information about a 3D model 819 that needs to be rendered and information on textures 820 that are needed for the model. The PE may also determine that two textures refer to a video source 821 , indicating that these textures should be accessed as a video based-asset pack. Based on the information in the scene description, the PE may request the MAF to allocate buffers for three textures (step 822). In other embodiments, the PE may allocate all buffers or a part of the buffers. Then, based on the scene description, the PE may load 3D model data, e.g. coordinates of vertices defining a 3D mesh, and to send the 3D model data to the Tenderer (step 824).
[0067] Further, the texture data may be loaded (step 826). To that end, the PE may instruct the MAF to retrieve the video-based asset pack and to allocate video frames in the video-based asset pack to specific buffers (step 828), e.g. frame 0 into buffer 0 and 1 and frame 1 into buffer 2. Further, the MAF may retrieve the video-based asset pack (a video file), e.g. by requesting a server to stream the video frames to the MAF (step 830).
[0068] Upon receiving the video frames, the MAF may start the decoding process of the first video frame (step 832), e.g. based a media pipeline. The decoded first video frame may be stored in the allocated buffers (step 834) buffer and the Tenderer may be informed that the buffers comprise the first video frame comprising the texture that is needed for the rendering process. In a similar way, the decoding process of the second frame may be started (step 840) which includes storing the decoded second video frame in the allocated buffer (step 842) and informing the Tenderer that the buffer comprises the second video frame. After, the decoding process and the storage of the decoded video frames in the allocated buffers, the PE may instruct the render to start the rendering process (step 846).
[0069] Hence, from the above, it follows that texture data for texturing a 3D object may be efficiently encoded in a sequence of video frames and packaging the sequence of encoded video frames into a video file of certain format. This video file may be referred to as a video-based asset pack. When constructing a scene based on 3D objects that use the video-based asset pack, the PE may request the MAF to retrieve the video-based asset pack and to store the video frames comprising the texture data into one or more separate (memory) buffers for each 3D object in the scene. The PE may then submit the one or more allocated buffers, each comprising one video frame, to the rendering engine so the scene can be rendered based on media assets that are packed as video.
[0070] A generic description of the process is provided in Fig. 9. In particular, this figure depicts flow diagram of a method of processing a video-based asset pack for rendering assets according to an embodiment. As shown in this figure, the method may include the step receiving a scene description file (step 902) wherein the scene description file may comprise information about at least one 3D object and one or more source locations for obtaining media assets for rendering the 3D object, the assets may include mesh data associated with the 3D object and a video file comprising a sequence of video frames, each video frame comprising encoded texture data for adding a texture to a surface of the 3D object. The mesh data and the video frames may be retrieved based on the information in the scene description (step 904). Then, the 3D object may be rendered by a rendering device on a display (step 908) based on the mesh data and the video frames, wherein rendering of the 3D object may include allocating one or more buffers, each buffer being configured to store a decoded video frame and each of the one or more allocated buffers being accessible by the rendering device; decoding the video frames into decoded video frames; and, storing at least part of the decoded video frames in the one or more allocated buffers.
[0071] Hence, the embodiments suggest using a video file comprising video frames comprising encoded texture data for adding a texture to a surface of the 3D object in the rendering process of a 3D object. Video coding is used for efficient compression of media assets such as textures that are used for rendering 3D objects in a scene based on a scene description. This approach is especially advantageous when using a large amount of texture images, for example in the case of physically-based rendering (PBR) methods. Further, adaptive streaming protocols such as DASH may be used for efficient retrieval of videobased asset packs for bandwidth-constrained client devices. The use of a single video file for retrieving assets allows optimizing the total amount of requests, e.g. HTTP requests, that are required to retrieve the media assets that are needed to render a scene. Existing HTTP adaptive streaming (HAS) CDN architecture may be used to efficiently distribute video-based asset packs to client devices. Further, in some embodiments, the GPU may be used for video decoding video so that in certain scenario’s the decoded output will be reused by the GPU.
[0072] To reference video-based asset packs in a scene the syntax of a scene description language may be extended. Below an example is provided wherein the syntax of the gITF “images” syntax is extended for addressing frames in a video-based texture pack:
[0073] "images" : [ {
[0074] "mimeType" : "image / mp4" ,
[0075] "name" : "asset pack frame 0" , "uri" : "asset pack . mp4" , "extensions" : {
[0076] "TNO asset packing" : { "track index" : 0 , "gop index" : 0 , "frame index" : 0
[0077] } "mimeType": "image / mp4", "name": "asset pack frame 1" , "uri": "asset pack. mp4", "extensions": {
[0078] "TNO asset packing": { "track index": 0, "gop index": 0, "frame index": 1 }
[0079] }
[0080] }, {
[0081] "mimeType": "image / mp4", "name": "asset pack frame 2", "uri": "asset pack. mp4", "extensions": {
[0082] "TNO asset packing": { "track index": 0, "gop index": 0, "frame index": 2 }
[0083] }
[0084] }]
[0085] The extension format specifies metadata that are needed as formal gITF extension, including a ‘gopjndex’ referring to the index of a particular group of pictures in a media stream, a ‘framejndex’ referring to the index of a particular frame within the selected GOP, and a ‘trackjndex’ referring to the index of a track in the referenced asset pack. This extension allows scene editors to specify a specific video frame within the sequence of video frames that define the asset pack as the source of an image for a certain asset, e.g. an object or the like.
[0086] Similarly, the Media Access Function needs to be adapted to retrieve and decode the asset pack. In particular, it needs an API to allow the PE to request buffers to be filled using an addressed frame. For this feature, an extension to the MAF API may be used, wherein the 'AlternativeLocation' interface (part of the MAFs initialize()' function arguments) is extended as follows:
[0087] Interface AlternativeLocation { attribute String mimeType; attribute Track tracks [] ; attribute uri;
[0088] }; interface Track { attribute String track; attribute integer id; attribute Selection selection [] ;
[0089] }; enum SelectionMode {
[0090] "NONE" , "FRAME" } interface Selection { attribute integer bufferld; attribute SelectionMode SelectionMode ; attribute integer goplndex ; attribute integer frameindex ;
[0091] }
[0092] Here, the default of a single buffer per track is replaced with an array structure that supports multiple buffers (while the use of a single buffer is still possible). Semantics for the additions are as follows:
[0093] • A track shall have at least 1 selection element.
[0094] • If the ‘SelectionMode’ is set to “NONE”, the ‘goplndex’ and ‘frameindex’ fields will be ignored, and the default behavior is expected.
[0095] • One of the buffers may have its SelectionMode field set to ‘NONE’ if and only if it is the only buffer in the list.
[0096] • If the ‘SelectionMode’ is set to “FRAME”, the ‘selection’ field will be parsed and interpreted as follows:
[0097] The MAF will decode up to the GOP indicated by the ‘goplndex’ field.
[0098] The MAF will then start decoding the indicated GOP up to the frame indicated by the ‘frameindex’ field.
[0099] The MAF will then decode the frame indicated by the ‘frameindex’ field, and place the resulting image into the buffer with the id specified by the ‘bufferld’ field.
[0100] In addition, the MAF may perform the following optimizations:
[0101] If possible (e.g., due to metadata present in the media container), the MAF will skip ahead until the indicated frame and / or GOP.
[0102] If the MAF detects that a particular media is requested multiple times, request it only once.
[0103] If the MAF detects that multiple selections are made from a particular media (e.g., when multiple selections are made), the MAF may spawn a single decoder and optimize which GOPs and frames to decode.
[0104] The render the asset pack the Presentation Engine needs to be adapted.
[0105] The PE will parse the gITF extension indicated above. Using this extension, for each image, when preparing the MAF ' initializeQ' request, the PE will add selections with the ‘FRAME’ selection mode to the ‘AlternativeLocation’ interface. Likewise, the PE will populate the ‘Selection’ interface using the values specified ‘gopjndex’ and ‘framejndex’ fields from the scene description and map them onto the ‘goplndex’ and ‘frameindex’ fields respectively.
[0106] MPEG SD requires in its MAF API that the PE provides a set of buffer identifiers. gITF semantics may be used for buffer references in order to generate these. The gITF specification (the Khronos® 3D Formats Working Group, "gITF™ 2.0 Specification," of 11 October 2021), specifies in section 3.6.1.1 the following method to reference buffers in a 'bufferView' declaration:
[0107] "bufferviews" : [ { "buffer" : 0 , "byteLength" : 25272 , "byteOffset" : 0 , "target" : 34963 } , ]
[0108] Here, the number ‘0’ is referencing the index of a buffer in the list of buffers of the gITF file where this bufferview (and the buffer) is defined. Using the above information, the PE will determine which buffers have been declared by parsing the gITF file. Then, for each buffer that needs to be provided to the MAF to implement the embodiments in this disclosure, the buffer index will be used as the buffer id in the call to the MAF’s initialize() method.
[0109] Fig. 10 depicts a schematic of an exemplary rendering system which is adapted to process and render video-based asset packets, i.e. assets that are formatted as a video file comprising a sequence of video frames. The system may include an asset preparation system 1002, a server system 1004, and media playout device 1006 comprising a client device 1042, which may communicate via a network 1036 (including the Internet) to the server system. In some embodiments the content preparation system may be part of the server system. In other embodiments, the content preparation system may be connected to the server system via e.g. the network 1036 or another network, or may be directly communicatively coupled. The server system 1004 may include a server processor 1032 and one or more network interfaces 1034 which are configured to send and receive data via network 1036. In some embodiment, functionalities of the server system and the asset preparation system may be implemented in the form of a distributed system including a plurality of communicatively connected network devices, including but not limited to routers, bridges, proxy devices, switches, etc. In an embodiment, the server system may be part of a CDN.
[0110] The asset preparation system 1002 may include a source of stored assets 1008 for 3D rendering, including different type of media data, e.g. 2D and 3D video data, point clouds, 3D objects including 3D meshes and textures, etc. At least part of the assets may be stored in the form of a video-based asset pack, i.e. sequence of video frames wherein the assets, e.g. textures or other type of asset data, are encoded in video frames so that these files can be compressed and streamed based on an adaptive streaming protocol.
[0111] The asset preparation system may include an encoder system 1014 comprising one or more encoder instances for encoding the media data. The encoder system may produce one or more encoded media data streams (in short media streams) representing assets, e.g. textures for one or more 3D models in a scene, that are needed for rendering a scene. In some embodiments, an individual stream may be referred to as an elementary stream representing a single, digitally coded component (e.g. video or audio) of a media representation. The asset preparation system may further include a packetizer 1016 for converting elementary streams comprising encoded media data into a packetized stream, e.g. a packetized elementary stream (PES). The PES streams may be formatted, e.g. encapsulated, by an encapsulator 1017 for transport so that that encoded media data can be transmitted to the server system using a suitable media streaming standard such as MPEG- DASH, HLS or CMAF and stored as one or more media files at a storage medium of the server system 1004. The encapsulator may be configured to generate media files, which are formatted according to a predetermined data format for example CMAF fragment or DASH segments.
[0112] Media data which are encoded in one or more elementary streams according to a certain bitrate or quality may form a media representation or in short a representation. The encoder system 1014 may be configured to encode media data of a media title in different ways using a video coding standard to produce different representations of a media title at various bitrates and various characteristics, such as pixel resolutions, frame rates, conformance to various coding standards, etc. These different representations may be used for adaptive bitrate streaming as known from streaming protocols such as DASH. In particular, the encoder system may encode the media data according to any suitable standardized coding scheme such as H.264 / AVC, HEVC or VVC and, in case of point cloud, the Geometry-based PCC (G-PCC) and Video-based PCC (V-PCC) compression standards.
[0113] The encapsulator 1016 may be configured may be configured to format packets of elementary systems into network abstraction layer (NAL) units. NAL units, which are defined as part of the H.264 / AVC and HEVC video coding standards, include Video Coding Layer (VCL) NAL units comprising video data payload and non-VCL NAL units, which may comprise metadata such as parameter sets (important header data that can apply to a large number of VCL NAL units) and supplemental enhancement information (timing information and other supplemental data that may enhance usability of the decoded video signal). Non-VCL NAL units may include sequence parameter sets (SPS), which apply to a series of consecutive coded video pictures called a coded video sequence and picture parameter sets (PPS), which apply to the decoding of one or more individual pictures within a coded video sequence. Non-VCL NAL units may further include Supplemental Enhancement Information (SEI) messages which may contain information for assisting the decoding process.
[0114] A set of NAL units may define a so-called access unit which together may form a coded picture (a video frame), This way, the decoding of an access unit generally results in one decoded picture (a decoded video frame). A coded video sequence consists of a series of access units that are sequential in the NAL unit stream and use only one sequence parameter set. Each coded video sequence can be decoded independently of any other coded video sequence, given the necessary parameter set information, which may be conveyed "in-band" or "out-of-band". At the beginning of a coded video sequence is an instantaneous decoding refresh (I DR) access unit. An I DR access unit comprises an intra picture (l-frame) which is a coded picture that can be decoded without decoding any previous pictures in the NAL unit stream. The presence of an I DR access unit indicates that no subsequent picture in the stream will require reference to pictures prior to the intra picture it contains in order to be decoded. The encapsulator may use coded video sequences to produce short non-overlapping short video files, such as DASH segments and CMAF fragments, that are used by adaptive streaming protocols to provide adaptive streaming functionality.
[0115] The encapsulator 1016 may be further configured to determine one or more manifest files 1024 and / or scene description files 1025. An example of a manifest file is a media presentation descriptor (MPD) in case MPEG-DASH is used for streaming the media data. The manifest file and / or scene description files may identify media assets 1026-1030, such as media objects, e.g. 3D objects, and textures. The manifest file and / or scene description file may comprise information formatted according to a certain syntax such as the extensible markup language (XML) or JSON. In some embodiments, media assets may be divided into so-called adaptation sets. An adaptation set may define media data associated with a common set of characteristics, including but not limited to e.g. codec, profile and level, resolution, number of views, file format for segments, etc. The manifest file and / or scene description file may include data identifying such adaptation sets and further information associated with characteristics, such as bitrates, of specific representations of adaptation sets. The packetized and encapsulated media files, e.g. CMAF fragments and / or DASH segments, prepared by the assets preparation system may be stored at the server system as tracks 1026-1030 and an associated manifest file 1024 and / or scene description 1025. These assets may include video-based asset packs as described with reference to the embodiments in this disclosure. The server processor 1032 may be configured to receive network requests from client devices, such as client device 1042. The client device may comprise a client processor 1048 configured to request media data and / or metadata, e.g. a manifest filed, that is stored at the server system. Based on a manifest file 1024 and / or scene description stored at the client device, the client device may request (via the client processor) media assets form the server system and store the media data in a client buffer 1044. The buffered (encoded) media assets may be provided to the to a rendering application 1050 comprising a media access function 1051 and a presentation engine 1052 as described with referend to Fig. 6. The MAF may retrieve media assets (upon instruction from the PE) based on the scene description. Functions of the client device, the MAF and the PE may overlap so that the distinction between the different entities is not so strict. In any case, the client device may be configured such that the MAF is capable of receiving video based asset packs using a HTTP adaptive streaming protocol such as DASH or the like.
[0116] Assets processed by the rendering application may be provided to a rendering device 1052, e.g. a GPU-based Tenderer comprising a frame buffer for a display 1056. The display device may be implemented as any type of display devices, including display devices such as a head-mounted device for rendering XR-type of media data (e.g. tiled 360 video data). Sensor information 1060 associated with the display device, e.g. viewing direction and pose information, may be used to control the rendering process and to select the assets that are needed for rendering.
[0117] The server processor 1032 and client processor 1048 may be implemented to process requests based on the hypertext transfer protocol (HTTP), for example HTTP version 1.1, which allow transmission of encoded media data based on the chunked transfer encoding mode. This way, the server request processor may be configured to receive HTTP messages, such as HTTP GET or partial GET requests and sent media data in response to the requests back to the client device. The requests may specify a video file or a part of a specific part, e.g. a fragment or a chunk of a fragment, of one of the tracks 1026-1030, e.g., using a resource locator, such as an URL. In some examples, the requests may also specify a byte range for identifying a chunk in a fragment. Instead of HTTP other client-server communication protocols may be used to handle request and response messages. For example, in an embodiment, a bi-directional communication channel between the client and the sever system may be realized based on a WebSocket protocol. In that case, a handshake request may be used to set up a WebSocket connection between the client and server. Request and response messages may be exchanged between the client and the server over the WebSocket connection. Other protocols that may be used to communicate between the server and the client include Long polling, WebRTC, SignaIR, or the like. The network interface 1040 of the client device may receive media and buffer media, e.g. encapsulated and packetized media assets such as CMAF fragments or chunks of a selected representation. The client processor 1048 may be configured to decapsulate the media files into PES streams and depacketize the PES streams into encoded media assets, e.g. a sequence of encoded video frames of a video based asset pack, which may be provided to the rendering application as described with reference to Fig. 6 .
[0118] The devices and modules depicted in Fig. 10 such as encoder, packetizer, encapsulator, server processor, client processor, etc. may be implemented as any of a variety of suitable processing circuitry, as applicable, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuitry, software, hardware, firmware or any combinations thereof. Alternatively and / or additionally these devices and modules may comprise an integrated circuit, a microprocessor, and / or a wireless communication device.
[0119] The devices and systems described with reference to embodiments in this disclosure, such as the client device, the server system and the content preparation system are typically implemented as one or more communicatively connected data processing systems. Fig. 11 is a block diagram illustrating an exemplary data processing system that may be used in as described in this disclosure. Data processing system 1100 may include at least one processor 1102 coupled to memory elements 1104 through a system bus 1106. As such, the data processing system may store program code within memory elements 1104. Further, processor 1102 may execute the program code accessed from memory elements 1104 via system bus 1106. In one aspect, data processing system may be implemented as a computer that is suitable for storing and / or executing program code. It should be appreciated, however, that data processing system may be implemented in the form of any system including a processor and memory that is capable of performing the functions described within this specification.
[0120] Memory elements 1104 may include one or more physical memory devices such as, for example, local memory 1108 and one or more bulk storage devices 1110. Local memory may refer to random access memory or other non-persistent memory device(s) generally used during actual execution of the program code. A bulk storage device may be implemented as a hard drive or other persistent data storage device. The data processing system 1100 may also include one or more cache memories (not shown) that provide temporary storage of at least some program code in order to reduce the number of times program code must be retrieved from bulk storage device 1110 during execution.
[0121] Input / output (I / O) devices depicted as input device 1112 and output device 1114 optionally can be coupled to the data processing system. Examples of input device may include, but are not limited to, for example, a keyboard, a pointing device such as a mouse, or the like. Examples of output device may include, but are not limited to, for example, a monitor or display, speakers, or the like. Input device and / or output device may be coupled to data processing system either directly or through intervening I / O controllers. A network adapter 1116 may also be coupled to data processing system to enable it to become coupled to other systems, computer systems, remote network devices, and / or remote storage devices through intervening private or public networks. The network adapter may comprise a data receiver for receiving data that is transmitted by said systems, devices and / or networks to said data and a data transmitter for transmitting data to said systems, devices and / or networks. Modems, cable modems, and Ethernet cards are examples of different types of network adapter that may be used with data processing system.
[0122] As pictured in Fig. 11, memory elements 1104 may store an application 1118. It should be appreciated that data processing system may further execute an operating system (not shown) that can facilitate execution of the application. Application, being implemented in the form of executable program code, can be executed by data processing system, e.g., by processor 1102. Responsive to executing application, data processing system may be configured to perform one or more operations to be described herein in further detail.
[0123] In one aspect, for example, data processing system may represent a client data processing system. In that case, application 1118 may represent a client application that, when executed, configures data processing system to perform the various functions described herein with reference to a "client". Examples of a client can include, but are not limited to, a personal computer, a portable computer, a mobile phone, or the like. In other aspects, data processing system may represent a server data processing system. In that case, application 1118 may represent a server application that, when executed, configures data processing system to perform the various functions described herein with reference to a "server".
[0124] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, various units may be combined in a codec hardware unit or provided by a collection of interoperative hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.
[0125] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0126] The corresponding structures, materials, acts, and equivalents of all means or step plus function elements in the claims below are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of the present invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the invention. The embodiment was chosen and described in order to best explain the principles of the invention and the practical application, and to enable others of ordinary skill in the art to understand the invention for various embodiments with various modifications as are suited to the particular use contemplated.
Claims
CLAIMS1. Method of processing media assets comprising: receiving a scene description file comprising information about at least one 3D object and one or more source locations for obtaining media assets for rendering the 3D object, the assets including mesh data associated with the 3D object and a video file comprising a sequence of video frames, each video frame comprising encoded texture data for adding a texture to a surface of the 3D object; retrieving the mesh data and the video frames based on the information in the scene description; and, rendering by a rendering device the 3D object on a display based on the mesh data and the video frames, the rending of the 3D object including:- allocating one or more buffers, each buffer being configured to store a decoded video frame and each of the one or more allocated buffers being accessible by the rendering device;- decoding the video frames into decoded video frames;- storing at least part of the decoded video frames in the one or more allocated buffers;- adding a plurality of textures to the surface of the 3D object on the basis of the texture data from the decoded video frames2. Method according to claim 1 wherein scene description file comprises information about the one or more buffers that should be allocated for rendering the 3D object and information about which decoded video frame from the decoded video frames should be stored in each of the one or more allocated buffers.
3. Method according to claim 2 wherein the one or more buffers that should be allocated are identified by a buffer identifier or a buffer index.
4. Method according to any of claims 1-3 wherein the scene description identifies each video frame in the video file, preferably the video frame being identified by a frame identifier or a frame index.
5. Method according to any of claims 1-4 wherein a presentation engine (PE) is configured to control the rendering of the 3D object based on the information in the scene description file, the presentation engine being further configured to instruct a media access function (MAF) to execute at least one of: the retrieval of the video frames, the allocation ofthe one or more buffers, the decoding of the video frames and the storage of at least part of the decoded video frames in the one or more allocated buffers.
6. Method according to claim 5 wherein the instruction of the MAF to the PE includes one of more of the following information items: information about one or more source locations for obtaining media assets for rendering the 3D object; information about one or more buffer identifiers or buffer indices for identifying the one or more buffers that should be allocated for rendering the 3D object; information about which decoded video frame from the decoded video frames should be stored in each of the one or more allocated buffers.
7. Method according to any of claims 1-6 wherein the rendering further includes: informing the rendering device that each of the one or more allocated buffers comprises a decoded video frame that comprises texture data.
8. Method according to any of claims 1-7 wherein the texture data of each video frame in the sequence of video frames is configured to add a different texture to the surface of the 3D object, the 3D object preferably being rendered with all textures of the sequence of video frames simultaneously.
9. Method according to any of claims 1-8 wherein the video frames are retrieved based on a streaming protocol, preferably an HTTP adaptive streaming protocol, such as MPEG-DASH.
10. Method according to any of claims 1-8 wherein the one or more buffers are GPU-based buffers.
11. An apparatus, preferably a client device, for processing media assets comprising: a computer readable storage medium having at least part of a program embodied therewith; and, a computer readable storage medium having computer readable program code embodied therewith, and a processor, preferably a microprocessor, coupled to the computer readable storage medium, wherein responsive to executing the computer readable program code, the processor is configured to perform executable operations comprising:receiving a scene description file comprising information about at least one 3D object and one or more source locations for obtaining media assets for rendering the 3D object, the assets including mesh data associated with the 3D object and a video file comprising a sequence of video frames, each video frame comprising encoded texture data for adding a texture to a surface of the 3D object; retrieving the mesh data and the video frames based on the information in the scene description; and, rendering, by a rendering device, the 3D object on a display based on the mesh data and the video frames, the rending of the 3D object including:- allocating one or more buffers, each buffer being configured to store a decoded video frame and each of the one or more allocated buffers being accessible by the rendering device;- decoding the video frames into decoded video frames;- storing at least part of the decoded video frames in the one or more allocated buffers.- adding a plurality of textures to the surface of the 3D object on the basis of the texture data from the decoded video frames12. An apparatus according to claim 11 wherein the processor is further configured to perform any of the method steps according to claims 2-10.
13. Computer program product comprising software code portions configured for, when run in the memory of a computer, executing the method steps according to any of 1-10.
Citation Information
Patent Citations
Using gltf2 extensions to support video and audio data
US20210099773A1