Method and apparatus for generating a scene description file

The method and apparatus generate a scene description file that integrates G-PCC coded point clouds into the MPEG scene description framework, addressing the lack of support for this media type and enhancing cross-platform compatibility and rendering efficiency.

JP2025529756AActive Publication Date: 2025-09-09HISENSE VISUAL TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025507589
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-27
Filing Date
2023-06-01
Publication Date
2025-09-09
Estimated Expiration
2043-06-01

AI Technical Summary

Technical Problem

The existing MPEG scene description standard does not support media files of the G-PCC coded point cloud type, which is an important 3D media format, necessitating a solution to integrate G-PCC coded point clouds into the scene description framework for cross-platform compatibility.

Method used

A method and apparatus for generating a scene description file that maps G-PCC coded point clouds to a point cloud based on description information, adding a target media description module to a media list in the scene description file, and utilizing MPEG scene description extensions to support dynamic scene updates and time-varying media.

Benefits of technology

Enables the integration of G-PCC coded point clouds into the scene description framework, supporting dynamic and time-varying media, thereby enhancing cross-platform compatibility and optimizing media rendering processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025529756000001_ABST
    Figure 2025529756000001_ABST
Patent Text Reader

Abstract

Some embodiments of the present application provide a method and apparatus for generating a scene description file in the video processing technical field, which includes determining a type of a media file in a 3D scene to be rendered, and if the type of the target media file in the 3D scene to be rendered is a geometry-based point cloud compressed G-PCC coded point cloud, generating a scene description file corresponding to the target media file based on description information of the target media file. Target Media Description Module and adding the target media description module to a media list of MPEG media in a scene description file of the three-dimensional scene to be rendered.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is This is a national phase application of a PCT application with application number PCT / CN2023 / 097873 filed on June 1, 2023, which was filed on January 10, 2023, with application number 202310036790.8 Chinese patent applications, and Priority is claimed to Chinese patent application No. 202310474240.4, filed on April 27, 2023, the contents of which are incorporated herein by reference. [Technical Field]

[0002] Some embodiments of the present application relate to the field of video processing, and more particularly to a method and apparatus for generating a scene description file. [Background technology]

[0003] A point cloud is a large collection of 3D points. There are two main compression standards for point clouds: Geometry-based Point Cloud Compression (G-PCC) and Video-based Point Cloud Compression (V-PCC).

[0004] With the development of immersive media and applications, the number of types of immersive media is increasing. Currently, mainstream immersive media mainly includes point clouds, 3D meshes, 6 DoF panoramic video, and MPEG immersive video (MIV). In 3D scenes, various types of immersive media often coexist. This requires rendering engines to support multiple different immersive media codecs. Different rendering engines are created depending on the type and number of supported codecs. Because rendering engines designed by different vendors support different media types, the Moving Picture Experts Group (MPEG) initiated the creation of an MPEG scene description standard, ISO / IEC 23090-14, to address the cross-platform description problem of 3D scenes in MPEG media (including the codecs, MPEG file formats, and MPEG transmission mechanisms established by MPEG). Although the first edition of the ISO / IEC 23090-14 MPEG-I scene description standard's extensions fulfilled an important need for immersive scene description solutions, the current scene description standard does not support media files of the G-PCC coded point cloud type. Point clouds are an important 3D media format, and G-PCC is one of the currently mainstream point cloud compression algorithms. Therefore, supporting media files of the G-PCC coded point cloud type in the scene description framework is of great significance and value. Summary of the Invention

[0005] According to a first aspect, some embodiments of the present application provide a method for generating a scene description file, the method comprising: determining the type of media files in the three-dimensional scene to be rendered; When the type of the target media file in the three-dimensional scene to be rendered is a G-PCC coded point cloud based on geometric point cloud compression, the target media file is mapped to a point cloud based on description information of the target media file. Target Media Description Module generating adding said target media description module to a media list of MPEG media in a scene description file of said three-dimensional scene to be rendered; A method for generating a scene description file including:

[0006] According to a second aspect, some embodiments of the present application provide a scene description file generation device, comprising: a memory configured to store a computer program; a processor configured to, when invoking a computer program, cause the scene description file generation device to implement the scene description file generation method according to the first aspect; The present invention provides a device for generating a scene description file including: [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 shows a schematic structural diagram of an immersive media description framework according to some embodiments. [Figure 2] FIG. 2 shows a schematic structural diagram of a scene description file according to some embodiments. [Figure 3] FIG. 3 shows a schematic structural diagram of a scene description file according to another embodiment of the present application. [Figure 4] FIG. 4 shows a schematic structural diagram of a G-PCC encoder according to some embodiments. [Figure 5] FIG. 5 shows a schematic diagram of the LOD splitting process according to some embodiments. [Figure 6] FIG. 6 shows a schematic diagram of a lifting transformation process according to some embodiments. [Figure 7] FIG. 7 shows a schematic diagram of the RAHT conversion process according to some embodiments. [Figure 8] FIG. 8 shows a schematic structural diagram of a G-PCC decoder according to some embodiments. [Figure 9] FIG. 9 shows a schematic structural diagram of a scene description file according to another embodiment. [Figure 10] FIG. 10 shows a schematic structural diagram of a scene description file according to another embodiment. [Figure 11] FIG. 11 shows a schematic diagram of a pipeline for handling G-PCC coded point cloud type media files, according to some embodiments. [Figure 12] FIG. 12 shows a flow chart of steps in a method for generating a scene description file according to some embodiments. [Figure 13] FIG. 13 shows a step flow diagram of a method for parsing a scene description file according to some embodiments. [Figure 14] FIG. 14 shows a step flow chart of a method for processing media files according to some embodiments. [Figure 15] FIG. 15 shows a step flow diagram of a method for rendering a three-dimensional scene according to some embodiments. [Figure 16] FIG. 16 illustrates a step flow diagram of a buffer management method according to some embodiments. [Figure 17] FIG. 17 shows an interaction flowchart of a method for rendering a three-dimensional scene according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0008] In order to clarify the purpose and embodiments of the present application, the exemplary embodiments of the present application will be described below clearly and completely with reference to the drawings in the exemplary embodiments of the present application, and obviously, the exemplary embodiments described are only some embodiments of the present application, but not all embodiments.

[0009] The brief explanation of terms in this application is intended to facilitate understanding of the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and customary meanings.

[0010] The terms "comprises" and "having" and any variations thereof are intended to cover but not exclusively include, for example, a product or device including a set of components is not necessarily limited to all components expressly listed, but may include other components not expressly listed or inherent to those products or devices.

[0011] References herein to "some implementations," "some embodiments," and the like indicate that the described implementations or embodiments may include a particular feature, structure, or characteristic, but do not necessarily mean that every embodiment includes that particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same implementation. Additionally, when a particular feature, structure, or characteristic is described in connection with one embodiment, it is believed to be within the knowledge of one skilled in the art that such feature, structure, or characteristic may also be implemented in connection with other implementations (whether or not expressly described herein).

[0012] The specification includes many parentheses, where the content in some parentheses is the English interpretation of the aforementioned terms, such as Media Access Function (MAF), Scene Description Documents, Application Programming Interface (API), etc., and the content in some parentheses indicates the abbreviation when the parameter is used in a computer program, code, or actual application, such as scene description module (scene), node description module (node), mesh description module (mesh), accessor description module (accessor), etc. Note that the above examples only explain some of the expressions in this disclosure, and understanding of more of the parenthesized content should be expressed according to the context.

[0013] Some embodiments of the present application relate to a scene description for immersive media. Referring to the immersive media scene description framework shown in FIG. 1 , the immersive media scene description framework decouples media file access and processing from media file rendering so that the display engine 11 can focus on media rendering, and designs a media access function (MAF) 12 to handle media file access and processing functions. At the same time, a media access function application programming interface (API) is designed, and command interaction is performed between the display engine 11 and the media access function 12 via the media access function API. The display engine 11 can send commands to the media access function 12 via the media access function API, and the media access function 12 can also request commands from the display engine 11 via the media access function API.

[0014] The general workflow of the immersive media scene description framework is as follows: 1) the display engine 11 reads a scene description file (Scene Description File) provided by the immersive media service provider. 2) the display engine 11 parses the scene description file to obtain parameters or information such as the access address of the media file, attribute information of the media file (such as the media type and codec parameters), and a format request for the processed media file, and calls the media access function API to transmit all or a part of the information obtained by parsing the scene description file to the media access function 12; 3) the media access function 12 requests download of the specified media file from the media resource server or obtains the specified media file locally based on the information transmitted from the display engine 11, establishes a corresponding pipeline for the media file, and then performs processes such as decapsulation, decryption, decoding, and post-processing on the media file in the pipeline to convert the media file from the encapsulated format to a format specified for the display engine 11; 4) the pipeline stores the output data after all processes have been completed in a specified buffer; and 5) finally, the display engine 11 reads the completely processed data from the specified buffer and renders the media file based on the data read from the buffer.

[0015] The files and functional modules associated with the immersive media scene description framework are further described below.

[0016] Scene description file

[0017] In the workflow of the immersive media scene description framework, a scene description file is used to describe the contents of a 3D scene, such as its structure (its features can be described by a 3D mesh), textures (e.g., texture mapping), animations (rotation, translation), and camera viewpoint (rendering angle).

[0018] In the related art, the GL Transmission Format 2.0 (glTF2.0) has been determined as a candidate format for scene description files, and can meet the requirements of Motion Picture Experts Group - Immersive (MPEG-I) and Six Degrees of Freedom (6DoF) applications. For example, the Khronos Group's GL Transmission Format (glTF) Version 2.0, available at github.com / KhronosGroup / glTF / tree / master / specification / 2.0#specifying-extensions, describes glTF2.0. Referring to Figure 2, Figure 2 is a schematic structural diagram of a scene description file in the glTF2.0 scene description standard (ISO / IEC 12113). As shown in FIG. 2, a scene description file in the glTF2.0 scene description standard includes, but is not limited to, a scene description module (scene) 201, a node description module (node) 202, a mesh description module (mesh) 203, an accessor description module (accessor) 204, a buffer slice description module (bufferView) 205, a buffer description module (buffer) 206, a camera description module (camera) 207, a light illumination description module (light) 208, a material description module (material) 209, a texture description module (texture) 210, a sampler description module (sampler) 211, a texture mapping description module (image) 212, an animation description module (animation) 213, and a skin description module (skin) 214.

[0019] 2 is used to describe a 3D scene included in the scene description file. One scene description file may contain any number of 3D scenes, each of which is represented by a scene description module 201. There is a parallel relationship between the scene description modules 201, i.e., there is a parallel relationship between 3D scenes.

[0020] The node description module (node) 202 in the scene description file shown in FIG. 2 is a hierarchical description module next to the scene description module 201 and is used to describe objects included in the 3D scene described by the scene description module 201. Each 3D scene may contain many concrete objects, such as a virtual digital human, nearby 3D objects, and distant background images, and the scene description file describes these concrete objects using the node description module 202. Each node description module 202 can represent one object or a set of objects, and the relationship between the node description modules 202 reflects the relationship between each component part in the 3D scene described by the scene description module 201. A scene described by one scene description module 201 may contain one or more nodes. Multiple nodes may have a parallel or hierarchical relationship, i.e., a "contains" and "is contained" relationship exists between the node description modules 202, thereby describing multiple concrete objects collectively or individually. When one node is contained within another, the contained node is called a subnode (children), and the subnode is referred to by replacing "node" with "children." Nodes and subnodes can be flexibly combined to form a hierarchical node structure, thereby expressing a rich scene content.

[0021] 2, the mesh description module (mesh) 203 is the next hierarchical description module after the node description module 202 and is used to describe the characteristics of the object represented by the node description module 202. The mesh description module 203 is a collection of one or more primitives, and each primitive may include one attribute, which defines the attributes that need to be used by a graphics processing unit (GPU) when rendering. The attributes may include position (three-dimensional coordinates), normal (normal vector), tangent (tangent vector), texcoord_n (texture coordinates), color_n (color: RGB or RGBA), joints_n (attribute related to the skin description module 214), and weights_n (attribute related to the skin description module 214). Because the mesh description module 203 contains a large number of vertices, each containing various attribute information, it is inconvenient to directly store the large amount of media data contained in the media file in the mesh description module 203 of the scene description file. Instead, the access address (Uniform Resource Identifier, URI) of the media file is indicated in the scene description file, and data in the media file can be downloaded when needed, thereby realizing the separation of the scene description file and the media file. Therefore, generally, the mesh description module 203 does not store media data, but stores the index value of the accessor description module (accessor) 204 corresponding to each attribute, and the accessor description module 204 points to the corresponding data in the buffer slice (bufferView) of the buffer (buffer).

[0022] In some embodiments, the scene description file and the media files may be merged to form one binary file, thereby reducing the number and variety of files.

[0023] Furthermore, a primitive in the mesh description module 203 may have a syntax element called mode, which is used to describe the topology structure when a graphics processing unit (GPU) draws a 3D mesh, for example, mode=0 represents a scattered point, mode=1 represents a line, and mode=4 represents a triangle.

[0024] For example, the following is a JSON example for the mesh description module 203: JPEG2025529756000070.jpg85164

[0025] In the above mesh description module 203, the value of “position” is 1, which points to the accessor description module 204 whose index is 1, and ultimately points to the vertex coordinate data stored in the buffer; the value of “color_0” is 2, which points to the accessor description module 204 whose index is 2, and ultimately points to the color data stored in the buffer.

[0026] The definitions of the syntax elements in the primitive attributes (mash.primitives.attributes) of the mesh description module 203 are shown in Table 1 below. [Table 1]

[0027] The definitions of the accessor types indexed in the primitive attributes (mash.primitives.attributes) of the mesh description module 203 are shown in Table 2 below. [Table 2]

[0028] The definition of the data types in the primitive attributes (mash.primitives.attributes) of the mesh description module 203 is shown in Table 3 below. [Table 3]

[0029] 2 , the accessor description module (accessor) 204, the buffer slice description module (bufferView) 205, and the buffer description module (buffer) 206 together realize a refined index for each layer of the data of the media file by the mesh description module 203. As described above, the mesh description module 203 does not store specific media data, but stores the index values ​​of the corresponding accessor description modules 204, and accesses the specific media data through the accessors described by the accessor description modules 204 indexed by the index values. The indexing process of media data by the mesh description module 203 includes: first, the index values ​​declared in the syntax elements in the mesh description module 203 point to the corresponding accessor description modules 204; second, the accessor description modules 204 point to the corresponding buffer slice description modules 205; and finally, the buffer slice description modules 205 point to the corresponding buffer description modules 206. 2, the buffer description module 206 mainly serves to point to the corresponding media file, and includes information such as the URI and byte length of the media file, and is used to describe a buffer that buffers the media data of the media file. A buffer may be divided into one or more buffer slices, and the buffer slice description module 205 mainly performs partial access to the media data in the buffer, and includes information such as the start byte offset of the accessed data and the byte length of the accessed data. The buffer slice description module 205 and the buffer description module 206 can realize partial access to the data of the media file. The accessor description module 204 mainly serves to add additional information to part of the data defined in the buffer slice description module 205, such as the data type, the number of data of that type, and the numerical range of that type of data.Such a three-layer structure can realize the function of retrieving part of the data from one media file, which is advantageous for accurate retrieval of data and for reducing the number of media files.

[0030] 2 is the next hierarchical description module after the node description module 202, and is used to describe visual information such as the viewpoint and viewing angle when a user views an object described by the node description module 202. In order to allow a user to be placed in a three-dimensional scene and to view the three-dimensional scene, the node description module 202 points to the camera description module 207, which can also describe visual information such as the viewpoint and viewing angle when a user views an object described by the node description module 202.

[0031] The light illumination description module (light) 208 in the scene description file shown in Figure 2 is the next hierarchical description module after the node description module 202, and is used to describe information about light illumination such as the light illumination intensity, ambient light color, light illumination direction, and light source position of the object described by the node description module 202.

[0032] The material description module (material) 209 in the scene description file shown in FIG. 2 is the next hierarchical description module after the mesh description module 203 and is used to describe material information of the 3D object described by the mesh description module 203. When describing a 3D object, describing the geometric information of the 3D object using the mesh description module 203 or simply defining the color and / or position of the 3D object cannot improve the realism of the 3D object; more information must be added to the surface of the 3D object. For 3D modeling technologies such as 3D mesh models, this process may be abbreviated as texture mapping or texture addition. Scene description files in the glTF 2.0 scene description standard also use this description module. The material description module 209 defines materials using a set of common parameters and describes the material information of geometric objects appearing in a 3D scene. The material description module 209 generally uses a metallic-roughness model to describe the material of a virtual object, and material property parameters based on the metallic-roughness model are expressed in the widely used physically based rendering (PBR) material. Based on this, the material description module 209 describes the metallic-roughness material attributes of the object in detail, and the definitions of the syntax elements in the material description module 209 are shown in Table 4. [Table 4]

[0033] In some embodiments, definitions of syntax elements in the metal-roughness (material.PbrMetarialRoughness) of the material description module 209 are shown in Table 5 below. [Table 5]

[0034] The value of each attribute in the metal-roughness material description module 209 may be defined using a coefficient and / or a texture (e.g., baseColorTexture and baseColorFactor). If no texture is provided, the values ​​of all corresponding texture components in this material model can be determined to be 1.0. If a coefficient and a texture are present at the same time, the coefficient value is a linear multiplier of the corresponding texture value. A texture binding is defined by a texture object index and a selectable texture coordinate index.

[0035] Illustratively, the following is a JSON example of the material description module 209: JPEG2025529756000076.jpg59164

[0036] By analyzing the material description module 209, it is possible to determine that the current material is named "gold" based on the material name syntax element and its value ("name":"gold"), and further determine that the base color value of the current material is [1.000, 0.766, 0.336, 1.0] based on the color syntax element and its value ("basecolorFactor":[1.000, 0.766, 0.336, 1.0]) in the pbrMetallicRoughness array, it is possible to determine that the metalness value of the current material is "1.0" based on the metallicity syntax element and its value ("metalnessFactor":1.0) in the pbrMetallicRoughness array, and it is possible to determine that the roughness value of the current material is "0.0" based on the roughness syntax element and its value ("roughnessFactor":0.0) in the pbrMetalRoughness array.

[0037] The texture description module (texture) 210 in the scene description file shown in FIG. 2 is the next hierarchical description module after the material description module 209. It is used to describe the color and other characteristics used in material definition of the three-dimensional object described by the material description module 209. Texture is one of the important aspects that gives an object a realistic appearance. By defining the object's primary color and other characteristics used in material definition through texture, the appearance of the rendered object can be accurately described. A material itself can define multiple texture objects, which can be used as textures for virtual objects during rendering and can be used to encode different material attributes. The texture description module 210 references one sampler description module (sampler) 211 and one texture mapping description module (image) 212 using sampler syntax element and texture mapping syntax element indexes. The texture mapping description module 212 contains a uniform resource identifier (URI) linked to the texture map or binary file capsule actually used by the texture description module 210. The sampler description module 211 is a filtering and encapsulation module for describing textures. The roles and cooperation of the material description module 209, texture description module 210, sampler description module 211, and texture mapping description module 212 include the following: The material description module 209 and texture description module 210 together define the color and physical information of the object surface; The sampler description module 211 defines how to apply texture mapping to the object surface.The texture description module 210 specifies the sampler description module 211 and the texture mapping description module 212, which realizes adding textures, which uses URIs for identification and indexing, and uses the accessor description module 204 for data access. The sampler description module 211 realizes the specific adjustment and encapsulation of textures. The definitions of the syntax elements in the texture description module 210 are shown in Table 6 below. [Table 6]

[0038] In some embodiments, the definitions of syntax elements in the texture description module 210 sample (texture.sample) are shown in Table 7 below. [Table 7]

[0039] Illustratively, the following are JSON examples of one material description module 209, texture description module 210, sampler description module 211, and texture mapping description module 212: JPEG2025529756000079.jpg162166

[0040] 2 is a hierarchical description module next to the node description module 202, and is used to describe animation information to be added to an object described by the node description module 202. Since animation can be added to an object described by the node description module 202 so that the object represented by the node description module 202 is not limited to a static state, the description hierarchy of the animation description module 213 in the scene description file is specified by the node description module 202; that is, the animation description module 213 is a hierarchical description module next to the node description module 202, and similarly has a corresponding relationship with the mesh description module 203. The animation description module 213 describes animation in three ways: position movement, angle rotation, and size scaling, and can specify the start and end times of the animation and the implementation method of the animation. For example, when an animation is added to the mesh description module 203 representing a three-dimensional object, the three-dimensional object represented by the mesh description module 203 can complete a predetermined animation process by combining position movement, angle rotation, and size scaling within a specified time window.

[0041] 2, the skin description module (skin) 214 is the next hierarchical description module after the node description module 202 and is used to describe the motion linkage relationship between the skeleton added to the node described by the node description module 202 and the mesh representing the object surface information. When the node described by the node description module 202 represents an object with a large degree of freedom of movement, such as a person, animal, or machine, the skeleton can be filled into the object to optimize the motion representation effect of the object, and the 3D mesh representing the object surface information then becomes the skin conceptually. The description hierarchy called the skin description module 214 is specified by the node description module 202; that is, the skin description module 214 is the next hierarchical description module after the node description module 202, and there is a correspondence between the skin description module 214 and the mesh description module 203. By moving the mesh on the surface of an object in conjunction with the movement of the skeleton and combining it with a bionic simulation design, it is possible to achieve a relatively realistic movement effect. For example, when a person's hand makes a fist, the skin on the surface stretches along with the internal skeleton, causing changes such as occlusion. At this time, by redefining the cooperative relationship between the skeleton and the skin for the skeleton pre-filled in the hand model, a realistic simulation of this movement can be achieved.

[0042] Each description module in the scene description file in the glTF 2.0 scene description standard only has the most basic capabilities for describing 3D objects, and has problems such as being unable to support dynamic 3D immersive media, audio files, or scene updates. glTF also states that each object attribute has one selectable extension object attribute, allowing any part of the object attribute to be extended using extensions to achieve more complete functionality. The scene description module (scene), node description module (node), mesh description module (mesh), accessor description module (accessor), buffer description module (buffer), animation description module (animation), and their internally defined syntax elements all have selectable extension object attributes to support certain functional extensions based on glTF 2.0.

[0043] Currently, rendering engines designed by different vendors support different media types. To address this issue, the Moving Picture Experts Group (MPEG) has initiated the development of the MPEG Scene Description standard, designated ISO / IEC 23090-14. This standard primarily addresses the cross-platform description problem of 3D scenes in MPEG media (including codecs developed by MPEG, MPEG file formats, and MPEG transmission mechanisms).

[0044] The MPEG#128th meeting resolution will create an MPEG-I Scene Description standard based on glTF2.0 (ISO / IEC 12113). The first version of the MPEG Scene Description standard is currently in the FDIS ballot stage. Based on the first version standard, the MPEG Scene Description standard adds corresponding extensions to address requirements not yet realized in cross-platform 3D scene description, including interactivity, AR anchors, user and avatar representations, haptic support, and extensions to support for immersive media codecs.

[0045] The first edition of the MPEG scene description standard created mainly creates the following contents:

[0046] The MPEG Scene Description standard defines a scene description file format for describing immersive 3D scenes, which combines the content of the original glTF2.0 (ISO / IEC 12113) and builds on it with a series of extensions.

[0047] MPEG scene description defines a scene description framework and an application programming interface (API) for inter-module communication within it, enabling decoupling of the immersive media acquisition and processing process from the media rendering process, and is beneficial for optimizing aspects such as adapting to different network conditions for immersive media, acquiring partial immersive media files, accessing different levels of immersive media detail, and adjusting content quality. Decoupling the immersive media acquisition and processing process from the immersive media rendering process is the key to realizing cross-platform description of 3D scenes.

[0048] MPEG scene description extensions based on the International Standardization Organization Base Media File Format (ISOBMFF) series (ISO / IEC 14496-12) have been proposed for use in the transmission of immersive media content.

[0049] As shown in Figure 3, based on the scene description file shown in Figure 2, the scene description file in the MPEG scene description standard is extended, and compared with the scene description file in the glTF2.0 scene description standard (the scene description file shown in Figure 2), the extensions of the scene description file in the MPEG scene description standard can be divided into two groups.

[0050] The first group of extensions includes MPEG media (MPEG_media) 301, MPEG time-varying accessor (MPEG_accessor_timed) 302, and MPEG ring buffer (MPEG_buffer_circular) 303. Here, MPEG media 301 is an independent extension for referencing external media sources, MPEG time-varying accessor 302 is an extension of the accessor hierarchy for accessing time-varying media, and MPEG ring buffer is an extension of the buffer hierarchy for supporting ring buffers. The first group of extensions provides basic descriptions and formats for media in a scene and meets the basic needs of describing time-varying immersive media in a scene description framework. Among them, MPEG time-varying accessor (MPEG_accessor_timed) 302 is for accessing time-varying media. Because the glTF 2.0 scene description standard does not support time-varying media, if media data needs to change over time, it must be achieved by updating the scene description file in the glTF 2.0 scene description standard. For example, in the glTF2.0 scene description standard, the texture mapping of an object surface can change over time, so if the texture mapping of the object surface needs to be updated, the scene description file in the glTF2.0 scene description standard must be updated. Frequent updates to the scene description file require frequent parsing, processing, and transmission of the scene description file, which increases performance overhead in the 3D scene rendering process. Based on this, MPEG designed an MPEG time-varying accessor (MPEG_accessor_timed) 302, which allows parameters in the MPEG time-varying accessor to change over time, changes the access method for media data, and enables the accessed data to change over time, thereby avoiding the frequent parsing, processing, and transmission of the scene description file.

[0051] The second group of extensions includes MPEG dynamic scene (MPEG_scene_dynamic) 304, MPEG texture (MPEG_texture_video) 305, MPEG audio spatial (MPEG_audio_spatial) 306, MPEG view recommendation (MPEG_viewport_recommended) 307, MPEG mesh mapping (MPEG_mesh_linking) 308, and MPEG animation timing (MPEG_animation_timing) 309. Here, MPEG_scene_dynamic 304 is a scene hierarchy extension to support dynamic scene updates, MPEG_texture_video 305 is a texture hierarchy extension to support video-style textures, MPEG_audio_spatial 306 is a node hierarchy and camera hierarchy extension to support spatial 3D audio, MPEG_viewport_recommended 307 is a scene hierarchy extension to support describing a recommended viewing angle when displaying in 2D, MPEG_mesh_linking 308 is a mesh hierarchy extension to support linking two meshes to provide mapping information, and MPEG_animation_timing 309 is a scene hierarchy extension to support controlling the animation timeline.

[0052] Each of the above extensions will be described in detail below.

[0053] The MPEG media in the MPEG scene description file is used to describe the type of media file, and the necessary descriptions are provided for the MPEG type media file to subsequently obtain these MPEG type media files. Among them, the definitions of the syntax elements of the first layer of MPEG media are shown in Table 8 below. [Table 8]

[0054] The definitions of the syntax elements in the MPEG media media list (MPEG_media.media) are shown in Table 9 below. [Table 9]

[0055] The definitions of the syntax elements within the MPEG media options in the media list (MPEG_media.alternatives) are shown in Table 10 below. [Table 10]

[0056] The definitions of the syntax elements in the MPEG media alternatives media list options tracks array (MPEG_media.alternatives.tracks) are shown in Table 11 below. [Table 11]

[0057] ISO / IEC 23090-14 also defines a transmission format for the delivery of scene description files and data related to extensions to glTF 2.0, based on ISOBMFF (ISO / IEC 14496-12). To facilitate the delivery of scene description files to clients, ISO / IEC 23090-14 defines how glTF files and associated data are encapsulated in ISOBMFF files as both non-time-varying and time-varying data (e.g., track samples). MPEG_scene_dynamic, MPEG_mesh_linking, and MPEG_animation_timing provide the display engine with specific time-varying data, and the display engine 11 should perform corresponding operations based on this changed information. ISO / IEC 23090-14 also defines the format of each extension's time-varying data and how it is encapsulated in ISOBMFF files. MPEF media (MPEG_media) can reference external media streams delivered via protocols such as RTP / SRTP and MPEG-DASH. To make it possible to address media flows without knowing the actual protocol solution, hostname, or port values, ISO / IEC 23090-14 defines a new Uniform Resource Locator (URL) scheme that requires the presence of one stream identifier in the query part, but does not specify a specific type of identifier; it allows the use of a Media Stream Identification scheme (RFC5888), a labeling scheme (RFC4575), or a zero-based indexing scheme.

[0058] Display Engine

[0059] 1 , the workflow of the immersive media scene description framework mainly includes: the display engine 11 acquiring a scene description file, parsing the acquired scene description file to acquire the structural structure of the 3D scene to be rendered and detailed information about the 3D scene to be rendered, and rendering and displaying the 3D scene to be rendered based on the information acquired by parsing the scene description file. The embodiments of the present application are not limited to the specific workflow and principles of the display engine 11, but are based on the display engine 11 parsing the scene description document, sending instructions to the media access function 12 via the media access function API, sending instructions to the buffer management module 13 via the buffer API, retrieving the processed data from the buffer, and completing the rendering and display of the 3D scene and the objects therein.

[0060] Media Access Functions

[0061] In the workflow of the immersive media scene description framework, the media access function 12 can receive commands from the display engine 11 and complete functions of accessing and processing media files according to the commands sent by the display engine 11. Specifically, this includes obtaining media files and then processing the media files. There are significant differences in the processing processes for different types of media files. In order to support a wide range of media types and also take into account the work efficiency of the media access function, various pipelines can be designed in the media access function, and the pipeline that matches the media type can be enabled in the processing process.

[0062] The input of the pipeline is media files downloaded from a server or read from a local storage control. These media files generally have complex structures and cannot be directly used by the display engine 11, so the main function of the pipeline is to process the data of such media files and make the data of the media files conform to the requirements of the display engine 11.

[0063] In the workflow of the immersive media scene description framework, the media data that has completed the pipeline processing must be passed to the display engine 11 in a standard array structure for use, which requires the participation of the buffer API and buffer management module 13. Module creates a corresponding buffer based on the format of the processed media data and is responsible for subsequent management of the buffer, such as operations to update and release. The buffer management module 13 may communicate with the media access function 12 or the display engine 11 via a buffer API, and the goal of communication with the display engine 11 and / or the media access function 12 is to realize buffer management. When the buffer management module 13 communicates with the media access function 12, the display engine 11 must first send buffer management-related commands to the media access function 12 via the media access function API, and the media access function 12 must then send buffer management-related commands to the buffer management module 13 via the buffer API. When the buffer management module 13 communicates with the display engine 11, the display engine 11 can directly send buffer management description information parsed from the scene description document to the buffer management module 13 via the buffer API.

[0064] The above embodiment introduces the basic flow of scene description framework rendering including a 3D scene of immersive media, and the content and role of each functional module or file in the scene description framework. The immersive media in the 3D scene may be a point cloud-based media file, a 3D mesh-based media file, a 6 DoF-based media file, an MIV media file, etc. Since some embodiments of the present application relate to rendering a 3D scene including a point cloud based on the scene description framework, the following description will first discuss point cloud-related content.

[0065] A point cloud is a collection of a large number of three-dimensional points. After obtaining the spatial coordinates of each sampling point on the surface of an object, the resulting set of points is called a point cloud. In addition to geometric coordinates, points in a point cloud may further include other attribute information such as color, normal vector, reflectance, transparency, and material type. Point clouds can be obtained in various ways. In some embodiments, the method for obtaining a point cloud includes observing an object using a camera array whose fixed position in space is known, and obtaining a three-dimensional representation of the object using several related algorithms based on two-dimensional images captured by the camera array to obtain a point cloud corresponding to the object. In other embodiments, the method for obtaining a point cloud includes obtaining a point cloud corresponding to the object using a laser radar scanning device. The sensor of the laser radar scanning device obtains volumetric information of the object by recording electromagnetic waves from a radar reflected by the surface of the object, and then obtaining a point cloud corresponding to the object based on the volumetric information of the object. In another embodiment, the method for obtaining a point cloud may further include obtaining a point cloud corresponding to the object by using an artificial intelligence or computer vision algorithm to create three-dimensional volumetric information based on the two-dimensional images.

[0066] Point clouds provide a highly accurate 3D representation method for detailed digitization of the physical world and are widely applied in fields such as 3D modeling, smart cities, autonomous navigation systems, and augmented reality. However, due to characteristics such as large data volume, unstructured nature, and uneven density, the storage and transmission of point clouds faces enormous challenges. Therefore, efficient compression of point clouds is necessary. Currently, there are two main compression standards for point clouds: geometry-based point cloud compression (G-PCC) and video-based point cloud compression (V-PCC). The principles and related algorithms of G-PCC are further explained below.

[0067] As shown in FIG. 4, the G-PCC encoder 400 may be divided into two parts: a geometric coding module 41 and an attribute coding module 42, and the geometric coding module 41 may be further divided into an octree-based geometric coding unit 411 and a prediction tree-based geometric coding unit 412.

[0068] As shown in FIG. 4, the main steps of the geometric encoding module 41 of the G-PCC encoder for encoding the geometric information of the point cloud to be encoded include step S401 of extracting geometric information (positions) from the point cloud to be encoded, step S402 of performing coordinate transformation on the geometric information to include all of the point cloud to be encoded in a single bounding box, and step S403 of voxelizing the geometric information after coordinate transformation. That is, first, the geometric information after coordinate transformation is quantized to scale the point cloud to be encoded. Because quantization rounding causes some points in the point cloud to have the same position, quantizing the geometric information after coordinate transformation requires determining whether or not to remove duplicate points based on a parameter. The process of quantization and removing duplicate points is called the voxelization process. After voxelization of the geometric information is completed, the octree-based geometric encoding unit 411 and the predictive tree-based geometric encoding unit 412 perform encoding, respectively, to obtain a geometric information codestream for the point cloud to be encoded.

[0069] The encoding process of the octree-based geometric encoding unit 411 includes tree partitioning (S404), which involves continuously performing tree partitioning (octree / quadtree / binary tree) on the bounding box in a breadth-first search order and encoding the bit code of each node. That is, the bounding box is sequentially partitioned to obtain subcubes, and non-empty subcubes (containing points in the point cloud) are partitioned until the resulting leaf node is a 1x1x1 unit cube. Next, the number of points contained in the leaf node is encoded, and finally the geometric octree encoding is completed, generating a binary code stream. S405: Surface fitting is performed on the geometric information based on triangle soup. Similarly, surface fitting first performs octree division, but there is no need to gradually divide the point cloud to be encoded into unit cubes with side lengths of 1x1x1. Instead, the division stops when the side length of a sub-block reaches a predetermined value. Next, based on the surface formed by the distribution of the point cloud in each sub-block, up to 12 intersections (vertices) generated by the surface and the 12 sides of the sub-block are obtained, and the intersection coordinates of each sub-block are sequentially encoded to generate a binary code stream.

[0070] The encoding process of the prediction tree-based geometric encoding unit 412 includes the following steps: S406: constructing a prediction tree structure, which includes sorting points in the point cloud to be encoded, including no order, Morton order, azimuth order, and radial distance order, and constructing a prediction tree structure using two different sorting methods (high-latency slow method and low-latency fast method); S407: traversing each node in the prediction tree based on the prediction tree structure, selecting a different prediction mode to predict the geometric position information of the node to obtain a prediction residual, and quantizing the geometric prediction residual using a quantization parameter; S408: arithmetic coding, which includes arithmetically coding the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameter through successive iterations to generate a binary geometric information code stream.

[0071] 4, the process of the G-PCC encoder attribute encoding module 42 encoding attribute information of a point cloud to be encoded mainly includes: S408: extracting attribute information (attributes) from the point cloud to be encoded; S409: performing attribute prediction on the attribute information; S410: performing a lifting transform on the attribute information; S411: performing a Region Adaptive Hierarchical Transform (RAHT) transform on the attribute information; S412: quantizing the RAHT transform coefficients and the lifting transform coefficients; and S413: performing arithmetic coding on the quantized RAHT transform coefficients and the lifting transform coefficients to obtain an attribute information code stream. Furthermore, since the attribute encoding module 42 performs processing based on the reconstructed geometric information, after the lossy geometric encoding is completed, it must perform steps S414: reconstructing geometric information based on the geometry code stream and matching the original attribute information (attributes) with the reconstructed geometric information; and S415: recoloring the geometric information. Here, the re-coloring part in step S415 is to assign attribute information to the reconstructed point group using the original point group, with the aim of making the attribute values ​​of the reconstructed point group as similar as possible to the attribute values ​​of the point group to be encoded, thereby minimizing errors.

[0072] The attribute prediction algorithm is an algorithm that obtains a predicted attribute value of a current prediction target point by weighting and adding the reconstructed attribute values ​​of reconstructed points in three-dimensional space. The attribute prediction algorithm can effectively remove redundancy in the attribute space to achieve the purpose of compressing attribute information. In some embodiments, the implementation of attribute prediction may include the following: First, a level of detail (LOD) algorithm is used to hierarchically divide the point cloud to be coded, thereby constructing a hierarchical structure of the point cloud to be coded. Next, points in lower layers are coded and decoded first, and points in higher layers are predicted using reconstructed points in the same layer as the points in the lower layers, thereby achieving progressive coding. Here, the implementation of hierarchically dividing the point cloud to be coded using the LOD algorithm may include the following: First, all points in the point cloud to be coded are marked as unaccessed, and the accessed point cloud is denoted as V. In the initial state, the accessed point cloud is empty. Circularly traverse all unvisited points in the point cloud to be coded, calculate the minimum distance D from the current point to the accessed point cloud V, and if D is smaller than the threshold distance, ignore the current point; otherwise, mark the current point as visited and add it to the accessed point cloud V and the current subspace. Finally, merge the points in each subspace and all subspaces before each subspace to obtain a hierarchical structure of the point cloud to be coded.

[0073] 5, the point cloud to be encoded includes points P1-P9. In the distance-based LOD division process, in the first circular traversal, points P0, P2, P4, and P5 are sequentially added to the accessed point cloud V and layer R0; in the second circular traversal, points P1, P3, and P8 are sequentially added to the accessed point cloud V and layer R1; in the third circular traversal, traversal of all points is completed, and points P6, P7, and P9 are sequentially added to the accessed point cloud V and layer R2; and finally, the points in each layer and all layers before each layer are merged to obtain a hierarchical structure of the point cloud to be encoded, which includes three layers. Here, the first layer is LOD0 and includes points P0, P2, P4, and P5, the second layer is LOD1 and includes points P0, P2, P4, P5, P1, P3, and P8, and the third layer is LOD2 and includes points P0, P2, P4, P5, P1, P3, P8, P6, P7, and P9.

[0074] The lifting transform is built on the predictive transform and includes three parts: partitioning, prediction, and update. As shown in FIG. 6, the partitioning module 61 spatially partitions the point cloud to be coded into two parts: a high-level point cloud H(N) and a low-level point cloud L(N), where there is a certain correlation between the high-level point cloud H(N) and the low-level point cloud L(N). The prediction module 62 performs attribute prediction on the high-level point cloud H(N) using the attributes of the low-level point cloud L(N), obtaining a prediction residual D(N) = H(N) - P(N). Here, P(N) is the feature output by the prediction module 62 after predicting the low-level point cloud L(N). During the partitioning module 61 and prediction module 62, points in the low-level LOD layer are more influential due to the prediction strategy in the LOD partitioning. Therefore, the update module 63 defines and recursively updates the influence weight of each point based on the prediction residual D(N) and the distance between the predicted point and its neighboring points. Here, the recursive update refers to performing multiple lifting transformations, and the output data of the previous lifting transformation is the input data of the next lifting transformation. The recursive update by defining the influence weight of each point based on the prediction residual D(N) and the distance between the predicted point and its adjacent points includes the recursive update by defining the influence weight of each point based on the prediction residual D(N), the distance between the predicted point and its adjacent points, and the formula L'(N) = L(N) + U(N), where U(N) is the feature output by the update module 63 after predicting the prediction residual D(N).

[0075] The RAHT transform is a hierarchical domain adaptive transformation algorithm based on the Haar wavelet transform. Based on a hierarchical tree structure, it recursively transforms occupied subnodes in the same parent node from bottom to top along each dimension, and then transmits the resulting low-frequency coefficients to the next layer of the transformation process, where the high-frequency coefficients are quantized and entropy coded.

[0076] In some embodiments, the RAHT transformation can be implemented using an RAHT transformation based on upsampling prediction. In the RAHT transformation based on upsampling prediction, the tree structure of the entire RAHT transformation is changed from bottom-up to top-down, and the transformation is still performed within a 2x2x2 block. As shown in FIG. 7, within a 2x2x2 block, the transformation flow includes the following: First, the RAHT transformation is performed on a voxel block 71 in the first direction. If there are adjacent voxel blocks in the first direction, the RAHT transformation is performed on both voxel blocks to obtain the weighted average (DC coefficient) and residual (AC coefficient) of the attribute values ​​of the two adjacent points. The obtained DC coefficient is then used as the attribute information of the voxel block 122 of the parent node, and the RAHT transformation of the next layer is performed. The AC coefficient is reserved for final encoding. If there are no adjacent points, the attribute value of the voxel block 71 is directly transmitted to the parent node in the second layer. The RAHT transformation for the second layer is performed along the second direction. If there are adjacent voxel blocks in the second direction, both undergo RAHT transformation to obtain the weighted average (DC coefficient) and residual (AC coefficient) of the attribute values ​​of the two adjacent points. The RAHT transformation for the third layer is then performed along the third direction. The parent node voxel block 73, where the three color depths correspond, is obtained as a subnode in the next layer of the octree. The RAHT transformation is then repeated along the first, second, and third directions until only one parent node exists in the entire point group to be coded.

[0077] As shown in FIG. 8, the G-PCC decoder 800 may be divided into a geometric decoding module 81 and an attribute decoding module 82, and the geometric decoding module 81 may be further divided into an octree-based geometric decoding unit 811 and a prediction tree-based geometric decoding unit 812.

[0078] 8, the main steps of the G-PCC decoder's octree-based geometric decoding unit 811 in the geometric decoding module 81 decoding the geometric information code stream include S801 (arithmetic decoding), S802 (octree synthesis), S803 (surface fitting), S804 (geometry reconstruction), and S805 (inverse coordinate transformation), to obtain the geometric information of the point cloud. Here, the geometric decoding of the octree-based geometric decoding unit 811 involves: obtaining the bit code of each node through successive analysis in breadth-first traversal order; successively dividing the node until a 1x1x1 unit cube is divided; obtaining the number of points contained in each leaf node through analysis; and finally restoring the geometric reconstruction point cloud information. The main steps of the G-PCC decoder's geometric information codestream decoding by the geometric decoding unit 812 based on the prediction tree of the geometric decoding module 81 include steps S801 (arithmetic decoding), S806 (reconstruction of the prediction tree), S807 (residual calculation), S804 (reconstruction of the geometry), and S805 (inverse coordinate transformation) to obtain the geometric information of the point cloud. The main steps of attribute decoding performed by the attribute decoding module 82 based on the G-PCC decoder 800 include steps S808 (arithmetic decoding), S809 (inverse quantization), S810 and S811 (or S812), S810 (attribute prediction), S811 (boosting transformation), S812 (inverse transformation based on RAHT), and S813 (inverse color transformation) to obtain the attribute information of the point cloud. Finally, a 3D image model of the point cloud data to be encoded is reconstructed based on the geometric information and attribute information. The main steps of the G-PCC decoder decoding the attribute information codestream based on the attribute decoding module 82 and the main steps of the G-PCC encoder encoding the attribute information based on the attribute encoding module 82 are reverse processes, so they will not be described here.

[0079] Currently, the first edition of the ISO / IEC 23090-14 MPEG-I scene description standard has been extended to address important needs for immersive scene description solutions, including virtual scene interaction, AR anchors, user virtual person display, haptic support, and immersive codec support. Point clouds are an important immersive 3D media format in 3D environments, so supporting point cloud media display in scene description standards is an important part of scene description. Geometry-based point cloud compression (G-PCC) is one of the mainstream point cloud compression algorithms, so supporting media files with G-PCC coded point clouds in scene descriptions is significant and valuable.

[0080] Some embodiments of the present application provide support for a scene description framework including a point cloud code stream obtained by the G-PCC compression standard, and specific contents include support for media files of type G-PCC coded point clouds by a scene description file, support for media files of type G-PCC coded point clouds by a media access function API, support for media files of type G-PCC coded point clouds by a media access function, support for media files of type G-PCC coded point clouds by a buffer API, and support for media files of type G-PCC coded point clouds by buffer management.

[0081] The process of rendering media files of type G-PCC coded point clouds in a 3D scene based on the scene description framework includes the following: First, the display engine obtains a scene description file by, for example, downloading or local reading. The scene description file includes description information for the entire 3D scene and the media files of type G-PCC coded point clouds included in the scene. The description information for the media files of type G-PCC coded point clouds may include the access addresses of the media files of type G-PCC coded point clouds, the storage format of the decoded data for the processed media files of type G-PCC coded point clouds, the playback time and playback frame rate of the media files of type G-PCC coded point clouds, etc. After analyzing the scene description file, the display engine transmits the description information for the media files of type G-PCC coded point clouds included in the scene description to the media access function via the media access function API. At the same time, the display engine may call a buffer management module via the buffer API to allocate buffers and transmit the buffer information to the media access function, which then calls the buffer management module via the buffer API to allocate buffers. After receiving the description information transmitted from the display engine, the media access function first requests the server to download the media file whose type is a G-PCC coded point cloud, or reads the media file whose type is a G-PCC coded point cloud from a local file. After obtaining the media file whose type is a G-PCC coded point cloud, the media access function creates and starts a corresponding pipeline to process the media file whose type is a G-PCC coded point cloud. The input of the pipeline is the capsule file of the media file whose type is a G-PCC coded point cloud. The pipeline sequentially performs processes such as decapsulation, G-PCC decoding, and post-processing, and then stores the processed data in a specified buffer.Finally, the display engine retrieves the decoded data of the media file of type G-PCC coded point cloud from the specified buffer, and renders and displays the three-dimensional scene based on the data retrieved from the buffer.

[0082] Below, we will explain the scene description file, media access function API, media access functions, buffer API, and buffer management that support media files whose type is G-PCC coded point cloud.

[0083] Scene description file supporting media files of type G-PCC coded point cloud

[0084] In order to enable a scene description file to accurately describe a media file of the G-PCC coded point cloud type, some embodiments of the present application extend the values ​​of syntax elements in the MPEG media (MPEG_media) of the scene description file, and specifically, the extensions include at least one of the following:

[0085] Extension 1: The media type syntax element (MPEG_media.media.alternatives.mimeType) has been extended to declare the encapsulation format of a media file in the options (MPEG_media.media.alternatives) of the media list (media) of the MPEG media (MPEG_media) in the scene description file. The extension of the media type syntax element (mime Type) includes extending the value "application / mp4" associated with G-PCC coded point clouds to the media type syntax element (mime Type). If the type of the media file is a G-PCC coded point cloud, the value of the media type syntax element (mimeType) is "application / mp4". For example, mimeType:application / mp4.

[0086] Extension 2 extends the value of the first track index syntax element (MPEG_media.media.alternatives.tracks.tracks) to declare track information of a media file in the optional track array (MPEG_media.media.alternatives.tracks) of the media list (media) of MPEG media (MPEG_media) in a scene description file. The extension of the first track index syntax element (MPEG_media.media.alternatives.tracks.track) includes the following: if G-PCC data is referenced by a scene description file as one of the optional track arrays of the media list of MPEG media, and the reference conforms to the specifications for tracks in the International Standardization Organization Base Media File Format (ISOBMFF), then for single-track encapsulated G-PCC data, the referenced track in the MPEG media is a G-PCC codestream track, and for multi-track encapsulated G-PCC data, the referenced track in the MPEG media is a G-PCC geometric codestream track.

[0087] Extension 3: Extends the codec parameter syntax element (MPEG_media.media.alternatives.tracks.codecs) to describe the codec parameters of media data included in codestream tracks in the tracks array (tracks) of the options (alternatives) of the media list (media) of the MPEG media (MPEG_media) in the scene description file. Specific extensions include extending the codec parameters of media files included in codestream tracks defined in IETF RFC6381. When a codestream track contains multiple codec parameters of different types (for example, when encapsulating a G-PCC coded point cloud using DASH, the AdaptationSet contains representations with different codecs), the codec parameters syntax element (codecs) can be represented by a comma-separated codec value list, so extending the retrieval of the value of the syntax element codec parameters syntax element (codecs) includes that when the type of the media file is a G-PCC coded point cloud, the value of the codec parameters syntax element (codecs) should be set in accordance with the provisions in the ISO / IEC23090-18 G-PCC Data Transmission (Carriage of Geometry-based Point Cloud Compression Data) standard. For example, if the G-PCC data employs DASH capsules, when G-PCC pre-selection signaling is used in the Media Presentation Description (MPD) file, the 'codecs' attribute of the pre-selection signaling should be set to 'gpc1' to indicate that the pre-selection media is based on geometric point clouds, and when multiple G-PCC Tile tracks exist in the G-PCC container, the 'codecs' attribute of the Main G-PCC Adaptation Set should be set to 'gpcb' or 'gpeb' to indicate that the adaptation set contains G-PCC Tile basic track data.If Tile Component Adaptation Sets signal only a single G-PCC component data, the 'codecs' attribute of the Main G-PCC Adaptation Set should be set to 'gpcb'. If Tile Component Adaptation Sets signal all G-PCC component data, the 'codecs' attribute of the Main G-PCC Adaptation Set should be set to 'gpeb'. If G-PCC Tile Preselection Signaling is used in the MPD file, the 'codecs' attribute of the Preselection Signaling should be set to 'gpt1', which indicates that the Preselection Media is a geometric point cloud fragment.

[0088] In summary, to enable a scene description file to accurately describe a media file of the G-PCC coded point cloud type, some implementations of the present application extend the values ​​of syntax elements in MPEG media (MPEG_media) in the scene description file, with specific extensions including one or more of the extensions shown in Table 12 below. [Table 12]

[0089] By implementing at least one of the above extensions 1 to 3 for obtaining values ​​of syntax elements in MPEG media in scene description files (MPEG_media), MPEG media in scene description files (MPEG_media) partially supports media files of type G-PCC coded point cloud.

[0090] In some embodiments, the scene and node description method in a scene description file including media files of type G-PCC coded point clouds includes, when a 3D scene includes media files of type G-PCC coded point clouds, using the scene and node description method to describe the overall structure of the 3D scene and the structural hierarchy and positions of the media files of type G-PCC coded point clouds in the 3D scene. Here, using the scene description module and node description module description method to describe the overall structure of the 3D scene and the structural hierarchy and positions of the media files of type G-PCC coded point clouds in the 3D scene includes describing one 3D scene using one scene description module. Each scene description file can describe one or more 3D scenes, and the 3D scenes have only a parallel relationship, not a hierarchical relationship. The nodes may have either a parallel or hierarchical relationship.

[0091] In some embodiments, a method for describing a 3D mesh in a scene description file supporting a media file of type G-PCC coded point cloud includes multiplexing various types of data in a media file whose syntax element description type is G-PCC coded point cloud in the primitive attributes (mesh.primitives.attributes) of a mesh description module. Specifically, a point cloud is a distributed data structure, and a collection of many scattered points is a point cloud. Therefore, describing a media file of type G-PCC coded point cloud corresponds to describing data at each point in the point cloud. Generally, each point in a media file of type G-PCC coded point cloud has two types of information: geometric information and attribute information. The geometric information represents the 3D coordinates of the point in space, and the attribute information represents information such as color, reflectance, and normal direction attached to the point. Since the data at the points of a media file whose type is G-PCC coded point cloud is similar to the attributes that can be expressed by the syntax elements included in the primitive attributes of the mesh description module, when describing the data at the points of a media file whose type is G-PCC coded point cloud in the mesh description module (mesh), the syntax elements in the primitive attributes (mesh.primitives.attribute) of the mesh description module (mesh) can be multiplexed to describe the data at the points of a media file whose type is G-PCC coded point cloud.

[0092] For example, the value of the position syntax element (position, the first entry in Table 1 above) in the attribute of a primitive in the mesh description module is a three-dimensional vector of floating-point numbers, and such a data structure can similarly represent the geometric information of a G-PCC-coded point cloud. Therefore, the position syntax element (position) in the attribute of a primitive in the mesh description module (mesh.primitives.attribute) is multiplexed to represent the geometric information of a point in a media file whose type is a G-PCC-coded point cloud. Also, for example, the color value of a point in a media file whose type is a G-PCC-coded point cloud may be represented by multiplexing the color syntax element (color_n, the fifth entry in Table 1 above) in the attribute of a primitive in the mesh description module (mesh.primitives.attribute). Also, for example, the normal vector of a point in a media file of type G-PCC coded point cloud may be represented by multiplexing the normal vector syntax element (normal, the third entry in Table 1 above) in the primitive attribute (mesh.primitives.attribute) of the mesh description module.

[0093] The first syntax element set is defined as a set of syntax elements supported in the primitive attributes of the mesh description module of a scene description file specified in the ISO / IEC 23090-14 MPEG-I Scene Description Standard. A method for describing a 3D mesh that supports a media file whose type is a G-PCC coded point cloud includes adding syntax elements corresponding to each type of data the 3D mesh has to the primitive attributes of the mesh description module corresponding to the 3D mesh based on the syntax elements in the first syntax element set. As shown in Table 13 below, Table 13 shows a method for describing some data of points in a media file whose type is a G-PCC coded point cloud by multiplexing syntax elements in the primitive attribute (mesh.primitives.attribute) of the mesh description module. [Table 13]

[0094] The above embodiment and Table 13 show a method in which only a portion of the G-PCC coded point cloud data is described by multiplexing syntax elements in the primitive attributes of the mesh description module, and the G-PCC coded point cloud data may further include other data, and the other data of the G-PCC coded point cloud may be described by multiplexing syntax elements in the primitive attributes of the mesh description module, such as texture coordinates (texcoord_n), joints (joints_n), and weights (weights_n).

[0095] In another embodiment, a 3D mesh description method supporting a media file whose type is a G-PCC coded point cloud includes adding a target extension array to a primitive extension list (mesh.primitives.extensions) of a mesh description module, adding syntax elements corresponding to each type of data included in the 3D mesh in the media file whose type is a G-PCC coded point cloud to the target extension array, and describing data such as geometric information, color data, and normal vectors associated with each vertex of the 3D mesh in the media file whose type is a G-PCC coded point cloud using the syntax elements corresponding to each type of data.

[0096] In some embodiments, adding syntax elements corresponding to each type of data included in the 3D mesh corresponding to the target extension array includes adding syntax elements corresponding to each type of data included in the 3D mesh corresponding to the target extension array based on syntax elements in a first set of syntax elements, where the first set of syntax elements is a set of syntax elements supported in attributes of primitives in a mesh description module of a scene description file specified in the ISO / IEC 23090-14 MPEG-I Scene Description Standard.

[0097] In some embodiments, adding syntax elements corresponding to each type of data included in the three-dimensional mesh corresponding to the target extension array includes adding syntax elements corresponding to each type of data included in the three-dimensional mesh corresponding to the target extension array based on a second syntax element set consisting of syntax elements corresponding to a predetermined G-PCC encoding point group.

[0098] If a syntax element representing geometric information associated with each vertex is defined as a first syntax element, a syntax element representing color data associated with each vertex is defined as a second syntax element, and a syntax element representing normal vectors associated with each vertex is defined as a third syntax element, then the syntax elements to be added to the target extension array of the primitive extension list (mesh.primitives.extensions) of some mesh description modules include the following, as shown in Table 14 below: [Table 14]

[0099] As shown in Figure 9, Figure 9 is a structural schematic diagram of the scene description file after adding a target extension array to the primitive extension list (mesh.primitives.extensions) of the mesh description module based on the above embodiment and expanding the first syntax element, the second syntax element and the third syntax element in the target extension array. The scene description file includes, but is not limited to, MPEG media (MPEG_media) 901, a scene description module (scene) 902, a node description module (node) 903, a mesh description module (mesh) 904, an accessor description module (accessor) 905, a buffer slice description module (bufferView) 906, a buffer description module (buffer) 907, a skin description module (skin) 908, an animation description module (animation) 909, a camera description module (camera) 910, a material description module (material) 911, a texture description module (texture) 912, a sampler description module (sampler) 913, and a texture mapping description module (image) 914. Here, the extension list of primitive attributes of the mesh description module 904 includes a target extension array 9000, and the extended syntax elements in the target extension array 9000 include a first syntax element 9001 for representing geometric information associated with each vertex, a second syntax element 9002 for representing color data associated with each vertex, and a third syntax element 9003 for representing a normal vector associated with each vertex. In addition to the above extensions, information such as the role, accessor type, and data type of other elements in the scene description file shown in Figure 9 are similar to those in the scene description file shown in Figure 3, and therefore will not be described here.

[0100] In some embodiments, a mesh description method supporting a media file whose type is a G-PCC coded point cloud includes: pre-setting syntax elements corresponding to each type of data of the G-PCC coded point cloud; and adding, based on the pre-set syntax elements corresponding to each type of data of the G-PCC coded point cloud, the syntax elements corresponding to each type of data to attributes of a primitive in a mesh description module corresponding to a 3D mesh in the G-PCC coded point cloud.

[0101] Illustratively, the syntax elements corresponding to each type of data in the pre-arranged G-PCC coded point cloud include a fourth syntax element for representing geometric information associated with each vertex, a fifth syntax element for representing color data associated with each vertex, and a sixth syntax element for representing a normal vector associated with each vertex, and adding syntax elements corresponding to each type of data to attributes of a primitive in a mesh description module corresponding to a 3D mesh in the G-PCC coded point cloud based on the syntax elements corresponding to each type of data in the pre-arranged G-PCC coded point cloud includes adding at least one of the fourth syntax element, the fifth syntax element, and the sixth syntax element to attributes of a primitive in a mesh description module corresponding to a 3D mesh in the G-PCC coded point cloud.

[0102] If a syntax element for representing geometric information associated with each vertex corresponding to the G-PCC coded point group is defined as the fourth syntax element, a syntax element for representing color data associated with each vertex corresponding to the G-PCC coded point group is defined as the fifth syntax element, and a syntax element for representing normal vectors associated with each vertex corresponding to the G-PCC coded point group is defined as the sixth syntax element, as shown in Table 15 below, the description method of syntax elements in the attributes of primitives of some mesh description modules includes the following: [Table 15]

[0103] As shown in Figure 10, Figure 10 is a structural diagram of a scene description file after extending the syntax elements in the primitive attribute (mesh.primitives.attribute) of the mesh description module based on the above embodiment. The scene description file is composed of MPEG media (MPEG_media) 101, a scene description module (scene) 102, and a node description module (node) 103. 103 , a mesh description module (mesh) 104, an accessor description module (accessor) 105, a buffer slice description module (bufferView) 106, a buffer description module (buffer) 107, a skin description module (skin) 108, an animation description module (animation) 109, a camera description module (camera) 110, a material description module (material) 111, a texture description module (texture) 112, a sampler description module (sampler) 113, and a texture mapping description module (image) 114. Here, the primitive attribute (mesh.primitives.attribute) of the mesh description module 104 includes an extended fourth syntax element 1041 for representing geometric information associated with each vertex, a fifth syntax element 1042 for representing color data associated with each vertex, and a fifth syntax element 1043 for representing a normal vector associated with each vertex. In addition to the above extensions, information such as the roles, accessor types, and data types of other elements in the scene description file shown in FIG. 10 is similar to that in the scene description file shown in FIG. 3, and therefore a description thereof will be omitted here.

[0104] In addition, when a scene description file describes a 3D scene including a media file of type G-PCC coded point cloud, whether the G-PCC coded point cloud data is described by multiplexing syntax elements in the attributes of primitives in the mesh description module, adding a target extension array to a primitive in the mesh description module, or extending a new syntax element in the attributes of primitives in the mesh description module to describe the media file of type G-PCC coded point cloud, the mesh description module (mesh) will contain a large number of points in the G-PCC coded point cloud, and each point will contain at least geometric information and attribute information. Therefore, it is inconvenient to directly store the data of the media file of type G-PCC coded point cloud in the scene description framework. Therefore, the scene description framework will point out the link of the media file of type G-PCC coded point cloud, and when it is necessary to obtain the data of the G-PCC coded point cloud, the media file will be downloaded.

[0105] In some embodiments, the scene description file and the media file of type G-PCC coded point cloud may be merged to form one binary file, reducing the number and type of files.

[0106] In some embodiments, the description methods of an accessor description module (accessor), a buffer slice description module (bufferView), and a buffer description module (buffer) that support media files whose type is a G-PCC coded point group include an index value declared by a media index syntax element (media) of an MPEG ring buffer (MPEG_buffer_circular) of the buffer description module (buffer) pointing to a media description module in MPEG media (MPEG_media) corresponding to a media file whose type is a G-PCC coded point group.

[0107] That is, a media file whose type is G-PCC coded point group needs to be specified in a buffer description module (buffer), but instead of directly adding the Uniform Resource Locator (URL) of the media file whose type is G-PCC coded point group to the buffer description module, the value of the media index syntax element (media) in the MPEG ring buffer (MPEG_buffer_circular) in the buffer description module (buffer) points to the media description module in MPEG media (MPEG_media) corresponding to the media file whose type is G-PCC coded point group.

[0108] For example, if the value of the uniform resource identifier syntax element (uri) in the options of a media description module corresponding to a media file whose type is G-PCC coded point group in the media list (media) of MPEG media (MPEG_media) is "http: / / www.example.com / G-PCCexample.mp4" and it is the first media description module in the MPEG media, the value of the media index syntax element (media) of the MPEG ring buffer (MPEG_buffer_circular) is set to "0" to index the link of the first media file in the MPEG media in the MPEG ring buffer in the buffer description module, and the media index syntax element (media) in the MPEG ring buffer (MPEG_buffer_circular.media) of the buffer description module (buffer) is used to index the media description module corresponding to the media file whose type is G-PCC coded point group in MPEG media (MPEG_media).

[0109] In some embodiments, an accessor, buffer slice, or buffer description method supporting a media file of type G-PCC coded point group includes track information of data buffered in the buffer by the value of the second track index syntax element (tracks) of the track array (tracks) of the MPEG ring buffer (MPEG_buffer_circular) of the buffer description module (buffer).

[0110] In the scene description technology proposed by MPEG on top of glTF2.0, an extension called MPEG ring buffer (MPEG_buffer_circular) has been proposed. MPEG ring buffers are used to reduce the number of buffers used while still guaranteeing data buffer capacity. MPEG ring buffers can be considered as connecting the head and tail of regular buffers to form a ring. Writing buffers to the ring buffer and reading data from the ring buffer are achieved by simultaneous writing and reading using write and read pointers. The syntax elements included in MPEG ring buffer (MPEG_buffer_circular) are shown in Table 16. [Table 16]

[0111] That is, based on the setting rule for the value of the syntax element “media” in Table 16, if the value of the media index syntax element (media) in Table 16 is set to the index value of the media description module corresponding to a media file whose type declared in MPEG media (MPEG_media) is G-PCC coded point group, a media file whose type is G-PCC coded point group can be indexed in the buffer description module (buffer); and based on the setting rule for the value of the track index syntax element (tracks) in Table 16, if the value of the track index syntax element (tracks) in Table 16 is set to the index value of one or more code stream tracks of a media file whose type is G-PCC coded point group, the decoded data of the one or more code stream tracks can be buffered in the corresponding buffer.

[0112] In some embodiments, a material, texture, sampler, and texture mapping (image) description method that supports media files of type G-PCC coded point cloud includes not describing the three-dimensional scene using materials, textures, samplers, and texture mapping (image) when a scene description file is used to describe the three-dimensional scene in G-PCC coded point cloud.

[0113] Since G-PCC coded point clouds have a scattered topology structure, they do not actually have the concept of a surface, and various additional information is displayed directly on the points. Material, texture, sampler, and image are all additional information for surfaces, so the definitions of material, texture, sampler, and image are reserved, but 3D scenes are not described using material, texture, sampler, and image.

[0114] In some embodiments, a camera description module (camera) description method supporting a media file of type G-PCC coded point cloud includes defining viewing-related visual information such as viewpoint, viewing angle, etc. of a node in a three-dimensional scene using the camera description module.

[0115] In some embodiments, an animation description method supporting a media file of type G-PCC coded point cloud includes animations added by the animation description module to node description modules in a three-dimensional scene.

[0116] In some implementations, the animation description module can describe the animation to be added to the node description module (node) by one or more of position translation, angle rotation, and size scaling.

[0117] In some embodiments, the animation description module may indicate at least one of the start time, end time, and animation implementation method of an animation added to a node description module (node).

[0118] That is, in a scene description file supporting a media file whose type is a G-PCC coded point cloud, animation may be added to a node representing an object in a 3D object. The animation may describe the animation added to the node in three ways: position movement, angle rotation, and size enlargement / reduction, and may also specify the start and end times of the animation and the implementation method of the animation.

[0119] In some embodiments, a method for describing a skin description module (skin) that supports a media file of type G-PCC coded point cloud includes defining a motion and distortion relationship between a mesh (mesh) in a node description module (node) and a corresponding skeleton by the skin description module (skin).

[0120] Based on the above embodiment, the MPEG_media, scene description module (scene), node description module (node), mesh description module (mesh), accessor description module (accessor), buffer slice description module (bufferView), buffer description module (buffer), skin description module (skin), animation description module (animation), camera description module (camera), material description module (material), texture description module (texture), sampler description module (sampler), and texture mapping description module (image) in the scene description file have been improved and extended, so that the scene description file can accurately describe media files whose type is G-PCC coded point cloud.

[0121] For illustrative purposes, the following describes a scene description file supporting a media file whose type is G-PCC coded point cloud according to an embodiment of the present application, with reference to one specific scene description file.

[0122] JPEG2025529756000089.jpg152166JPEG2025529756000090.jpg148164JPEG2025529756000091.jpg148170JPEG2025529756000092.jpg150164

[0123] In the above example, the pair of square brackets on lines 1 and 118 contain the main contents of a scene description file supporting a media file whose type is a G-PCC coded point cloud, which includes a digital asset description module (asset), a used extension description module (extension Used), MPEG media (MPEG_media), a scene statement (scene), a scene list (scenes), a node list (nodes), a mesh list (meshes), an access list (accessors), a buffer slice list (buffer Views), and a buffer list (buffers). The contents of each part and the information contained in each list in terms of analysis angle are explained below.

[0124] Digital asset description module (asset): The digital asset description module is on lines 2-4. The "version":"2.0" on line 3 of the digital asset description module determines that the scene description file is created based on the glTF2.0 version, which is also the reference version of the scene description standard. From the perspective of analysis, the display engine can determine which parser to select to parse the scene description file according to the digital asset description module.

[0125] Used extension description module (extension Used): The used extension description module is from line 6 to line 10. The used extension description module contains three syntax elements: MPEG media (MPEG_media), MPEG ring buffer (MPEG_buffer_circular), and MPEG time-varying accessor (MPEG_accessor_timed). This determines that the scene description file uses three MPEG extensions: MPEG media, MPEG ring buffer, and MPEG time-varying accessor. From the perspective of analysis, the display engine can know in advance based on the contents of the used extension description module that the extension items to be analyzed subsequently include MPEG media, MPEG ring buffer, and MPEG time-varying accessor.

[0126] MPEG Media (MPEG_media): MPEG media is from lines 12 to 34. The MPEG media implements the statement for the media file whose type is G-PCC coded point cloud included in the 3D scene. The media type syntax element and its value "mimeType": "application / mp4" on line 21 indicate the encapsulation format of the media file containing the media file whose type is G-PCC coded point cloud. The "uri": "http: / / www.exp.com / G-PCCexp.mp4" on line 22 indicates the access address of the media file whose type is G-PCC coded point cloud. The "track": "trackIndex=1" on line 25 indicates the track information of the media file whose type is G-PCC coded point cloud. The "codecs": "gpc1" on line 26 indicates the codec parameters of the media file whose type is G-PCC coded point cloud. The "name": "G-PCCexample" on line 16 indicates the name of the media file whose type is G-PCC coded point cloud. The "autoplay": "true" indicates that media files of type G-PCC coded point cloud should be automatically played, and "loop": true" on line 18 indicates that files of type G-PCC coded point cloud should be cyclically played. From an analysis perspective, the display engine can analyze the MPEG media to determine that media files of type G-PCC coded point cloud exist in the 3D scene to be rendered, and obtain a method for accessing and analyzing the media files of type G-PCC coded point cloud.

[0127] Scene statement (scene): The scene statement is on line 36. Since one scene description file can theoretically contain multiple 3D scenes, in the above scene description file, the scene statement on line 36 and its "scene": 0 indicate that the 3D scene to be subsequently processed and rendered based on the scene description file is the first 3D scene in the scene list, i.e., the 3D scene enclosed in square brackets on lines 39-43.

[0128] Scene List (scenes): The scene list is located on lines 38-44. The scene list contains only one bracket, which means that the scene list contains only one scene description module, and the scene description file contains only one 3D scene. Within the brackets, "nodes":[0] on lines 40-42 indicates that the 3D scene contains only one node, and the index value of the node description module corresponding to the node is 0. From an analytical perspective, the contents of the scene list make it clear that the entire scene description framework should select the first 3D scene in the scene list (the 3D scene with index 0) for subsequent processing and rendering, clarifying the overall structure of the 3D scene and pointing to the more detailed node description module (node) in the next layer.

[0129] Node List (nodes): The node list is located in lines 46-51. The node list contains only one bracket, which means that the node list contains only one node description module, and the 3D scene contains only one node. This node and the node in the scene description module whose node description module has an index value of 0 are the same node, and are related by index. In the brackets representing the node, "name": "G-PCCexample_node" on line 48 indicates that the node's name is "G-PCC example_node," and "mesh": 0 on line 49 indicates that the content mounted on the node is a 3D mesh corresponding to the first mesh description module in the mesh list, which corresponds to the mesh description module in the next layer. From an analytical perspective, the contents of this node list indicate that the content mounted on the node is a 3D mesh, and that the 3D mesh is a 3D mesh corresponding to the first mesh description module in the mesh list.

[0130] Mesh list (meshes): The mesh list is from line 53 to line 66. The fact that the mesh list contains only one square bracket indicates that the mesh list contains only one mesh description module, that the 3D scene has only one 3D mesh, and that this 3D mesh and the 3D mesh with an index value of 0 in the node description module are the same 3D mesh. In the square brackets (mesh description module) describing this 3D mesh, "name":"G-PCC example_mesh" on line 55 indicates that the name of this 3D mesh is "G-PCC example_mesh", and this name is used only as an identification mark. "primitives" on line 56 indicates that this 3D mesh has primitives. "attributes" on line 58 and "mode" on line 62 respectively indicate that the primitives Attribute toThe "position" on line 59 and the "color_0" on line 60 indicate that the 3D mesh has geometric coordinates and color data. The "position":0 on line 59 and the "color_0":1 on line 60 indicate that the accessor corresponding to the geometric coordinates is the accessor corresponding to the first accessor description module in the accessory list, and the accessor corresponding to the color data is the accessor corresponding to the second accessor description module in the accessory list. Also, the "mode":0 on line 62 can determine that the topology of the 3D mesh is a scattered structure. From an analytical perspective, the mesh list clarifies the actual data type and topology type of the 3D mesh in the scene description file.

[0131] Buffer List (buffers): The buffer list is on lines 106-117. The fact that the buffer list contains only one bracket indicates that the scene description file contains only one buffer description module and that the display of the 3D scene requires access to only one media file. The brackets use an extension called MPEG ring buffer (MPEG_buffer_circular), indicating that this buffer is a ring buffer modified using MPEG extensions. The "media":0 on line 112 indicates that the data source in the ring buffer is the media file corresponding to the first media description module declared in the MPEG media file. The "tracks":#trackIndex=1 on line 113 indicates that the track with index 1 should be referenced when accessing the media file. The track with index 1 is not limited to this and may be the only track of a media file whose single-track encapsulation type is a G-PCC coded point cloud, or it may be a geometric codestream track of a media file whose multi-track encapsulation type is a G-PCC coded point cloud. Also, the syntax element "count": 5 in the MPEG ring buffer can be used to determine that the MPEG ring buffer has 5 storage segments, and the syntax element "byteLength": 15000 in the MPEG ring buffer can be used to determine that the byte length (capacity) of the MPEG ring buffer is 15000 bytes. From an analytical perspective, the buffer list allows media files whose declared type in the MPEG media is G-PCC coded point cloud to correspond to buffers, or allows buffers to reference previously declared but unused media files whose type is G-PCC coded point cloud.Note that the media file of type G-PCC coded point cloud cited here is an unprocessed G-PCC capsule file, and a G-PCC capsule file cannot extract information that can be directly used for rendering, such as the position coordinates (position) and color value (color_0) mentioned in the mesh description module, without undergoing processing by the media access function.

[0132] Buffer Slice List (Buffer Views): The buffer slice list is on lines 93-104. The buffer slice column contains two parallel brackets, and there is only one buffer defined by the buffer description module. This indicates that the buffer for storing a media file whose type is G-PCC coded point cloud is divided into two buffer slices, and that the point cloud data for the media file whose type is G-PCC coded point cloud is stored in the two buffer slices. In the first bracket (first buffer slice description module), buffer:0 on line 95 first points to the buffer description module with index 0, i.e., the only buffer description module mentioned in the buffer list. Next, the two parameters byte length and byte offset on lines 96 and 94 limit the data slice range of the corresponding buffer slice to the first 12,000 bytes. The second bracket (second buffer slice description module) is similar to the first bracket, but defines the data slice range as the last 3,000 bytes. From an analytical point of view, the buffer slice list groups the point cloud data in a media file of type G-PCC coded point cloud and contributes to the detailed definition of the subsequent accessor description modules.

[0133] Accessors: The accessories list is on lines 68-91. The accessories list has a similar structure to the buffer slice list, and each contains two parallel brackets, which indicates that the accessories list contains two accessor description modules and that media data must be accessed via two accessors to display the 3D scene. Additionally, both brackets (accessor description modules) have an extension called MPEG time-varying accessor (MPEG_accessor_timed), which explains that both accessors point to MPEG-defined time-varying media. In the first bracket, the contents of the MPEG time-varying accessor point to the buffer slice description module with an index value of 0. In the first accessor description module, "componentType":5126 on line 70 and "type":"VEC3" on line 71 further indicate that the data format stored in the accessor is a three-dimensional vector consisting of 32-bit floating-point numbers, and "count":1000 indicates that there are 1000 pieces of data that need to be accessed by the accessor in this format, and each 32-bit floating-point number occupies 4 bytes. Therefore, the accessor corresponding to the accessor description module contains 12000 bytes of data, which corresponds to the setting in the buffer slice description module with an index value of 0. The second accessor description module has similar content, changing the index value of the buffer slice description module to 1 and redefining the data type. From an analysis perspective, accessors improve the complete definition of data required for rendering. For example, data types missing in the buffer slice description module and buffer description module are defined in the corresponding accessor description module.

[0134] Display engines that support media files of type G-PCC coded point clouds

[0135] In the workflow of the immersive media scene description framework, the main functions of the display engine are to support the display engine function for media files whose type is G-PCC coded point clouds, which are similar to the main functions of the display engine in the workflow of the immersive media scene description framework described above: 1. to parse the scene description file for the media file whose type is G-PCC coded point clouds and obtain the corresponding 3D scene rendering method; 2. to transmit media access commands or media data processing commands to the media access function via the media access function API, where the media access commands or media data processing commands are obtained from the analysis result of the scene description file for the media file whose type is G-PCC coded point clouds; 3. to send buffer management commands to the buffer management module via the buffer API; 4. to obtain the processed G-PCC coded point cloud data from the buffer, and complete the rendering and display of the 3D scene and the objects in the 3D scene based on the read data. Note that the details of the processing process will not be described in detail here.

[0136] Media access function API that supports media files of type G-PCC coded point cloud

[0137] In the workflow of the immersive media scene description framework, the display engine can obtain a method for rendering a three-dimensional scene including a media file whose type is a G-PCC media file by analyzing the scene description file, and needs to transmit the method for rendering the three-dimensional scene to a media access function or send an instruction to the media access function based on the method for rendering the three-dimensional scene, and the process of transmitting the method for rendering the three-dimensional scene to the media access function or sending an instruction to the media access function based on the method for rendering the three-dimensional scene is realized by the media access function API.

[0138] In some embodiments, the display engine may send media access instructions or media data processing instructions to the media access function via the media access function API, where the media access instructions or media data processing instructions sent by the display engine to the media access function via the media access function API are derived from analysis results of a scene description file for a media file whose type is a G-PCC coded point cloud, and the media access instructions or media data processing instructions may include an index for the media file whose type is a G-PCC coded point cloud, a URL for the media file whose type is a G-PCC coded point cloud, attribute information for the media file whose type is a G-PCC coded point cloud, a display time window for the media file whose type is a G-PCC coded point cloud, a format request for the processed media file whose type is a G-PCC coded point cloud, etc.

[0139] In some embodiments, the media access function may request media access instructions or media data processing instructions from the display engine via the media access function API.

[0140] Media access functions that support media files of type G-PCC coded point cloud

[0141] In the workflow of the immersive media scene description framework, the media access function receives a media access instruction or a media data processing instruction sent by the display engine through the media access function API, and then executes the media access instruction or the media data processing instruction sent by the display engine through the media access function API, such as obtaining a media file whose type is a G-PCC coded point cloud, establishing an appropriate pipeline for the media file whose type is a G-PCC coded point cloud, and allocating an appropriate buffer for the processed media file whose type is a G-PCC coded point cloud.

[0142] In some embodiments, the media access function obtaining the media file of type G-PCC coded point cloud includes downloading the media file of type G-PCC coded point cloud from a server using a network transmission service.

[0143] In some embodiments, the media access function obtaining the media file of the G-PCC coded point cloud type includes reading the media file of the G-PCC coded point cloud type from a local storage space.

[0144] After receiving a media file whose type is a G-PCC-encoded point cloud, the media access function must process the media file whose type is a G-PCC-encoded point cloud. Because the processing process for different types of media files differs significantly, in order to support a wide range of media types and also take into account the efficiency of the media access function, various pipelines are designed in the media access function, and only the pipeline that matches the media type needs to be enabled during the media file processing process. If a media file is a G-PCC-encoded point cloud media file, the media access function must establish a corresponding pipeline for the G-PCC-encoded point cloud media file. The established pipeline must then perform processes such as decapsulation, G-PCC decoding, and post-processing on the G-PCC-encoded point cloud media file, completing the processing of the G-PCC-encoded point cloud media file and converting the G-PCC-encoded point cloud media file data into a data format suitable for direct rendering by a display engine.

[0145] 11, which is a structural schematic diagram of a pipeline corresponding to G-PCC coded point clouds in some embodiments of the present application. As shown in FIG. 11, a pipeline 1100 supporting media files whose type is G-PCC coded point clouds includes: an input module 111, a decapsulation module 112, a geometric decoder 113, an attribute decoder 114, a first post-processing module 115, and a second post-processing module 116.

[0146] The input module 111 receives a G-PCC capsule file and inputs the G-PCC capsule file to the decapsulation module 112. Here, the G-PCC capsule file is a file obtained by encapsulating a G-PCC code stream obtained by G-PCC encoding point cloud data. Since the G-PCC capsule file is expressed in the form of a track, what is received by the input module 111 is a track stream of the G-PCC capsule file. Furthermore, as can be seen from the capsule rules of the G-PCC code stream, the G-PCC capsule file may be a single track or a multi-track. Therefore, in the embodiment of the present application, the G-PCC capsule file received by the input module 111 may be a single track or a multi-track, and the embodiment of the present application is not limited thereto.

[0147] The decapsulation module 112 decapsulates the G-PCC capsule file input from the input module 111 to obtain a G-PCC code stream (including a geometric information code stream and an attribute information code stream), inputs the geometric information code stream to the geometric decoder 113, and inputs the attribute information code stream to the attribute decoder 114. It should be noted that, with the development of related technologies, code streams of other information may be added to the G-PCC code stream, and when the G-PCC code stream further includes code streams of other information, the decapsulation module 112 decapsulates the G-PCC capsule file to obtain the code streams of the other information, and inputs the code streams of the other information to corresponding decoders.

[0148] The geometric decoder 113 decodes the geometric information code stream output from the decapsulation module 112 to obtain the geometric information of the point cloud. Here, the main steps of the geometry decoder 113 decoding the geometric information code stream include obtaining the geometric information of the point cloud through arithmetic decoding, octree synthesis, surface fitting, geometric reconstruction, inverse coordinate transformation, etc. The specific implementation of the geometry decoder 113 decoding the geometric information code stream can refer to the workflow of the geometric decoding module 81 in Figure 8, and detailed description will be omitted here.

[0149] The attribute decoder 114 decodes the attribute information code stream input from the decapsulation module 112 to obtain the attribute information of the point cloud. Here, the main steps of the attribute decoder 114 decoding the geometric information code stream include attribute prediction, enhancement, and the inverse operation of the RAHT transform, etc., to obtain the attribute information code stream. For the specific implementation of the attribute decoder 114 decoding the attribute information code stream, please refer to the workflow of the attribute decoding module 82 in Figure 8, and detailed description will be omitted here.

[0150] The first post-processing module 115 processes the geometric information output from the geometry decoder 113. After completing the decoding of the geometric information codestream, the geometric information of the points in the G-PCC-encoded point cloud can be obtained. In some cases, the obtained geometric information can be directly used by a display engine. However, because the scene description framework does not impose excessive restrictions or special definitions on display engines, various types of display engines may appear. Since these different display engines may have different requirements for input data, adding the first post-processing module 115 after completing the decoding of the geometric information codestream ensures that the geometric information output from the pipeline can be used by any display engine. In some embodiments, the processing of the geometric information by the first post-processing module 115 includes performing format conversion on the geometric information.

[0151] The second post-processing module 116 is configured to process the attribute information output by the attribute decoder 114. After completing decoding of the attribute information codestream, attribute information of points in the G-PCC-encoded point cloud can be obtained, and in some cases, the attribute information can be directly used by a display engine. However, since the scene description frame does not impose excessive restrictions or special definitions on display engines, various types of display engines may appear. Since these different display engines may have different requirements for input data, adding the second post-processing module 116 after completing decoding of the attribute information codestream ensures that the attribute information output from the pipeline can be used by any display engine. In some embodiments, processing the geometric information by the first post-processing module 115 includes performing format conversion on the attribute information.

[0152] Finally, the processed geometric information output by the first post-processing module 115 and the processed attribute information output by the second post-processing module 116 are written to a buffer 117, whereby a display engine 118 reads the geometric information and attribute information from the buffer as needed, and renders and displays the G-PCC encoded point cloud in the 3D scene based on the read geometric information and attribute information.

[0153] Buffer API supporting media files of type G-PCC coded point cloud

[0154] After the media access function completes processing of the G-PCC encoded point cloud data through the pipeline, the media access function must transmit the processed data to the display engine in a standard array structure, which requires the processed G-PCC encoded point cloud data to be accurately stored in a buffer. This work is completed by the buffer management module, which must obtain buffer management commands from the media access function or the display engine via the buffer API.

[0155] In some embodiments, the media access function can send buffer management instructions to the buffer management module via a buffer API, where the buffer management instructions are buffer management instructions sent by the display engine to the media access function via the media access function API.

[0156] In some embodiments, the display engine can send buffer management instructions to the buffer management module via a buffer API.

[0157] That is, the buffer management module may communicate with the media access function via a buffer API, or may communicate with the display engine via the buffer API, and the purpose of communicating with the media access function or the display engine is to realize buffer management. When the buffer management module communicates with the media access function via the buffer API, the display engine must first send buffer management commands to the media access function via the media access function API, and the media access function must then send the buffer management commands to the buffer management module via the buffer API. When the buffer management module communicates with the display engine via the buffer API, the display engine may generate buffer management commands based on buffer management information parsed from the scene description file and send them to the buffer management module via the buffer API.

[0158] In some embodiments, the buffer management instructions may include one or more of a create buffer instruction, a update buffer instruction, and a release buffer instruction.

[0159] A buffer management module that supports media files of type G-PCC coded point cloud.

[0160] In the workflow of the immersive media scene description framework, after the media access function completes processing of the G-PCC encoded point cloud data through the pipeline, the processed G-PCC encoded point cloud data needs to be passed to the display engine in a standard array structure, which means that the processed G-PCC encoded point cloud data needs to be accurately stored in a buffer, and this work is handled by the buffer management module.

[0161] The buffer management module implements management operations such as creating, updating, and releasing buffers, and operation commands are received via the buffer API. Buffer management rules are recorded in a scene description document, analyzed by the display engine, and finally sent to the buffer management module by the display engine or media access function. After media files are processed by the media access function, they must be stored in appropriate buffers and used by the display engine. The role of buffer management is to manage these buffers to match the format of the processed media data without disrupting the processed media data. For specific design methods for the media management module, please refer to the designs of the display engine and media access function.

[0162] Based on the above, some embodiments of the present application provide a scene description file generating method, as shown in FIG. 12, the scene description file generating method includes the following steps S121-S123.

[0163] S121, determining the type of media file in the three-dimensional scene to be rendered;

[0164] In embodiments of the present application, the types of media files may include one or more of G-PCC coded point clouds, V-PCC coded point clouds, haptic media files, 6DoF videos, MIV videos, etc., and any number of media files of the same type may be included. For example, the 3D scene to be rendered may include only one media file of type G-PCC coded point clouds. Alternatively, for example, the 3D scene to be rendered may include one media file of type G-PCC coded point clouds and one media file of type V-PCC coded point clouds. Alternatively, for example, the 3D scene to be rendered may include two media files of type G-PCC coded point clouds and one haptic media file.

[0165] In the above step S121, if the type of the target media file in the 3D scene to be rendered is a G-PCC coded point group, the following step S122 is executed.

[0166] S122, according to the description information of the target media file, Target Media Description Module Generate.

[0167] In some embodiments, the descriptive information of the target media file includes one or more of the following: a name of the target media file; whether the target media file needs to be automatically played; whether the target media file needs to be played in a circular manner; an encapsulation format of the target media file; a codestream type of the target media file; encoding parameters of the target media file; etc.

[0168] In some embodiments, in step S122 (selecting a target media file corresponding to the target media file based on the description information of the target media file), Target Media Description Module ) includes at least one of the following steps 1221-1229:

[0169] In step 1221, add a media name syntax element (name) to the target media description module, and set the value of the media name syntax element based on the name of the target media file.

[0170] For example, if the media name syntax element in the target media description module is "name" and the name of the target media file is "G-PCC example", add the syntax element "name" to the target media description module and set the value of the syntax element "name" to "G-PCC example".

[0171] In step 1222, add an autoplay syntax element (autoplay) to the target media description module, and set the value of the autoplay syntax element according to whether the target media file needs to be autoplayed.

[0172] For example, if the autoplay syntax element in the target media description module is "autoplay" and the target media file needs to be autoplayed, then the syntax element "autoplay" is added to the target media description module and the value of the syntax element "autoplay" is set to "true".

[0173] Further, for example, if the autoplay syntax element in the target media description module is "autoplay" and the target media file does not need to be autoplayed, then the syntax element "autoplay" is added to the target media description module and the value of the syntax element "autoplay" is set to "false".

[0174] In step 1223, in the target media description module circulation regeneration Syntax element (loop) addition and sets the value of the cyclic playback syntax element according to whether the target media file needs to be cyclically played.

[0175] For example, if the autoplay syntax element in the target media description module is "loop" and the target media file needs to be played in a circular manner, add the syntax element "loop" to the target media description module and set the value of the syntax element "loop" to "true".

[0176] Further, for example, if the autoplay syntax element in the target media description module is "loop" and the target media file does not need to be played in a circular manner, add the syntax element "loop" to the target media description module and set the value of the syntax element "loop" to "false".

[0177] In step 1224, options (alternatives) are added to the target media description module.

[0178] In step 1225, a media type syntax element (mime Type) is added to the options (alternatives), and the value of the media type syntax element is set to the capsule format value corresponding to the G-PCC coded point group.

[0179] In some embodiments, the capsule format corresponding to the G-PCC encoded point group is MP4, and the capsule format value corresponding to the G-PCC encoded point group is application / mp4.

[0180] For example, if the media type syntax element is "mimeType" and the capsule format value corresponding to the G-PCC coding point group is "application / mp4", add the syntax element "mimeType" to the options of the target media description module and set the value of the syntax element "mimeType" to "application / mp4".

[0181] In step 1226, a uniform resource identifier syntax element (uri) is added to the options (alternatives), and the value of the uniform resource identifier syntax element is set to the access address of the target media file.

[0182] For example, if the uniform resource identifier syntax element is "uri" and the access address of the target media file is "http: / www.exp.com / G-PCCexp.mp4", add the syntax element "uri" to the options of the target media description module and set the value of the syntax element "uri" to http: / www.exp.com / G-PCCexp.mp4.

[0183] In step 1227, a track arrangement is added to the alternatives.

[0184] In step 1228, a first track index syntax element (track) is added to the track array (track) of the options (alternatives) of the target media description module, and the value of the first track index syntax element (track) is set according to the encapsulation method of the target media file.

[0185] In some embodiments, the step of setting the value of the first track index syntax element (track) based on the encapsulation format of the target media file includes setting the value of the first track index syntax element to an index value of a codestream track of the target media file if the target media file is a single-track capsule file, and setting the value of the first track index syntax element to an index value of a geometric codestream track of the target media file if the target media file is a multi-track capsule file.

[0186] That is, if the encoded G-PCC coded point cloud is referenced by the scene description file as one of MPEG_media.alternative.tracks and the referenced track satisfies the track specification in ISOBMFF, for single-track encapsulated G-PCC data, the track referenced in MPEG_media is the G-PCC code stream track. For example, if G-PCC data is encapsulated in one MIHS track by ISOBMFF, the track referenced in MPEG_media is this barcode stream track. For multi-track encapsulated G-PCC data, the track referenced in MPEG_media is the G-PCC geometric code stream track.

[0187] In the embodiment of the present application, the encapsulation method of the G-PCC coding point group includes single-track encapsulation and multi-track encapsulation, where single-track encapsulation refers to an encapsulation method in which the geometric code stream and attribute code stream of the G-PCC coding point group are encapsulated in the same code stream track, and multi-track encapsulation refers to an encapsulation method in which the geometric code stream and attribute code stream of the G-PCC coding point group are encapsulated in multiple code stream tracks, respectively.

[0188] In step 1229, a codec parameter syntax element (codecs) is added to the optional track array of the target media description module, and the value of the codec parameter syntax element is set based on the encoding parameters of the target media file, the type of the codestream of the target media file, and the ISO / IEC 23090-18G-PCC data transmission standard.

[0189] For example, the ISO / IEC 23090-18 G-PCC data transmission standard specifies that when a G-PCC coded point cloud employs DASH encapsulation, if G-PCC pre-selection signaling is used in an MPD file, the "codecs" attribute of the pre-selection signaling should be set to 'gpc1', indicating that the pre-selection media is a geometric point cloud. When multiple G-PCC Tile tracks exist in a G-PCC container, the "codecs" attribute of the Main G-PCC Adaptation Set should be set to 'gpcb' or 'gpeb', indicating that the adaptation set contains G-PCC Tile basic track data. If Tile Component Adaptation Sets signal only a single G-PCC component data, the "codecs" attribute of the Main G-PCC Adaptation Set should be set to 'gpcb'. If Tile Component Adaptation Sets signal all G-PCC component data, the "codecs" attribute of the Main G-PCC Adaptation Set should be set to 'gpeb'. When G-PCC Tile preselection signaling is used in an MPD file, the "codecs" attribute of the preselection signaling should be set to 'gpt1', indicating that the preselection media is a geometric point cloud fragment. When the G-PCC coded point cloud adopts a DASH capsule and G-PCC preselection signaling is used in an MPD file, the value of "codecs" in the "tracks" of the "alternatives" of the target media description module can be set to "gpc1".

[0190] For example, if the media files in the 3D scene to be rendered only include a target media file whose type is a G-PCC coded point cloud, the capsule format value corresponding to the G-PCC coded point cloud is "application / mp4", the name of the target media file is "G-PCC example", the target media file is auto-played and cyclically played, the access address of the target media file is http: / / www.exp.com / G-PCCexp.mp4, the target media file is a single-track capsule file, and the index value of the codestream track of the target media file is 1, the target media file adopts DASH capsules, and uses G-PCC preselection signaling in the MPD file, the target media description module corresponding to the target media file can be shown as follows: JPEG2025529756000093.jpg96164

[0191] In S123, the target media description module is added to the media list (media) of the MPEG media (MPEG_media) in the scene description file of the three-dimensional scene to be rendered.

[0192] Here, the target media description module is a media description module generated based on description information of the target media file.

[0193] For example, if the media files in the 3D scene to be rendered only include a target media file whose type is a G-PCC coded point cloud, the capsule format value corresponding to the G-PCC coded point cloud is application / mp4, the name of the target media file is "G-PCCexample1", the target media file is auto-played and cyclically played, the access address of the target media file is "uri":http: / www.exp.com / G-PCCexp.mp4, the target media file is a single-track capsule file, and the index value of the codestream track of the target media file is 1, the target media file is DASH encapsulated, and G-PCC preselection signaling is used in the MPD file, the MPEG media of the scene description file can be shown as follows: JPEG2025529756000094.jpg122165

[0194] In some embodiments, the three-dimensional scene to be rendered may further include a plurality of media files, and one or more of the plurality of media files is of the type of G-PCC coded point cloud. When generating the scene description file, it is necessary to add a media description module corresponding to the media file of the type of G-PCC coded point cloud based on the above embodiment, and add a media description module corresponding to other types of media files based on the scene description file generation method for other types of media files.

[0195] For example, if the media files in the 3D scene to be rendered include a target media file whose type is a G-PCC coded point cloud and one haptic media file, the capsule format value corresponding to the G-PCC coded point cloud is "application / mp4", the name of the target media file is "G-PCC example", the target media file is auto-playable and cyclically played, the access address of the target media file is "uri":http: / www.exp.com / G-PCCexp.mp4, the target media file is a single-track capsule file, and the index value of the codestream track of the target media file is 1, the target media file is DASH encapsulated, and G-PCC preselection signaling is used in the MPD file, the MPEG media of the scene description file can be shown as follows: JPEG2025529756000095.jpg111164JPEG2025529756000096.jpg92164

[0196] In the above example, the media list (media) for MPEG media contains two brackets, the first bracket (lines n+2-n+18) contains a media description module corresponding to a target media file of type G-PCC coded point cloud, and the second bracket (lines n+19-n+35) contains a media description module corresponding to a haptic media file.

[0197] In a method for generating a scene description file according to an embodiment of the present application, when generating a scene description file of a 3D scene to be rendered, first, a type of a media file in the 3D scene to be rendered is determined; if the type of the target media file in the 3D scene to be rendered is a G-PCC coded point group, a scene description file corresponding to the target media file is generated based on the description information of the target media file. Target Media Description Moduleand add the target media description module to a media list of MPEG media in a scene description file of the 3D scene to be rendered. In an embodiment of the present application, if the media files in the 3D scene to be rendered include a target media file of a G-PCC coded point cloud type, a target media description module corresponding to the target media file is generated based on description information of the target media file, and the target media description module is added to a media list of MPEG media in a scene description file of the 3D scene to be rendered. Target Media Description Module and add a media description module corresponding to the target media file to the media description module list of the MPEG media in the scene description file. Thus, the embodiment of the present application can generate a scene description file containing a 3D scene whose type is a G-PCC coded point cloud, and realize the support of media files whose type is a G-PCC coded point cloud by the scene description file.

[0198] In some embodiments, the method for generating a scene description file further comprises the following steps:

[0199] A target scene description module (scene) corresponding to the three-dimensional scene to be rendered is added to the scene list (scenes) of the scene description file, and the index value of the node description module corresponding to the node in the scene to be rendered is added to the node list (nodes) of the target scene description module.

[0200] For example, if the 3D scene to be rendered includes two nodes and the index values ​​of the node description modules (node) corresponding to the two nodes are 0 and 1, respectively, the target scene description module corresponding to the 3D scene to be rendered added to the scene description file can be shown as follows: JPEG2025529756000097.jpg59165

[0201] In the above example, the 3D scene to be rendered contains two nodes, and the index values ​​of the node description modules corresponding to the two nodes are 0 and 1, respectively, so the two index values ​​0 and 1 are added to the node list (nodes) of the scene description module corresponding to the 3D scene to be rendered.

[0202] In some embodiments, the method for generating a scene description file further comprises the following steps:

[0203] A node description module corresponding to a node in the scene to be rendered is added to the node list (nodes) of the scene description file, and the index value of a mesh description module corresponding to the 3D mesh mounted on the node is added to the mesh index list (mesh) of the node description module.

[0204] In some embodiments, the method for generating a scene description file further comprises the following steps:

[0205] A node name syntax element (name) is added to the node description module, and the value of the node name syntax element (name) in the corresponding node description module is set based on the name of the node.

[0206] For example, the three-dimensional scene to be rendered includes two nodes, the names of which are G-PCCexp_node1 and G-PCCexp_node2, respectively, the index values ​​of the mesh description modules corresponding to the three-dimensional meshes included in node G-PCCexp_node1 are 0 and 1, respectively, and the index value of the mesh description module corresponding to the three-dimensional mesh included in node G-PCCexp_node2 is 2, and the node list (nodes) portion of the scene description file can be shown as follows: JPEG2025529756000098.jpg65165

[0207] In the above example, the node list (nodes) of the scene description file corresponding to the 3D scene to be rendered includes two node description modules: the first node description module is the content enclosed in brackets on lines n+2-n+5, and the second node description module is the content enclosed in brackets on lines n+6-n+9. The value of the node name syntax element (name) in the first node description module is set to correspond to the name of the node "G-PCCexp_node1", the value of the mesh index syntax element (mesh) in the first node description module is set to correspond to the index values ​​0 and 1 of the mesh description module of the 3D mesh mounted on the node, the value of the node name syntax element (name) in the second node description module is set to correspond to the name of the node "G-PCCexp_node2", and the value of the mesh index syntax element (mesh) in the second node description module is set to correspond to the index value 2 of the mesh description module of the 3D mesh mounted on the node.

[0208] In some embodiments, the method for generating a scene description file further comprises the following steps:

[0209] A mesh description module (mesh) corresponding to a 3D mesh in the scene to be rendered is added to the mesh list (meshes) of the scene description file, syntax elements corresponding to each type of data included in the 3D mesh corresponding to the mesh description module are added to the mesh description module, and the values ​​of the syntax elements corresponding to each type of data are set to the index values ​​of accessor description modules corresponding to accessors for accessing each type of data.

[0210] In an embodiment of the present application, the data contained in the 3D mesh may include one or more of geometric coordinates (position), color values ​​(color), normal vectors (normal), tangent vectors (tangent), texture coordinates (texcoord), joints (joints), and weights (weights).

[0211] In some embodiments, adding syntax elements to the mesh description module corresponding to each type of data included in the 3D mesh corresponding to the mesh description module comprises the following steps:

[0212] An extension list is added to the primitives of a mesh description module corresponding to the 3D mesh in the target media file, a target extension array is added to the extension list, and syntax elements corresponding to each type of data contained in the 3D mesh corresponding to the target extension array are added.

[0213] In some embodiments, the target extension sequence may be MPEG_primitive_GPCC.

[0214] In some embodiments, adding syntax elements corresponding to each type of data included in the 3D mesh corresponding to the target extension array includes adding syntax elements corresponding to each type of data included in the 3D mesh corresponding to the target extension array based on syntax elements in a first set of syntax elements, where the first set of syntax elements is a set of syntax elements supported in attributes of primitives in a mesh description module of a scene description file specified in the ISO / IEC 23090-14 MPEG-I Scene Description Standard.

[0215] Specifically, the syntax elements supported by the attributes of primitives in the mesh description module of the scene description file defined in the ISO / IEC23090-14 MPEG-I scene description standard include position, color_n, normal, tangent, texcoord, joints, and weights, and therefore the first syntax element set is {position, color_n, normal, tangent, texcoord, joints, weights}.

[0216] For example, if a 3D mesh includes geometric coordinates and color data, and the index value of an accessor description module corresponding to an accessor for accessing the geometric coordinates is 0, and the index value of an accessor description module corresponding to an accessor for accessing the color data is 1, after adding syntax elements corresponding to each type of data included in the 3D mesh corresponding to the target extension array based on the first syntax element set, the mesh description module corresponding to the 3D mesh can be expressed as follows: JPEG2025529756000099.jpg75164

[0217] In some embodiments, adding syntax elements corresponding to each type of data included in the three-dimensional mesh corresponding to the target extension array includes adding syntax elements corresponding to each type of data included in the three-dimensional mesh corresponding to the target extension array based on a second syntax element set consisting of syntax elements corresponding to a predetermined G-PCC encoding point group.

[0218] Illustratively, syntax elements corresponding to the G-PCC coded point group may include G-PCC_position, G-PCC_color_n, G-PCC_normal, G-PCC_tangent, G-PCC_texcoord, G-PCC_joints, and G-PCC_weights, and correspondingly, the second syntax element set is {G-PCC_position, G-PCC_color_n, G-PCC_normal, G-PCC_tangent, G-PCC_texcoord, G-PCC_joints, G-PCC_weights}.

[0219] For example, if a 3D mesh includes geometric coordinates and color data, and the index value of an accessor description module corresponding to an accessor for accessing the geometric coordinates is 0, and the index value of an accessor description module corresponding to an accessor for accessing the color data is 1, after adding syntax elements corresponding to each type of data included in the 3D mesh corresponding to the target extension array based on the second syntax element set, the mesh description module corresponding to the 3D mesh can be expressed as follows: JPEG2025529756000100.jpg75164

[0220] In some embodiments, adding syntax elements to the mesh description module corresponding to each type of data included in the three-dimensional mesh corresponding to the mesh description module includes adding syntax elements corresponding to each type of data included in the three-dimensional mesh corresponding to the mesh description module to attributes of primitives in the mesh description module.

[0221] In some embodiments, adding syntax elements corresponding to each type of data included in the 3D mesh corresponding to the mesh description module to attributes of primitives in the mesh description module includes adding syntax elements corresponding to each type of data included in the 3D mesh corresponding to the mesh description module to attributes of primitives in the mesh description module based on the first set of syntax elements, where the first set of syntax elements is a set of syntax elements supported in attributes of primitives in the mesh description module of a scene description file defined in the ISO / IEC 23090-14 MPEG-I Scene Description Standard.

[0222] That is, for all 3D meshes in the scene description file (including 3D meshes in media files of type G-PCC and 3D meshes in media files of other types), add syntax elements to the attributes of the primitives of the corresponding mesh description module based on the syntax elements in the same syntax element set.

[0223] For example, a 3D mesh includes geometric coordinates and color data, and the index value of the accessor description module corresponding to the accessor for accessing the geometric coordinates is 1, and the index value of the accessor description module corresponding to the accessor for accessing the color data is 2. After adding syntax elements corresponding to each type of data included in the 3D mesh corresponding to the target extension array based on the first syntax element set, the mesh description module corresponding to the 3D mesh can be expressed as follows: JPEG2025529756000101.jpg76165

[0224] In some embodiments, adding syntax elements corresponding to each type of data included in the 3D mesh corresponding to the mesh description module to attributes of primitives of the mesh description module includes adding syntax elements corresponding to each type of data included in the corresponding 3D mesh to attributes of primitives of the first mesh description module based on syntax elements in a first set of syntax elements, and adding syntax elements corresponding to each type of data included in the corresponding 3D mesh to attributes of primitives of the second mesh description module based on syntax elements in a second set of syntax elements.

[0225] Here, the first mesh description module is a mesh description module corresponding to a three-dimensional mesh in a media file whose type is a G-PCC coded point cloud, and the second mesh description module is not a mesh description module corresponding to a three-dimensional mesh in a media file whose type is a G-PCC coded point cloud.

[0226] In some embodiments, the first set of syntax elements is a set formed by syntax elements supported in attributes of primitives of a mesh description module of a scene description file defined in the ISO / IEC 23090-14 MPEG-I Scene Description Standard, and the second set of syntax elements is a set formed by syntax elements corresponding to a pre-established G-PCC coded point group.

[0227] That is, when adding syntax elements corresponding to each type of data contained in the corresponding 3D mesh to the attributes of the primitive in the mesh description module, the 3D mesh in the scene description file is divided into two types depending on whether the 3D mesh belongs to the 3D mesh in the G-PCC type media file or not. For the 3D mesh in the media file whose type is not G-PCC coded point cloud, syntax elements corresponding to each type of data contained therein are added to the attributes of the primitive in the corresponding mesh description module based on the syntax elements in the first syntax element set, and for the 3D mesh in the media file whose type is G-PCC coded point cloud, syntax elements corresponding to each type of data contained therein are added to the attributes of the primitive in the corresponding mesh description module based on the syntax elements in the second syntax element set.

[0228] Illustratively, the scene description file includes two 3D meshes, named GPCCexample_mesh1 and GPCCexample_mesh2, respectively. Here, GPCCexample_mesh1 does not belong to the 3D mesh in the media file of type G-PCC, and includes geometric coordinates and color data, the index value of the accessor description module corresponding to the accessor for accessing the geometric coordinates of GPCCexample_mesh1 is 0, and the index value of the accessor description module corresponding to the accessor for accessing the color data of GPCCexample_mesh1 is 1; GPCCexample_mesh2 belongs to the 3D mesh in the media file of type G-PCC, and includes geometric coordinates and color data, the index value of the accessor description module corresponding to the accessor for accessing the geometric coordinates of GPCCexample_mesh2 is 2, and the index value of the accessor description module corresponding to the accessor for accessing the color data of GPCCexample_mesh2 is 3; after adding syntax elements corresponding to each type of data included in the 3D mesh corresponding to the target extension array based on the above embodiment, the mesh list (meshes) in the scene description file can be shown as follows: JPEG2025529756000102.jpg140164

[0229] In some embodiments, the method for generating a scene description file further comprises the following steps:

[0230] Based on the name of the 3D mesh, the value of the mesh name syntax element (name) in the mesh description module corresponding to the 3D mesh is set.

[0231] In some embodiments, the method for generating a scene description file further comprises the following steps:

[0232] Depending on the type of data included in the three-dimensional mesh, syntax elements included in the attributes of the primitives in the mesh description module corresponding to the three-dimensional mesh are set.

[0233] In some embodiments, the method for generating a scene description file further comprises the following steps:

[0234] Based on the type of the topology structure of the three-dimensional mesh, a value of a syntax element for describing the topology type of the three-dimensional mesh in a mesh description module corresponding to the three-dimensional mesh is set.

[0235] In some embodiments, the syntax element for describing the topology type of a 3D mesh in a mesh description module corresponding to the 3D mesh is "mode".

[0236] In some embodiments, the method for generating a scene description file further comprises the following steps:

[0237] Add an accessor description module (accessor) corresponding to a target accessor to a buffer list (accessor) of the scene description file, where the target accessor is an accessor for accessing decoded data of the target media file.

[0238] In some embodiments, the method for generating a scene description file further includes adding a buffer description module (buffer) corresponding to a target buffer to a buffer list (buffers) of the scene description file, where the target buffer is a buffer for storing decoded data of the target media file.

[0239] In some embodiments, adding a buffer description module (buffer) corresponding to a target buffer to a buffer list (buffers) of the scene description file includes at least one of the following steps a1-a5:

[0240] In step a1, a byte length syntax element (byteLength) is added to a buffer description module corresponding to the target buffer, and the value of the byte length syntax element is set to the byte length of the target media file.

[0241] For example, if the data amount of the G-PCC coded point group is 15000, the value of "byteLenth" in the buffer description module is set to "15000".

[0242] In step a2, an MPEG ring buffer (MPEG_buffer_circular) is added to the buffer description module corresponding to the target buffer.

[0243] In step a3, a segment count syntax element (count) is added to the MPEG ring buffer, and the value of the corresponding segment count syntax element (count) is set based on the number of stored segments in the target buffer.

[0244] For example, if the number of storage segments in the ring buffer is 8, the "count" in the ring buffer and its value are set to "count":8.

[0245] In step a4, a media index syntax element (media) is added to the MPEG ring buffer, and the value of the media index syntax element (media) is set according to the index value of the target media description module.

[0246] For example, if the index value of the target media description module is 0, "media" and its value in the ring buffer description module are set to "media":0.

[0247] In step a5, a second track index syntax element (tracks) is added to the MPEG ring buffer, and the value of the second track index syntax element (tracks) is set according to the track index value of the source data of the data stored in the target buffer.

[0248] For example, if the index value of the codestream track to which the data stored in the ring buffer belongs is 1, the "tracks" and its value in the description module of the ring buffer can be set to "tracks":"#trackIndex=1".

[0249] For example, adding a buffer description module corresponding to a target buffer to the buffer list of the scene description file includes any of the above steps a1-a5, and if the byte length of the target media file is 9000, the number of storage segments of a certain target buffer is 8, the index value of the media description module corresponding to the target media file is 1, and the track index value of the source data of the data stored in the MPEG ring buffer is 1, the buffer description module corresponding to the target buffer to be added to the buffer list of the scene description file can be shown as follows: JPEG2025529756000103.jpg55164

[0250] In some embodiments, the method for generating a scene description file further includes adding a buffer slice description module corresponding to a buffer slice of a target buffer to a buffer slice list (buffer Views) of the scene description file.

[0251] In some embodiments, adding a buffer slice description module corresponding to a buffer slice of the target buffer to a buffer slice list of the scene description file includes at least one of the following steps b1-b3:

[0252] In step b1, add a buffer index syntax element (buffer) to a buffer slice description module corresponding to a buffer slice of the target buffer, and set the value of the buffer index syntax element (buffer) based on the index value of the buffer description module corresponding to the target buffer to which the buffer slice belongs.

[0253] For example, if the index value of the buffer description module corresponding to a certain buffer is 2, then "buffer" and its value in the buffer slice description module are set to "buffer":2.

[0254] In step b2, a second byte length syntax element (byte Length) is added to a buffer slice description module corresponding to a buffer slice of the target buffer, and the value of the second byte length syntax element (byte Length) is set based on the capacity of the buffer slice.

[0255] In step b3, an offset syntax element (byte Offset) is added to a buffer slice description module corresponding to a buffer slice of the target buffer, and the value of the offset syntax element is set according to the offset of the stored data of the corresponding buffer slice.

[0256] For example, if the data range of a certain buffer slice of the buffer is [1, 12000], based on the above steps b2 and b3, "byteLenth" and its value in the buffer slice description module corresponding to the buffer slice are set to "byteLenth":12000, and "byteOffset" and its value in the buffer slice description module corresponding to the buffer slice are set to "byteOffset":0; if the data range of a certain buffer slice of the buffer is [12001, 15000], based on the above steps b2 and b3, "byteLenth" and its value in the buffer slice description module corresponding to the buffer slice are set to "byteLenth":3000, and "byteOffset" and its value in the buffer slice description module corresponding to the buffer slice are set to "byteOffset":12000.

[0257] For example, adding a buffer slice description module corresponding to a buffer slice of the target buffer to the buffer slice list (buffer Views) of the scene description file includes any of steps b1-b3. If the index value of the buffer description module corresponding to a target buffer is 1, the capacity of the target buffer is 8000, and the target buffer includes two buffer slices, the capacity of the first buffer slice is 6000 and the offset amount is 0, and the capacity of the second buffer slice is 2000 and the offset amount is 6001, adding a buffer slice description module corresponding to the buffer slice of the target buffer to the buffer slice list of the scene description file is as follows: JPEG2025529756000104.jpg63165

[0258] In some embodiments, the method for generating a scene description file further includes adding an accessor description module corresponding to a target accessor to an accessors list of the scene description file, where the target accessor is an accessor for accessing decoded data of the target media file.

[0259] In some embodiments, adding an accessor description module corresponding to a target accessor to the accessories list of the scene description file includes at least one of the following steps c1-c6:

[0260] In step c1, a data type syntax element (component Type) is added to the accessor description module corresponding to the target accessor, and the value of the corresponding data type syntax element is set according to the type of data accessed by the target accessor.

[0261] For example, if the type of data accessed by a certain accessor is 5126, the data type syntax element and its value in the accessor description module corresponding to that accessor are set to "componentType":5126.

[0262] In step c2, an accessor type syntax element (type) is added to the accessor description module corresponding to the target accessor, and the value of the accessor type syntax element is set based on a preset accessor type.

[0263] For example, if the accessor type accessed by a certain accessor is "VEC3", the accessor type syntax element (type) and its value in the accessor description module corresponding to that accessor are set to "type":"VEC3".

[0264] In step c3, a data count syntax element (count) is added to the accessor description module corresponding to the target accessor, and the value of the corresponding accessor type syntax element is set based on the type of the target accessor.

[0265] In step c4, an MPEG time-varying accessor (MPEG_accessor_timed) is added to the accessor description module corresponding to the target accessor.

[0266] In step c5, a buffer slice index syntax element (bufferView) is added to the MPEG time-varying accessor, and the value of the corresponding slice index syntax element is set based on the index value of the buffer slice description module corresponding to the buffer slice that stores the data accessed by the target accessor.

[0267] For example, if the index value of the buffer slice description module corresponding to the buffer slice to which the data accessed by a certain accessor belongs is 3, the buffer slice index syntax element and its value in the MPEG time-varying accessor of the accessor description module corresponding to the target accessor are set to "bufferView":3.

[0268] In step c6, a time-varying syntax element (immutable) is added to the MPEG time-varying accessor, and the value of the time-varying syntax element is set based on whether the value of the syntax element in the corresponding target accessor changes over time.

[0269] In some embodiments, if the value of a syntax element in a target accessor does not change over time, the time-varying syntax element and its value in the MPEG time-varying accessor of the accessor description module corresponding to the target accessor are set to "immutable":true, and if the value of a syntax element in a target accessor does change over time, the time-varying syntax element and its value in the MPEG time-varying accessor of the accessor description module corresponding to the target accessor are set to "immutable":false.

[0270] For example, adding an accessor description module corresponding to the target accessor for accessing data in a buffer slice of the target buffer to the accessories list of the scene description file includes any of steps c1-c6 above, and the type of data accessed by a certain target accessor is 5121, the accessor type of the target accessor is VEC2, the number of data accessed by the target accessor is 4000, the index value of the buffer slice description module corresponding to the buffer slice that stores the data that the target accessor needs to access is 1, and the value of the syntax element in the corresponding accessor does not change over time, the accessor description module corresponding to the target accessor added to the accessories list of the scene description file can be shown as follows: JPEG2025529756000105.jpg58164

[0271] In some embodiments, the method for generating a scene description file further comprises the following steps:

[0272] A digital asset description module (asset) is added to the scene description file, a version syntax element (version) is added to the digital asset description module, and if the scene description file creates a scene description statement based on the glTF2.0 version, the value of the version syntax element is set to 2.0.

[0273] Exemplarily, the digital asset description module added to the scene description file can be shown as follows: JPEG2025529756000106.jpg30164

[0274] In some embodiments, the method for generating a scene description file further comprises the following steps:

[0275] An extensions usage description module (extensionsUsed) is added to the scene description file, and extensions for the scene description file of MPEG glTF2.0 version used by the scene description file are added to the extensions usage description module.

[0276] For example, MPEG extensions used in the scene description file include MPEG media (MPEG_media), MPEG ring buffer (MPEG_buffer_circular), and MPEG time-varying accessor (MPEG_accessor_timed), and the extension usage description module added to the scene description file can be shown as follows: JPEG2025529756000107.jpg41164

[0277] In some embodiments, the method for generating a scene description file further comprises the following steps:

[0278] A scene statement (scene) is added to the scene description file, and the value of the scene statement is set to the index value of the scene description module corresponding to the scene to be rendered.

[0279] For example, if the index value of the scene description module corresponding to the scene to be rendered is 0, adding a scene statement to the scene description file can be expressed as follows: JPEG2025529756000108.jpg18164

[0280] Some embodiments of the present application further provide a scene description file parsing method, as shown in FIG. 13, the scene description file parsing method includes the following steps S131-S133.

[0281] In S131, a scene description file of the three-dimensional scene to be rendered is obtained.

[0282] Here, the 3D scene to be rendered includes a target media file of type G-PCC coded point cloud.

[0283] In the embodiments of the present application, the 3D scene to be rendered may include one or more media files, and when the 3D scene to be rendered includes multiple media files, the type of one or more of the media files may be G-PCC coded point cloud. When the 3D scene to be rendered includes multiple target media files of the type of G-PCC coded point cloud, the analysis method according to the embodiments of the present application may be performed on each of the target media files of the type of G-PCC coded point cloud.

[0284] In S132, a target media description module corresponding to the target media file is obtained from the media list (media) of the MPEG media (MPEG_media) of the scene description file.

[0285] Exemplarily, the target media description module corresponding to the target media file can be shown as follows: JPEG2025529756000109.jpg78164

[0286] In S133, description information of the target media file is obtained based on the target media description module.

[0287] In some embodiments, the above step S133 (obtaining description information of the target media file based on the target media description module) includes at least one of the following steps 1331-1337:

[0288] In step 1331, obtain the name of the target media file based on the value of the media name syntax element (name) in the target media description module.

[0289] For example, if the media name syntax element in the target media description module and its value are "name":"GPCC example", it can be determined that the name of the target media file is GPCC example.

[0290] In step 1332, it is determined whether the target media file needs to be autoplayed based on the value of an autoplay syntax element (autoplay) in the target media description module.

[0291] In some embodiments, determining whether the target media file needs to be autoplayed based on the value of an autoplay syntax element (autoplay) in the target media description module includes determining that the target media file needs to be autoplayed if the autoplay syntax element (autoplay) in the target media description module and its value is "autoplay":true, and determining that the target media file does not need to be autoplayed if the autoplay syntax element (autoplay) in the target media description module and its value is "autoplay":false.

[0292] In step 1333, it is determined whether the target media file needs to be played in a circular manner based on the value of a loop syntax element (loop) in the target media description module.

[0293] In some embodiments, determining whether the target media file needs to be played in a circular fashion based on the value of a circular playback syntax element (loop) in the target media description module includes determining that the target media file needs to be played in a circular fashion if the circular playback syntax element (loop) in the target media description module and its value is "loop":true, and determining that the target media file does not need to be played in a circular fashion if the circular playback syntax element (loop) in the target media description module and its value is "loop":false.

[0294] In step 1334, the encapsulation format of the target media file is obtained based on the value of a media type syntax element (mime Type) in the options (alternatives) of the target media description module.

[0295] If the type of a media file is a G-PCC coded point group, the value of the media type syntax element (mime Type) in the media description module corresponding to the media file is set to the capsule format value corresponding to the G-PCC coded point group, and the capsule format value corresponding to the G-PCC coded point group may be "application / mp4". Therefore, if the capsule format value corresponding to the G-PCC coded point group is "application / mp4", it can be obtained that the capsule format of the target media file is MP4.

[0296] In step 1335, obtain an access address of the target media file based on the value of a unique address identifier syntax element (uri) in the options (alternatives) of the target media description module.

[0297] For example, if there is only one address identifier syntax element (uri) in the options (alternatives) of the target media description module and its value is "uri": "http: / www.example.com / GPCCexample.mp4", it can be determined that the access address of the target media file is http: / www.example.com / GPCCexample.mp4.

[0298] In step 1336, track information of the target media file is obtained based on the value of a first track index syntax element (track) in a track array (track) of options (alternatives) of the target media description module.

[0299] In some embodiments, obtaining track information for the target media file based on the value of a first track index syntax element (track) in a track array (tracks) of options (alternatives) of the target media description module includes: determining the value of the first track index syntax element as an index value of a codestream track of the target media file if the capsule file of the target media file is a single-track capsule file; and determining the value of the first track index syntax element as an index value of a geometric codestream track of the target media file if the target media file is a multi-track capsule file.

[0300] In step 1337, the codestream type and decoding parameters of the target media file are determined based on the value of the codec parameter syntax element (codecs) in the track array (tracks) of the options (alternatives) of the target media description module and the ISO / IEC 23090-18G-PCC data transmission standard.

[0301] In some embodiments, the above step 1337 (determining the type and decoding parameters of the codestream of the target media file based on the value of the codec parameters syntax element (codecs) in the tracks array (tracks) of the options (alternatives) of the target media description module and the ISO / IEC 23090-18G-PCC data transmission standard) includes the following steps 13371 and 13372.

[0302] In step 13371, the type and encoding parameters of the codestream of the target media file are determined based on the value of the codec parameter syntax element (codecs) in the track array (tracks) of the options (alternatives) of the target media description module and the ISO / IEC 23090-18G-PCC data transmission standard.

[0303] As specified in the ISO / IEC 23090-18 G-PCC Data Transmission Standard, when a G-PCC coded point cloud employs DASH encapsulation, when using G-PCC pre-selection signaling in an MPD file, the codecs attribute of the pre-selection signaling should be set to 'gpc1' to indicate that the pre-selection media is based on a geometric point cloud. When multiple G-PCC Tile tracks exist in a G-PCC container, the "codecs" attribute of the Main G-PCC Adaptation Set should be set to 'gpcb' or 'gpeb' to indicate that the adaptation set contains G-PCC Tile basic track data. When Tile Component Adaptation Sets signal only a single G-PCC component data, the "codecs" attribute of the Main G-PCC Adaptation Set should be set to 'gpcb'. When Tile Component Adaptation Sets signal all G-PCC component data, the "codecs" attribute of the Main G-PCC Adaptation Set should be set to 'gpeb'. When G-PCC Tile preselection signaling is used in an MPD file, the "codecs" attribute of the preselection signaling should be set to 'gpt1', indicating that the preselection media is based on a geometric point cloud fragment. When a G-PCC coded point cloud employs DASH encapsulation and G-PCC preselection signaling is used in an MPD file, the value of "codecs" in "tracks" of "alternatives" of the target media description module can be set to 'gpc1'. Therefore, the encapsulation method and encoding parameters of the target media file can be determined based on the value of the codec parameter syntax element (codecs) in the track array (tracks) of the options (alternatives) of the target media description module and the ISO / IEC 23090-18 G-PCC data transmission standard.

[0304] In step 13372, decoding parameters of the target media file are determined based on the encoding parameters of the target media file.

[0305] Since the process of decoding the target media file and the process of encoding the target media file are inverse operations, the decoding parameters of the target media file can be determined based on the encoding parameters of the target media file.

[0306] Exemplarily, the target media description module corresponding to the target media file can be shown as follows: JPEG2025529756000110.jpg80164

[0307] In this case, the description information of the target media file obtained by the target media description module includes: the name of the target media file is AAAA; the target media file does not need to be automatically played but needs to be played in a circular manner; the encapsulation format of the target media file is MP4; the access address of the target media file is http: / / www.bbb.com / AAAA.mp4; the reference track of the target media file is a codestream track with an index value of 0; the encapsulation / decapsulation method of the target media file is MP4; and the codec parameters of the target media file are gpc1.

[0308] A scene description file analysis method according to an embodiment of the present application can obtain a scene description file of a 3D scene to be rendered, which includes a target media file whose type is a G-PCC coded point cloud, then obtain a target media description module corresponding to the target media file from a media list of MPEG media in the scene description file, and obtain descriptive information of the target media file based on the target media description module. The scene description file analysis method according to an embodiment of the present application can obtain the descriptive information of the target media file based on the target media description module, and further render and display a 3D scene to be rendered, which includes the target media file whose type is a G-PCC coded point cloud, based on the descriptive information of the target media file. Therefore, an embodiment of the present application provides a method for analyzing a scene description file of a 3D scene including a media file whose type is a G-PCC coded point cloud, and realizes analysis of a scene description file of a 3D scene including a G-PCC coded point cloud.

[0309] In some embodiments, the method for analyzing a scene description file according to the above embodiments further includes:

[0310] A target scene description module (scene) corresponding to the three-dimensional scene to be rendered is obtained from a scene list (scenes) of the scene description file, and description information of the three-dimensional scene to be rendered is obtained based on the target scene description module.

[0311] In some embodiments, a scene statement and an index value of the statement can be obtained from the scene description file, and a target scene description module corresponding to the 3D scene to be rendered can be obtained from the scene list of the scene description file based on the scene statement and the index value of the statement.

[0312] For example, if a scene statement and its index value are "scene":0, the first scene description module from the scene list of the scene description file can be obtained as the target scene description module corresponding to the 3D scene to be rendered based on the scene statement and its index value.

[0313] In some embodiments, obtaining description information of the 3D scene to be rendered based on the target scene description module includes determining an index value of a node description module corresponding to a node in the 3D scene to be rendered based on an index value declared in a node index list (nodes) of the target scene description module.

[0314] Exemplarily, the target scene description module is as follows: JPEG2025529756000111.jpg39165

[0315] In this case, based on the index values ​​stated in the node index list (nodes) of the target scene description module, it can be determined that the 3D scene to be rendered includes two nodes, and the index value of the node description module corresponding to one node is 0 (the first node description module in the node list), and the index value of the node description module corresponding to the other node is 1 (the second node description module in the node list).

[0316] In some embodiments, after determining the index value of the node description module corresponding to the node in the 3D scene to be rendered based on the index value declared in the node index list (nodes) of the target scene description module, the scene description file analysis method provided in the above embodiments includes:

[0317] Based on the index value of the node description module corresponding to the node in the three-dimensional scene to be rendered, the node description module corresponding to the node in the three-dimensional scene to be rendered is obtained from the node list (nodes) of the scene description file, and based on the node description module corresponding to the node in the three-dimensional scene to be rendered, description information of the node in the three-dimensional scene to be rendered is obtained.

[0318] For example, if the index value stated in the node index list of the target scene description module contains only 0, the first node description module from the node list of the scene description file is obtained as the node description module corresponding to the node in the 3D scene to be rendered.

[0319] Furthermore, for example, if the index values ​​stated in the node index list of the target scene description module include 0 and 1, the first node description module and the second node description module are obtained from the node list of the scene description file and are used as the node description modules corresponding to the nodes in the 3D scene to be rendered.

[0320] In some embodiments, obtaining description information of a node in the three-dimensional scene to be rendered based on a node description module corresponding to the node in the three-dimensional scene to be rendered includes at least one of the following steps a1 and a2:

[0321] In step a1, a name of a node in the three-dimensional scene to be rendered is obtained based on a value of a node name syntax element (name) in a node description module corresponding to the node in the three-dimensional scene to be rendered.

[0322] In step a2, an index value of a mesh description module corresponding to a 3D mesh mounted on a node in the 3D scene to be rendered is determined based on the index value stated in the mesh index list in the node description module corresponding to the node in the 3D scene to be rendered.

[0323] For example, the node description module corresponding to a node is shown below: JPEG2025529756000112.jpg35165

[0324] In this case, based on step a1 above, it can be determined that the name of the node is GPCC example_node, and based on step a2 above, it can be determined that the index values ​​of the mesh description modules corresponding to the 3D meshes mounted on the node are 0 and 1, respectively.

[0325] In some embodiments, after determining the index value of the mesh description module corresponding to the 3D mesh mounted on the node in the 3D scene to be rendered, the method for analyzing a scene description file according to the above embodiments further includes: obtaining a mesh description module corresponding to the 3D mesh mounted on the node in the 3D scene to be rendered from a mesh list (meshes) of the scene description file based on the index value of the mesh description module corresponding to the 3D mesh mounted on the node in the 3D scene to be rendered; and obtaining description information of the 3D mesh mounted on the node in the 3D scene to be rendered based on the mesh description module corresponding to the 3D mesh mounted on the node in the 3D scene to be rendered.

[0326] For example, if the index value declared in the mesh index list of a node description module contains only 0, the first mesh description module from the mesh list of the scene description file is obtained as the mesh description module corresponding to the 3D mesh mounted on the node corresponding to the node description module.

[0327] Furthermore, for example, if the index values ​​stated in the mesh index list of a node description module include 1 and 2, the second mesh description module and the third mesh description module are obtained from the mesh list of the scene description file as mesh description modules corresponding to the 3D mesh mounted on the node corresponding to the node description module.

[0328] In some embodiments, obtaining description information of a 3D mesh mounted on a node in the 3D scene to be rendered based on a mesh description module corresponding to the 3D mesh mounted on a node in the 3D scene to be rendered includes at least one of the following steps b1 to b4:

[0329] In step b1, the name of the 3D mesh is obtained based on the mesh name syntax element (name) in the mesh description module corresponding to the 3D mesh.

[0330] In step b2, the data type included in the three-dimensional mesh is obtained based on the data type syntax element in the mesh description module corresponding to the three-dimensional mesh.

[0331] In some embodiments, step b2 (obtaining the data type included in the 3D mesh based on the data type syntax element in the mesh description module corresponding to the 3D mesh) includes obtaining the data type included in the 3D mesh based on the data type syntax element in the target extension array of the extension list of primitives in the mesh description module corresponding to the 3D mesh.

[0332] In some embodiments, the target extension sequence may be MPEG_primitive_GPCC.

[0333] For example, the list of primitive extensions for a mesh description module corresponding to a 3D mesh is as follows: JPEG2025529756000113.jpg44164

[0334] In this case, it is possible to determine that the three-dimensional mesh includes a position coordinate based on the position coordinate syntax element (position) in the target extension array (MPEG_primitives_GPCC) of the extension list of primitives of the mesh description module corresponding to the three-dimensional mesh, to determine that the three-dimensional mesh includes a color value based on the color value syntax element (color_0) in the target extension array (MPEG_primitives_GPCC) of the extension list of primitives of the mesh description module corresponding to the three-dimensional mesh, and to determine that the three-dimensional mesh includes a normal vector based on the normal vector syntax element (normal) in the target extension array (MPEG_primitives_GPCC) of the extension list of primitives of the mesh description module corresponding to the three-dimensional mesh.

[0335] Further, for example, the extension list of primitives of a mesh description module corresponding to a certain 3D mesh is shown as follows: JPEG2025529756000114.jpg44164

[0336] In this case, it can be determined that the three-dimensional mesh includes a position coordinate based on the position coordinate syntax element (G-PCC_position) in the target extension array (MPEG_primitives_GPCC) of the extension list of primitives of the mesh description module corresponding to the three-dimensional mesh, that the three-dimensional mesh includes a color value based on the color value syntax element (G-PCC_color_0) in the target extension array (MPEG_primitives_GPCC) of the extension list of primitives of the mesh description module corresponding to the three-dimensional mesh, and that the three-dimensional mesh includes a normal vector based on the normal vector syntax element (G-PCC_normal) in the target extension array (MPEG_primitives_GPCC) of the extension list of primitives of the mesh description module corresponding to the three-dimensional mesh.

[0337] In some embodiments, step b2 (obtaining the data type included in the 3D mesh based on a data type syntax element in a mesh description module corresponding to the 3D mesh) includes obtaining the data type included in the 3D mesh based on a data type syntax element in attributes of primitives in the mesh description module corresponding to the 3D mesh.

[0338] For example, the attributes of the primitives in the mesh description module corresponding to a 3D mesh are as follows: JPEG2025529756000115.jpg39165

[0339] In this case, it is determined that the 3D mesh includes position coordinates based on the position coordinate syntax element (position) in the attributes of the primitives of the mesh description module corresponding to the 3D mesh, it is determined that the 3D mesh includes color values ​​based on the color value syntax element (color_0) in the attributes of the primitives of the mesh description module corresponding to the 3D mesh, and it is determined that the 3D mesh includes normal vectors based on the normal vector syntax element (normal) in the attributes of the primitives of the mesh description module corresponding to the 3D mesh.

[0340] Further, for example, the attributes of the primitives in the mesh description module corresponding to a certain 3D mesh are as follows: JPEG2025529756000116.jpg40164

[0341] In this case, it can be determined that the three-dimensional mesh includes position coordinates based on the position coordinate syntax element (G-PCC_position) in the attributes of the primitives of the mesh description module corresponding to the three-dimensional mesh, that the three-dimensional mesh includes color values ​​based on the color value syntax element (G-PCC_color_0) in the attributes of the primitives of the mesh description module corresponding to the three-dimensional mesh, and that the three-dimensional mesh includes normal vectors based on the normal vector syntax element (G-PCC_normal) in the attributes of the primitives of the mesh description module corresponding to the three-dimensional mesh.

[0342] In step b3, an index value of an accessor description module corresponding to an accessor for accessing data of the 3D mesh type is obtained based on the value of the data type syntax element.

[0343] As described in the above example, since the value of the position coordinate syntax element (G-PCC_position) is 0, the index value of the accessor description module corresponding to the accessor for accessing the position coordinates of the 3D mesh is determined to be 0 (the first accessor in the accessory list); since the value of the color value syntax element (G-PCC_color_0) is 1, the index value of the accessor description module corresponding to the accessor for accessing the color value of the 3D mesh is determined to be 1 (the second accessor in the accessory list); and since the value of the normal vector syntax element (G-PCC_normal) is 2, the index value of the accessor description module corresponding to the accessor for accessing the normal vector of the 3D mesh is determined to be 2 (the third accessor in the accessory list).

[0344] In step b4, the type of the topology structure of the three-dimensional mesh is obtained based on the value of the mode syntax element (mode) in the mesh description module corresponding to the three-dimensional mesh.

[0345] Illustratively, if the value of the mode syntax element is 0, the type of the topological structure of the 3D mesh can be determined to be a scattered point; if the value of the mode syntax element is 1, the type of the topological structure of the 3D mesh can be determined to be a line; and if the value of the mode syntax element is 4, the type of the topological structure of the 3D mesh can be determined to be a triangle.

[0346] By way of example, a mesh description module corresponding to a 3D mesh is as follows: JPEG2025529756000117.jpg78164

[0347] In this case, the description information of the three-dimensional mesh obtained based on the mesh description module corresponding to the three-dimensional mesh includes: the name of the three-dimensional mesh is G-PCC example_mesh; the topology type of the three-dimensional mesh is scattered point; the three-dimensional mesh includes three types of data, which are position coordinates, color values, and normal vectors, respectively; the index value of the accessor description module corresponding to the accessor for accessing the position coordinates of the three-dimensional mesh is 0; the index value of the accessor description module corresponding to the accessor for accessing the color values ​​of the three-dimensional mesh is 1; and the index value of the accessor description module corresponding to the accessor for accessing the normal vector of the three-dimensional mesh is 2.

[0348] In some embodiments, after obtaining the index value of the accessor description module corresponding to the accessor for accessing the data of the 3D mesh type based on the value of the data type syntax element, the method further includes:

[0349] Based on the index value of the accessor description module corresponding to the accessor for accessing the data of each type of three-dimensional mesh, an accessor description module corresponding to the accessor for accessing the data of each type of three-dimensional mesh is obtained from the accessory list of the scene description file, and based on the accessor description module corresponding to the accessor for accessing the data of each type of three-dimensional mesh, description information of the accessor for accessing the data of each type of three-dimensional mesh is obtained.

[0350] For example, if the index value of the accessor description module corresponding to the accessor for accessing the color value of the three-dimensional mesh is 1, a second accessor description module is obtained from the accessory list of the scene description file as the accessor description module corresponding to the accessor for accessing the color value of the three-dimensional mesh.

[0351] In some embodiments, obtaining description information of an accessor for accessing each type of data of the three-dimensional mesh based on an accessor description module corresponding to the accessor for accessing each type of data of the three-dimensional mesh includes at least one of the following steps c1 to c6:

[0352] In step c1, the type of data accessed by the accessor is determined based on the value of the data type syntax element (component Type) in the accessor description module.

[0353] For example, if the data type syntax element in an accessor description module corresponding to an accessor for accessing the normal vector of a 3D mesh is "componentType":5126, it can be determined that the type of the data accessed by the accessor corresponding to the accessor description module (the normal vector of the 3D mesh) is a 32-bit floating-point number (float).

[0354] In step c2, the type of the accessor is determined based on the value of the accessor type syntax element (type) in the accessor description module.

[0355] For example, if the accessor type syntax element in an accessor description module corresponding to an accessor for accessing the position coordinates of a certain 3D mesh is "type":VEC3, it can be determined that the type of the accessor corresponding to the accessor description module is a 3D vector.

[0356] In step c3, the quantity of data accessed by the accessor is determined based on the value of the data quantity syntax element (count) in the accessor description module.

[0357] For example, if the data count syntax element in an accessor description module corresponding to an accessor for accessing the color values ​​of a certain three-dimensional mesh is "count":1000, it can be determined that the number of data (color values ​​of the three-dimensional mesh) accessed by the accessor corresponding to the accessor description module is 1000.

[0358] In step c4, it is determined whether the accessor is a time-varying accessor based on MPEG extension modification based on whether the accessor description module includes an MPEG time-varying accessor (MPEG_accessor_timed).

[0359] In some embodiments, determining whether the accessor is a time-varying accessor based on MPEG extended adaptation based on whether the accessor description module includes an MPEG time-varying accessor includes determining that the accessor is a time-varying accessor based on MPEG extended adaptation if the accessor description module includes an MPEG time-varying accessor, and determining that the accessor is not a time-varying accessor based on MPEG extended adaptation if the accessor description module does not include an MPEG time-varying accessor.

[0360] In step c5, based on the value of the buffer slice index syntax element (buffer View) in the MPEG time-varying accessor (MPEG_accessor_timed) of the accessor description module, determine the index value of the buffer slice description module corresponding to the buffer slice that stores the data accessed by the accessor.

[0361] For example, if the buffer slice index syntax element in the MPEG time-varying accessor of an accessor description module corresponding to an accessor for accessing the normal vector of a certain 3D mesh is "bufferView":0, it can be determined that the data accessed by the accessor corresponding to the accessor description module (the normal vector of the 3D mesh) is stored in the buffer slice corresponding to the first buffer slice description module in the buffer slice list.

[0362] In step c6, it is determined whether the value of the syntax element in the accessor changes over time based on the value of the time-varying syntax element (immutable) in the MPEG time-varying accessor of the accessor description module.

[0363] In some embodiments, determining whether the value of the syntax element in the accessor changes over time based on the value of a time-varying syntax element (immutable) in the MPEG time-varying accessor of the accessor description module includes determining that the value of the syntax element in the accessor does not change over time if the time-varying syntax element in the MPEG time-varying accessor of the accessor description module and its value is "immutable":true, and determining that the value of the syntax element in the accessor changes over time if the time-varying syntax element in the MPEG time-varying accessor of the accessor description module and its value is "immutable":false.

[0364] For example, the accessor description module corresponding to a certain accessor is as follows: JPEG2025529756000118.jpg53164

[0365] In this case, the description information of the accessor obtained based on the accessor description module corresponding to the accessor includes: the type of data accessed by the accessor is 5123; the type of the accessor is SCALAR; the number of data accessed by the accessor is 1000; the accessor is a time-varying accessor based on MPEG extended modifications; the data accessed by the accessor is buffered in the buffer slice corresponding to the second buffer slice description module in the buffer slice list; and the value of the syntax element in the accessor does not change over time.

[0366] In some embodiments, the scene description file analysis method according to the above embodiments further includes the following steps d to g.

[0367] In step d, a buffer description module in the buffer list (buffers) of the scene description file is obtained.

[0368] In step e, the value of the media index syntax element (media) in the buffer description module is obtained.

[0369] In step f, a buffer description module whose value of the media index syntax element is the same as the index value of the target media description module is determined as a target buffer description module corresponding to a target buffer for buffering the decoded data of the target media file.

[0370] Exemplarily, if the index value of the target media description module is 0, the buffer description module whose media index syntax element value is 0 is determined as the target buffer description module corresponding to the target buffer for buffering the decoded data of the target media file.

[0371] It should be noted that the number of target buffers for buffering the decoded data of the target media file may be one or more, and the embodiment of the present application is not limited thereto.

[0372] In step g, description information of the target buffer is obtained based on the target buffer description module.

[0373] In some embodiments, obtaining description information of the target buffer based on the target buffer description module includes at least one of the following steps g1 to g4:

[0374] In step g1, the capacity of the target buffer is obtained based on the value of a first byte length syntax element (byte Length) in the target buffer description module.

[0375] For example, if the first byte length syntax element in the target buffer description module and its value is "byteLength":15000, it can be determined that the capacity of the target buffer is 15000 bytes.

[0376] In step g2, it is determined whether the target buffer is a ring buffer based on MPEG extension modification according to whether the target buffer description module includes an MPEG ring buffer (MPEG_buffer_circular).

[0377] In some embodiments, determining whether the target buffer is a ring buffer based on MPEG extension modification based on whether the target buffer description module includes an MPEG ring buffer includes determining that the target buffer is a ring buffer based on MPEG extension modification if the target buffer description module includes an MPEG ring buffer, and determining that the target buffer is not a ring buffer based on MPEG extension modification if the target buffer description module does not include an MPEG ring buffer.

[0378] In step g3, the number of storage segments in the MPEG ring buffer is obtained based on the value of the segment count syntax element (count) in the MPEG ring buffer of the target buffer description module.

[0379] For example, if the segment count syntax element in the MPEG ring buffer of the target buffer description module and its value is "count":8, it can be determined that the MPEG ring buffer includes 5 storage segments.

[0380] In step g4, the track index value of the source data of the data buffered in the MPEG ring buffer is obtained according to the value of the second track index syntax element (tracks) in the MPEG ring buffer of the target buffer description module.

[0381] Exemplarily, the buffer description module corresponding to a certain buffer is as follows: JPEG2025529756000119.jpg53164

[0382] In this case, obtaining description information of a buffer based on a buffer description module corresponding to the buffer includes: the capacity of the buffer is 8000 bytes; the buffer is a ring buffer based on MPEG extension modification; the number of storage segments of the ring buffer is 5; the media file stored in the ring buffer is the second media file declared in the MPEG media; and the track index value of the source data of the data buffered in the ring buffer is 1.

[0383] In some embodiments, the scene description file analysis method according to the above embodiments further includes the following steps h to k.

[0384] In step h, a buffer slice description module in the buffer slice list (buffer Views) of the scene description file is obtained.

[0385] In step i, the value of the buffer index syntax element (buffer) in the buffer slice description module is obtained.

[0386] In step j, a buffer slice description module whose value of the buffer index syntax element is the same as the index value of the target buffer description module is determined as a buffer slice description module corresponding to a buffer slice of the target buffer.

[0387] Exemplarily, when the index value of the target media description module is 1, the buffer slice description module whose buffer index syntax element has a value of 1 is determined as the buffer slice description module corresponding to the buffer slice of the target buffer.

[0388] It should be noted that the number of buffer slices in the target buffer may be one or more, and the embodiment of the present application is not limited thereto.

[0389] In step k, description information of the buffer slice of the target buffer is obtained according to a buffer slice description module corresponding to the buffer slice of the target buffer.

[0390] In some embodiments, obtaining description information of a buffer slice of the target buffer based on a buffer slice description module corresponding to the buffer slice of the target buffer includes at least one of the following steps k1 and k2:

[0391] In step k1, the capacity of the buffer slice of the target buffer is obtained based on the value of a second byte length syntax element (byte Length) in a buffer slice description module corresponding to the buffer slice of the target buffer.

[0392] For example, if the second byte length syntax element in the buffer slice description module corresponding to a buffer slice of the target buffer and its value is "byteLength":12000, the capacity of the buffer slice of the target buffer can be determined to be 12000 bytes.

[0393] In step k2, an offset of the buffer slice of the target buffer is obtained based on the value of an offset amount syntax element (byte offset) in a buffer slice description module corresponding to the buffer slice of the target buffer.

[0394] For example, if the offset syntax element in the buffer slice description module corresponding to a buffer slice of the target buffer and its value is "byteOffset":0, it may be determined that the offset of the buffer slice of the target buffer is 0 bytes.

[0395] Exemplarily, the buffer slice description module corresponding to a certain buffer slice is as follows: JPEG2025529756000120.jpg39165

[0396] In this case, obtaining description information of the buffer slice based on the buffer slice description module corresponding to the buffer slice includes: the buffer slice is a buffer slice of a buffer corresponding to the second buffer description module in the buffer list, the capacity of the buffer slice is 8000 bytes, and the offset amount of the buffer slice is 0, i.e., the data range buffered in the buffer slice is the previous 8000 bytes.

[0397] In some embodiments, the scene description file analysis method according to the above embodiments further includes the following steps l-o:

[0398] In step l, an accessor description module in the accessor list of the scene description file is obtained.

[0399] In step m, the value of the buffer slice index syntax element (buffer View) in the accessor description module is obtained.

[0400] In step n, an accessor description module whose value of the buffer slice index syntax element is the same as the index value of the buffer slice description module corresponding to the buffer slice of the target buffer is determined as the accessor description module corresponding to the accessor for accessing data in the buffer slice of the target buffer.

[0401] For example, if the index value of a buffer slice description module corresponding to a buffer slice of the target buffer is 2, the accessor description module whose buffer slice index syntax element has a value of 2 is determined as the accessor description module corresponding to the accessor for accessing data in that buffer slice of the target buffer.

[0402] In step o, description information of an accessor for accessing data in a buffer slice of the target buffer is obtained based on an accessor description module corresponding to the accessor for accessing data in a buffer slice of the target buffer.

[0403] In some embodiments, obtaining description information of an accessor for accessing data in a buffer slice of the target buffer based on an accessor description module corresponding to the accessor for accessing data in a buffer slice of the target buffer includes at least one of the following steps o1 to o6:

[0404] In step o1, the type of data accessed by the accessor is determined based on the value of the data type syntax element (component Type) in the accessor description module.

[0405] In step o2, the type of the accessor is determined based on the value of the accessor type syntax element (type) in the accessor description module.

[0406] In step o3, the quantity of data accessed by the accessor is determined based on the value of the data quantity syntax element (count) in the accessor description module.

[0407] In step o4, it is determined whether the accessor is a time-varying accessor based on the MPEG extension modification based on whether the accessor description module includes an MPEG time-varying accessor (MPEG_accessor_timed).

[0408] In step o5, based on the value of the buffer slice index syntax element (buffer View) in the MPEG time-varying accessor of the accessor description module, determine the index value of the buffer slice description module corresponding to the buffer slice that stores the data accessed by the accessor.

[0409] In step o6, it is determined whether the value of the syntax element in the accessor changes over time based on the value of the time-varying syntax element (immutable) in the MPEG time-varying accessor of the accessor description module.

[0410] The implementation of the above steps o1 to o6 can refer to the implementation of the above steps c1 to c6, and detailed explanations will be omitted here to avoid duplication.

[0411] Some embodiments of the present application further provide a method for rendering a three-dimensional scene, where the execution entity of the three-dimensional scene rendering method is a display engine in an immersive media description framework, and as shown in FIG. 14 , the three-dimensional scene rendering method includes the following steps:

[0412] In S141, a scene description file of the three-dimensional scene to be rendered is obtained.

[0413] Here, the 3D scene to be rendered includes a target media file of type G-PCC coded point cloud.

[0414] In some embodiments, an implementation method for obtaining a scene description file of a 3D scene to be rendered includes sending request information to a media resource server to request a scene description file of the 3D scene to be rendered, and receiving a request response sent from the media resource server, the request response including the scene description file of the 3D scene to be rendered.

[0415] In S142, description information of the target media file is obtained based on the media description module corresponding to the target media file in the media list (media) of the MPEG media (MPEG_media) of the scene description file.

[0416] In some embodiments, the descriptive information of the target media file includes one or more of the following: a name of the target media file; whether the target media file needs to be automatically played; whether the target media file needs to be played in a circular manner; an encapsulation format of the target media file; a codestream type of the target media file; and encoding parameters of the target media file.

[0417] The implementation method for obtaining the description information of the target media file based on the media description module corresponding to the target media file can refer to the implementation method of the media description module for analyzing the target media file in the above-mentioned scene description file analysis method, and detailed description will be omitted here to avoid duplication.

[0418] In S143, the description information of the target media file is sent to the media access function.

[0419] After the display engine sends the description information of the target media file to the media access function, the media access function can obtain the target media file based on the description information of the target media file, process the target media file to obtain decoded data of the target media file, and write the decoded data of the target media file to a target buffer.

[0420] In some embodiments, the display engine sending the descriptive information of the target media file to the media access function includes the display engine sending the descriptive information of the target media file to the media access function via a media access function API.

[0421] In some embodiments, the display engine sending the descriptive information of the target media file to the media access function includes the display engine sending the media access function media file processing instructions that include the descriptive information of the target media file.

[0422] At S144, the decoded data of the target media file is read from the target buffer.

[0423] That is, it reads data from the target buffer that has been fully processed by the media access functions and can be used directly to render the 3D scene to be rendered.

[0424] In S145, the to-be-rendered 3D scene is rendered based on the decoded data of the target media file.

[0425] In a 3D scene rendering method according to an embodiment of the present application, after obtaining a scene description file of a 3D scene to be rendered, which includes a target media file whose type is a G-PCC coded point cloud, the method first obtains description information of the target media file based on a media description module corresponding to the target media file in a media list of MPEG media in the scene description file, and sends the description information of the target media file to a media access function, which then obtains the target media file based on the description information of the target media file, processes the target media file to obtain decoded data of the target media file, writes the decoded data of the target media file to a target buffer, and further reads the decoded data of the target media file from the target buffer, and renders the 3D scene to be rendered based on the decoded data of the target media file. In the 3D scene rendering method according to the embodiment of the present application, the display engine obtains the description information of the target media file based on the target media description module, sends the description information of the target media file to a media access function, reads the decoded data of the target media file whose type is a G-PCC coded point cloud, and renders the 3D scene to be rendered based on the decoded data of the target media file. Therefore, the embodiment of the present application provides a method for rendering a 3D scene to be rendered that includes a media file whose type is a G-PCC coded point cloud, and realizes rendering the media file whose type is a G-PCC coded point cloud based on a scene description file.

[0426] Some embodiments of the present application further provide a media file processing method, where the execution body of the media file processing method is a media access function in an immersive media description framework. Referring to FIG. 15 , the media file processing method includes the following steps:

[0427] In S151, description information of a target media file, description information of a target buffer, and description information of a buffer slice of the target buffer sent by a display engine are received.

[0428] Here, the target media file is a media file whose type is G-PCC coded point group, and the target buffer is a buffer for buffering decoded data of the target media file.

[0429] In some embodiments, the description information of the target media file may include at least one of the name of the target media file, whether the target media file needs to be automatically played, whether the target media file needs to be played in a circular manner, the capsule format of the target media file, the type of codestream of the target media file, and the encoding parameters of the target media file.

[0430] In some embodiments, the description information of the target buffer may include at least one of the following: the capacity of the buffer, whether it is an MPEG ring buffer, the number of storage segments in the ring buffer, an index value of a media description module corresponding to the target media file, and a track index value of source data of the data buffered in the ring buffer.

[0431] In some embodiments, the description information of the buffer slice of the target buffer may include at least one of the buffer to which the buffer slice belongs, the capacity of the buffer slice, and the offset amount of the buffer slice.

[0432] In some embodiments, receiving the descriptive information of the target media file, the descriptive information of the target buffer, and the descriptive information of the buffer slices of the target buffer sent from the display engine includes receiving, by a media access function API, the descriptive information of the target media file, the descriptive information of the target buffer, and the descriptive information of the buffer slices of the target buffer sent by the display engine.

[0433] In S152, the decoded data of the target media file is obtained based on the description information of the target media file.

[0434] In some embodiments, the media access function obtaining decoded data for the target media file based on the description information of the target media file includes creating a target pipeline for processing the target media file based on the description information of the target media file, obtaining the target media file by the target pipeline, and decapsulating and decodes the target media file to obtain the decoded data for the target media file.

[0435] In some embodiments, obtaining the target media file through the target pipeline and decapsulating and decoding the target media file to obtain decoded data for the target media file includes obtaining the target media file through an input module of the target pipeline and inputting the target media file to a decapsulation module of the target pipeline, decoding the target media file through the decapsulation module to obtain a geometry code stream and an attribute code stream for the target media file, decoding the geometry code stream through a geometry decoder of the target pipeline to obtain geometry decoded data for the target media file, and decoding the attribute code stream through an attribute decoder of the target pipeline to obtain attribute decoded data for the target media file.

[0436] In some embodiments, obtaining the target media file through the target pipeline and decapsulating and decoding the target media file to obtain decoded data for the target media file further includes: after obtaining geometric decoded data for the target media file, processing the geometric decoded data by a first post-processing module of the target pipeline; and after obtaining attribute decoded data for the target media file, processing the attribute decoded data by a second post-processing module of the target pipeline.

[0437] Illustratively, processing the geometry-decoded data by a first post-processing module of the target pipeline may include performing a format conversion on the geometry-decoded data by a first post-processing module of the target pipeline, and processing the attribute-decoded data by a second post-processing module of the target pipeline may include performing a format conversion on the attribute-decoded data by a second post-processing module of the target pipeline.

[0438] In S153, the decoded data of the target media file is written into the target buffer according to the description information of the target buffer and the description information of the buffer slice of the target buffer.

[0439] After writing the decoded data of the target media file to the target buffer, the display engine can read the decoded data of the target media file from the target buffer based on the description information of the target buffer and the description information of the buffer slices of the target buffer, and render a three-dimensional scene to be rendered including the target media file based on the decoded data of the target media file.

[0440] A media file processing method according to an embodiment of the present application receives, from a display engine, description information of a target media file whose type is a G-PCC coded point cloud, description information of a target buffer for buffering decoded data of the target media file, and description information of a buffer slice of the target buffer. The method obtains decoded data corresponding to the target media file based on the description information of the target media file, and writes the decoded data of the target media file to the target buffer based on the description information of the target buffer and the description information of the buffer slices of the target buffer. Thus, the display engine can read the decoded data of the target media file from the target buffer based on the description information of the target buffer and the description information of the buffer slices of the target buffer, and render a to-be-rendered 3D scene including the target media file based on the decoded data of the target media file. Thus, the embodiment of the present application can support rendering of media files whose type is a G-PCC coded point cloud in a scene description framework.

[0441] Some embodiments of the present application further provide a buffer management method, where the execution body of the buffer management method is a buffer management module in an immersive media description framework. Referring to FIG. 16 , the buffer management method includes the following steps:

[0442] In S161, description information of a target buffer and description information of a buffer slice of the target buffer are received.

[0443] Here, the target buffer is a buffer for buffering a target media file, and the target media file is a media file whose type is a G-PCC coded point cloud.

[0444] In some embodiments, the description information of the target buffer may include at least one of the capacity of the buffer, whether it is an MPEG ring buffer, the number of storage segments of the ring buffer, an index value of a media description module corresponding to the media file buffered in the ring buffer (the target media file), and a track index value of the source data of the data buffered in the ring buffer.

[0445] In some embodiments, the description information of the buffer slice of the target buffer may include at least one of the buffer to which the buffer slice belongs, the capacity of the buffer slice, and the offset amount of the buffer slice.

[0446] In S162, the target buffer is created based on the description information of the target buffer.

[0447] For example, the description information of the target buffer includes: the capacity of the target buffer is 8000 bytes; the target buffer is a ring buffer based on MPEG extended modification; the number of storage segments of the ring buffer is 3; the media file stored in the ring buffer is the first media file described in the MPEG media; and the track index value of the source data of the data buffered in the ring buffer is 1; the buffer management module creates a ring buffer with a capacity of 8000 bytes and including 3 storage segments as the target buffer.

[0448] In S163, the target buffer is divided into buffer slices based on the description information of the buffer slices of the target buffer.

[0449] As described in the above embodiment, if the ring buffer includes two buffer slices, and the description information of the first buffer slice includes a capacity of 6000 bytes and an offset of 0, and the description information of the second buffer slice includes a capacity of 2000 bytes and an offset of 6001, the target buffer is divided into two buffer slices, the capacity of the first buffer slice is 6000 bytes and buffers the first 6000 bytes of the decoded data of the target media file, and the capacity of the second buffer slice is 2000 bytes and buffers the 6001-8000 bytes of the decoded data of the target media file.

[0450] After the buffer management module divides the target buffer into buffer slices based on the description information of the buffer slices of the target buffer, a media access function writes the decoded data of the target media file to the target buffer, and a display engine reads the decoded data of the target media file from the target buffer and renders a three-dimensional scene to be rendered including the target media file based on the decoded data of the target media file.

[0451] A buffer management method according to an embodiment of the present application may receive description information of a target buffer and description information of buffer slices of the target buffer, and then create the target buffer based on the description information of the target buffer and divide buffer slices for the target buffer based on the description information of the buffer slices of the target buffer. Thus, a media access function may write decoded data of the target media file, which is decoded data of a media file whose type is a G-PCC coded point cloud, to the target buffer. A display engine may read the decoded data of the target media file from the target buffer and render a to-be-rendered 3D scene including the target media file based on the decoded data of the target media file. Thus, an embodiment of the present application may support rendering of media files whose type is a G-PCC coded point cloud in a scene description framework.

[0452] Some embodiments of the present application further provide a three-dimensional scene rendering method, including a scene description file parsing method and a three-dimensional scene rendering method performed by a display engine, a media file processing method performed by a media access function, and a buffer management method performed by a buffer management module. As shown in Figure 17, the method includes the following steps:

[0453] In S1701, the display engine obtains a scene description file of the three-dimensional scene to be rendered.

[0454] Here, the 3D scene to be rendered includes a target media file of type G-PCC coded point cloud.

[0455] In some embodiments, the display engine obtaining a scene description file for the scene to be rendered includes the display engine downloading the scene description file from a server using a network transmission service.

[0456] In some embodiments, the display engine obtaining a scene description file for the scene to be rendered comprises reading the scene description file from a local storage space.

[0457] In S1702, the display engine obtains a media description module corresponding to each media file from the media list (media) of the MPEG media (MPEG_media) of the scene description file (including obtaining a media description module corresponding to the target media file from the media list of the MPEG media of the scene description file).

[0458] In S1703, the display engine obtains description information for each media file based on a media description module corresponding to each media file (including obtaining description information for the target media file based on a media description module corresponding to the target media file).

[0459] In some embodiments, the description information of the media file includes at least one of the name of the media file, whether the media file is auto-played, whether the media file is cyclically played, the capsule format of the media file, the access address of the media file, track information of the capsule file of the media file, and codec parameters of the media file.

[0460] The implementation method by which the display engine obtains the description information of the target media file based on the media description module corresponding to the target media file can refer to the implementation method of the media description module that analyzes the target media file in the above-mentioned scene description analysis method, and detailed description will be omitted here to avoid duplication.

[0461] In S1704, the display engine sends description information of each media file to a media access function (including sending description information of the target media file to the media access function).

[0462] In response, the media access function receives description information for each media file sent from the display engine (including receiving description information for the target media file sent from the display engine).

[0463] In some embodiments, the display engine sending the descriptive information for each media file to the media access function includes the display engine sending the descriptive information for each media file to the media access function via a media access function API.

[0464] In some embodiments, the media access function receiving descriptive information for each media file sent from the display engine includes the media access function receiving descriptive information for each media file sent from the display engine via a media access function API.

[0465] In S1705, the media access function creates a corresponding pipeline for processing each media file based on the description information of each media file (including creating a target pipeline for processing the target media file based on the description information of the target media file).

[0466] In some embodiments, the target pipeline includes an input module, a decapsulation module, and a decoding module, wherein the input module is configured to obtain the target media file (capsule file), and the decapsulation module is configured to decapsulate the target media file to obtain a code stream of the target media file (which may be a single-track encapsulated G-PCC code stream, or may be a multi-track encapsulated G-PCC geometric code stream and G-PCC attribute code stream), and the decoding module includes a decoder, a geometry decoder, and an attribute decoder, wherein if the code stream of the target media file is a single-track encapsulated G-PCC code stream, the decoding module decodes the G-PCC code stream through a decoder to obtain decoded data of the target media file, and if the code stream of the target media file is a multi-track encapsulated G-PCC geometric code stream and G-PCC attribute code stream, the decoding module decodes the G-PCC geometric code stream and G-PCC attribute code stream through a geometry decoder and an attribute decoder, respectively, to obtain geometry data and attribute data of the target media file, and obtains the decoded data of the target media file.

[0467] In some embodiments, the target pipeline further includes a first post-processing module and a second post-processing module, wherein the first post-processing module performs post-processing such as format conversion on geometric data obtained by decoding a G-PCC geometric codestream, and the second post-processing module performs post-processing such as format conversion on attribute data obtained by decoding a G-PCC attribute codestream.

[0468] In S1706, the media access function obtains each media file through a pipeline process corresponding to each media file, decapsulates and decodes each media file, and obtains decoded data corresponding to each media file (including obtaining the target media file through the target pipeline, and decapsulating and decodes the target media file to obtain decoded data corresponding to the target media file).

[0469] In some embodiments, the description information of the target media file includes an access address of the target media file, and the media access function obtaining the decrypted data of the target media file based on the description information of the target media file includes the media access function obtaining the target media file based on the access address of the target media file.

[0470] In some embodiments, the media access function obtaining the target media file based on the access address of the target media file includes the media access function sending a media resource request to a media resource server based on the access address of the target media file, and receiving a media resource response including the target media file sent from the media server.

[0471] In some embodiments, the media access function obtaining the target media file based on the access address of the target media file includes the media access function reading the target media file from a predetermined storage space based on the access address of the target media file.

[0472] In some embodiments, the description information of the target media file further includes an index value of each code stream track of the target media file, and the media access function obtaining the decoded data of the target media file based on the description information of the target media file includes the media access function decapsulating the target media file based on an encapsulation format of the target media file and obtaining the code stream of each code stream track of the target media file.

[0473] In some embodiments, the description information of the target media file further includes a codestream type and codec parameters of the target media file, and the media access function obtaining the decoded data of the target media file based on the description information of the target media file includes the media access function decoding the codestream of each codestream track of the target media file based on the codestream type and codec parameters of the target media file to obtain the decoded data of the target media file.

[0474] In S1707, the display engine obtains each buffer description module in the buffer list (buffers) of the scene description file (including obtaining, from the buffer list of the scene description file, a buffer description module corresponding to the target buffer used to buffer the decoded data of the target media file).

[0475] In S1708, the display engine obtains description information for each buffer based on a buffer description module corresponding to each buffer (including obtaining description information for the target buffer based on a buffer description module corresponding to the target buffer).

[0476] In some embodiments, the description information of the buffer may include at least one of the capacity (byte length) of the buffer, the access address of the data buffered in the buffer, whether it is an MPEG ring buffer, the number of storage segments in the ring buffer, the index value of the media description module corresponding to the media file buffered in the ring buffer, and the track index value of the source data of the data buffered in the ring buffer.

[0477] In S1709, the display engine obtains each buffer slice description module in the buffer slice list (buffer Views) of the scene description file (including obtaining a buffer slice description module corresponding to the buffer slice description of the target buffer from the buffer slice list of the scene description file).

[0478] In S1710, the display engine obtains description information for the buffer slices of each buffer based on a buffer slice description module corresponding to the buffer slices of each buffer (including obtaining description information for the buffer slices of the target buffer based on a buffer slice description module corresponding to the buffer slices of the target buffer).

[0479] In some embodiments, the buffer description information may include at least one of the following: the buffer to which the buffer slice belongs, the capacity of the buffer slice, and the offset amount of the buffer slice.

[0480] In S1711, the display engine obtains each accessor description module in the accessories list of the scene description file (including obtaining an accessor description module corresponding to a target accessor for accessing the decoded data of the target media file from the accessories list of the scene description file).

[0481] In S1712, the display engine obtains description information for each accessor based on the accessor description module corresponding to each accessor (including obtaining description information for the target accessor for accessing the decoded data of the target media file based on the accessor description module corresponding to the target accessor).

[0482] In some embodiments, the descriptive information of the accessor may include at least one of: a buffer slice accessed by the accessor; a data type of data accessed by the accessor; a type of accessor; a number of data accessed by the accessor; whether it is an MPEG time-varying accessor; a buffer slice accessed by the time-varying accessor; and whether the accessor parameters change over time.

[0483] In some embodiments, after the above steps S1707-S1712, the embodiments of the present application may use the following method 1 to send description information of each buffer, description information of buffer slices of each buffer, and description information of each access unit to the media access function and the buffer management module.

[0484] In some embodiments, the implementation of method 1 (sending description information of each buffer, description information of buffer slices of each buffer, and description information of each accessor to the media access function and the buffer management module) includes the following steps a and b:

[0485] In step a, the display engine sends description information of each buffer, description information of a buffer slice of each buffer, and description information of each accessor to a media access function (including the display engine sending description information of the target buffer, description information of a buffer slice of the target buffer, and description information of the target accessor to a media access function).

[0486] Correspondingly, the media access function receives description information of each buffer and description information of buffer slices of each buffer sent by the display engine (the media access function includes receiving description information of the target buffer, description information of buffer slices of the target buffer, and description information of the accessor sent by the display engine).

[0487] In some embodiments, the implementation of step a above (the display engine sending description information of each buffer, description information of buffer slices of each buffer, and description information of each accessor to the media access function) may be such that the display engine sends description information of each buffer, description information of buffer slices of each buffer, and description information of each accessor to the media access function via the media access function API.

[0488] Correspondingly, the implementation manner in which the media access function receives the description information of each buffer sent from the display engine may be that the media access function receives the description information of each buffer, the description information of the buffer slices of each buffer, and the description information of each accessor sent from the display engine via the media access function API.

[0489] In step b, the media access function sends description information of each buffer, description information of a buffer slice of each buffer, and description information of each accessor to the buffer management module (the media access function includes sending description information of the target buffer, description information of a buffer slice of the target buffer, and description information of the target accessor to the buffer management module).

[0490] Correspondingly, the buffer management module receives descriptive information of each buffer, descriptive information of buffer slices of each buffer, and descriptive information of each access sent by the media access function (including the media access function receiving descriptive information of the target buffer, descriptive information of buffer slices of the target buffer, and descriptive information of the target accessor sent by the display engine).

[0491] In some embodiments, the implementation of step b above (the media access function sending the descriptive information of each buffer, the descriptive information of each buffer slice, and the descriptive information of each access to the buffer management module) may include the media access function sending the descriptive information of each buffer, the descriptive information of each buffer slice, and the descriptive information of each access to the buffer management module via a buffer API. Correspondingly, the implementation of the buffer management module receiving the descriptive information of each buffer, the descriptive information of each buffer slice, and the descriptive information of each access sent by the media access function may include the buffer management module receiving the descriptive information of each buffer, the descriptive information of each buffer slice, and the descriptive information of each access sent by the media access function via the buffer API.

[0492] In some embodiments, the implementation of method 1 (sending description information of each buffer, description information of buffer slices of each buffer, and description information of each accessor to the media access function and the buffer management module) includes the following steps c and d:

[0493] In step c, the display engine sends description information of each buffer, description information of a buffer slice of each buffer, and description information of each accessor to a media access function (including the display engine sending description information of the target buffer, description information of a buffer slice of the target buffer, and description information of the target accessor to a media access function).

[0494] Correspondingly, the media access function receives description information of each buffer, description information of buffer slices of each buffer, and description information of each accessor sent by the display engine (including receiving description information of the target buffer, description information of buffer slices of the target buffer, and description information of the accessor sent by the display engine).

[0495] In step d, the display engine sends description information of each buffer, description information of a buffer slice of each buffer, and description information of each accessor to the buffer management module (including the display engine sending description information of the target buffer, description information of a buffer slice of the target buffer, and description information of the target accessor to the buffer management module).

[0496] Correspondingly, the buffer management module receives the description information of each buffer, the description information of the buffer slices of each buffer, and the description information of each access sent by the display engine.

[0497] In some embodiments, the implementation of step d above (the display engine sending the description information of each buffer, the description information of each buffer slice, and the description information of each access to the buffer management module) may include the display engine sending the description information of each buffer, the description information of each buffer slice, and the description information of each access to the buffer management module via a buffer API.

[0498] Correspondingly, an implementation manner in which the buffer management module receives the descriptive information of each buffer, the descriptive information of the buffer slices of each buffer, and the descriptive information of each access sent by the display engine may include the buffer management module receiving the descriptive information of each buffer, the descriptive information of the buffer slices of each buffer, and the descriptive information of each access sent from the display engine via a buffer API.

[0499] In some embodiments, after the above steps S1707-S1712, the embodiments of the present application may use the following method 2 to send description information of each buffer, description information of buffer slices of each buffer, and description information of each accessor to a media access function, and send description information of each buffer and description information of buffer slices of each buffer to a buffer management module.

[0500] In some embodiments, the implementation of method 2 (sending description information of each buffer, description information of each buffer's buffer slice, and description information of each accessor to the media access function, and sending description information of each buffer and description information of each buffer's buffer slice to the buffer management module) includes the following steps e and f:

[0501] In step e, the display engine sends description information of each buffer, description information of a buffer slice of each buffer, and description information of each accessor to the media access function (including the display engine sending description information of the target buffer, description information of a buffer slice of the target buffer, and description information of the target accessor to the media access function).

[0502] Correspondingly, the media access function receives description information of each buffer and description information of buffer slices of each buffer sent by the display engine (the media access function includes receiving description information of the target buffer, description information of buffer slices of the target buffer, and description information of the accessor sent by the display engine).

[0503] In step f, the display engine sends description information of each buffer and description information of buffer slices of each buffer to a buffer management module (including the display engine sending description information of the target buffer and description information of the buffer slices of the target buffer to a media access function).

[0504] Correspondingly, the buffer management module receives description information of each buffer and description information of buffer slices of each buffer sent by the display engine (including the media access function receiving description information of the target buffer and description information of buffer slices of the target buffer sent by the display engine).

[0505] In some embodiments, the implementation of method 2 (sending description information of each buffer, description information of each buffer's buffer slice, and description information of each accessor to the media access function, and sending description information of each buffer and description information of each buffer's buffer slice to the buffer management module) includes the following steps g and h:

[0506] In step g, the display engine sends description information of each buffer, description information of a buffer slice of each buffer, and description information of each accessor to a media access function (including the display engine sending description information of the target buffer, description information of a buffer slice of the target buffer, and description information of the target accessor to a media access function).

[0507] Correspondingly, the media access function receives description information of each buffer and description information of buffer slices of each buffer sent by the display engine (the media access function includes receiving description information of the target buffer, description information of buffer slices of the target buffer, and description information of the accessor sent by the display engine).

[0508] In step f, the media access function sends description information of each buffer and description information of buffer slices of each buffer to the buffer management module (including the media access function sending description information of the target buffer and description information of buffer slices of the target buffer to the media access function).

[0509] Correspondingly, the buffer management module receives descriptive information for each buffer and descriptive information for buffer slices of each buffer sent by the media access function (including the media access function receiving descriptive information for the target buffer and descriptive information for the buffer slices of the target buffer sent by the display engine).

[0510] According to the above method 1, the description information of each buffer, the description information of the buffer slices of each buffer, and the description information of each access are sent to the media access function and the buffer management module, or according to the above method 2, the description information of each buffer, the description information of the buffer slices of each buffer, and the description information of each access are sent to the media access function, and the description information of each buffer and the buffer slices of each buffer are sent to the buffer management module, and then the following steps are continued to be executed.

[0511] In S1713, the buffer management module creates each buffer based on the description information of each buffer (including creating the target buffer based on the description information of the target buffer).

[0512] In S1714, the buffer management module divides buffer slices for each buffer based on description information of the buffer slices of each buffer (including dividing buffer slices for the target buffer based on description information of the buffer slices of the target buffer).

[0513] In S1715, the media access function writes decoded data corresponding to each media file to a buffer corresponding to each media file based on the description information of each buffer, the description information of the buffer slice of each buffer, and the description information of each accessor (the media access function includes writing the decoded data of the target media file to the target buffer based on the description information of the target buffer, the description information of the buffer slice of the target buffer, and the description information of the target accessor).

[0514] That is, information such as the buffer capacity in the description information of the buffer of the media access function, the buffer slice capacity in the description information of the buffer slice of the buffer, the accessor type in the description information of the accessor, and the data type in the description information of the accessor allows the decoded data corresponding to the media file to be written into the buffer in an accurate arrangement manner.

[0515] In S1716, the display engine obtains a scene description module corresponding to the three-dimensional scene to be rendered from the scene list of the scene description file.

[0516] In S1717, the display engine obtains description information of the three-dimensional scene to be rendered based on a scene description module corresponding to the three-dimensional scene to be rendered.

[0517] Here, the description information of the three-dimensional scene to be rendered includes an index value of a node description module corresponding to each node in the three-dimensional scene to be rendered.

[0518] In S1718, the display engine obtains the node description module corresponding to each node in the three-dimensional scene to be rendered from the node list of the scene description file based on the index value of the node description module corresponding to each node in the three-dimensional scene to be rendered.

[0519] In S1719, the display engine obtains description information of each node in the three-dimensional scene to be rendered based on a node description module corresponding to each node in the three-dimensional scene to be rendered.

[0520] Here, the description information of any node includes the index value of the mesh description module corresponding to the 3D mesh mounted by that node.

[0521] In some embodiments, the descriptive information of any node further includes the name of the node.

[0522] In S1720, the display engine obtains mesh description modules corresponding to the 3D meshes in the 3D scene to be rendered from the mesh list in the scene description file based on the index values ​​of the mesh description modules corresponding to the 3D meshes mounted on each node in the 3D scene to be rendered.

[0523] In S1721, the display engine obtains, based on a mesh description module corresponding to a 3D mesh in the 3D scene to be rendered, the data types contained in the 3D mesh in the 3D scene to be rendered and accessors for accessing each type of data of each 3D mesh in the 3D scene to be rendered.

[0524] In some embodiments, the method further comprises obtaining a name and a topology type of a 3D mesh in the to-be-rendered 3D scene based on a mesh description module corresponding to the 3D mesh in the to-be-rendered 3D scene.

[0525] In S1722, the display engine creates each accessor based on the description information of each accessor (including creating an accessor for accessing each type of data of each 3D mesh in the 3D scene to be rendered based on the description information of the accessor for accessing each type of data of each 3D mesh in the 3D scene to be rendered).

[0526] In S1723, the display engine reads decoded data of each media file from a buffer corresponding to each media file by each accessor (including reading each type of data of each 3D mesh in the 3D scene to be rendered from the target buffer memory by an accessor for accessing each type of data of each 3D mesh in the 3D scene to be rendered).

[0527] In S1724, the display engine renders the three-dimensional scene to be rendered based on the decoded data of each media file.

[0528] In some embodiments, the present application provides an apparatus for generating a scene description file, the apparatus comprising: a memory configured to store a computer program; a processor configured to, when calling a computer program, cause the scene description file generation device to implement the scene description file generation method according to any one of the above embodiments; The present invention provides a device for generating a scene description file including:

[0529] In some embodiments, some embodiments of the present application provide a computer-readable storage medium having stored thereon a computer program that, when executed by a computing device, causes the computing device to implement the scene description file generation method described in any of the above embodiments.

[0530] In some embodiments, some embodiments of the present application provide a computer program product that, when executed on a computer, causes the computer to implement the scene description file generation method described in any of the above embodiments.

[0531] Finally, it should be noted that the above embodiments are merely for explaining the technical solutions of the present application, and are not intended to limit the same. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they may modify the technical solutions described in the above embodiments or make equivalent substitutions for some or all of the technical features therein, and such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0532] For the sake of convenience, the above description has been given in combination with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations are possible in light of the above teachings. The selection and description of the above embodiments are intended to better explain the principles and practical applications, so that those skilled in the art can better use the above embodiments and various different variations thereof that are considered for specific applications.

[0533] The claims may include the following items: (Item 1) 1. A method for generating a scene description file, comprising: determining the type of media files in the three-dimensional scene to be rendered; When the type of the target media file in the three-dimensional scene to be rendered is a geometry-based point cloud compressed G-PCC coded point cloud, generating a target description module corresponding to the target media file based on description information of the target media file; and adding said target media description module to a media list of MPEG media in a scene description file of said three-dimensional scene to be rendered; How to generate a scene description file including: (Item 2) generating a target description module corresponding to the target media file based on description information of the target media file, Item 10. The method of claim 1, further comprising adding a media type syntax element to the options of the target media description module and setting the value of the media type syntax element to an encapsulation format value corresponding to a G-PCC coded point group. (Item 3) generating a target description module corresponding to the target media file based on description information of the target media file, Item 10. The method of claim 1, further comprising adding a first track index syntax element to an optional track array of the target media description module and setting a value of the first track index syntax element based on an encapsulation method of the target media file. (Item 4) setting a value of the first track index syntax element based on an encapsulation type of the target media file, If the target media file is a single-track capsule file, setting the value of the first track index syntax element to the index value of a codestream track of the target media file; Item 4. The method of item 3, comprising: if the target media file is a multi-track capsule file, setting the value of the first track index syntax element to an index value of a geometric codestream track of the target media file. (Item 5) generating a target description module corresponding to the target media file based on description information of the target media file, Item 10. The method of claim 1, further comprising adding a codec parameters syntax element to an optional track array of the target media description module, and setting a value of the codec parameters syntax element based on encoding parameters of the target media file, a codestream type of the target media file, and the ISO / IEC 23090-18G-PCC data transmission standard. (Item 6) generating a target description module corresponding to the target media file based on description information of the target media file, Item 10. The method of item 1, comprising adding a uniform resource identifier syntax element to an option of the target media description module, and setting the value of the uniform resource identifier syntax element as an access address of the target media file. (Item 7) adding a target scene description module corresponding to the three-dimensional scene to be rendered to a scene list in the scene description file; and adding index values ​​of node description modules corresponding to nodes in the scene to be rendered to a node index list of the target scene description module; The method according to item 1, further comprising: (Item 8) adding node description modules corresponding to nodes in the scene to be rendered to a node list in the scene description file; and adding to a mesh index list of said node description module an index value of a mesh description module corresponding to a three-dimensional mesh mounted on said node; The method according to item 1, further comprising: (Item 9) adding a mesh description module corresponding to a three-dimensional mesh in the scene to be rendered to a mesh list in the scene description file; adding syntax elements to the mesh description module corresponding to each type of data included in the three-dimensional mesh corresponding to the mesh description module; and The value of the syntax element corresponding to each type of data is set as the index value of the accessor description module corresponding to the accessor for accessing each type of data; The method according to item 1, further comprising: (Item 10) adding syntax elements to the mesh description module corresponding to each type of data included in the three-dimensional mesh corresponding to the mesh description module, adding an extension list to a primitive of a mesh description module corresponding to a three-dimensional mesh in the target media file; adding a target extension sequence to the extension list; and adding to the target extension array a syntax element corresponding to each type of data included in the corresponding 3D mesh; Item 10. The method according to item 9, comprising: (Item 11) adding syntax elements to the mesh description module corresponding to each type of data included in the three-dimensional mesh corresponding to the mesh description module, 10. The method of claim 9, further comprising adding syntax elements corresponding to each type of data included in the three-dimensional mesh corresponding to the mesh description module to attributes of primitives in the mesh description module. (Item 12) Adding syntax elements corresponding to each type of data included in the three-dimensional mesh corresponding to the mesh description module to the attributes of the primitives of the mesh description module includes: According to the syntax elements in the first syntax element set, add syntax elements corresponding to each type of data included in the corresponding three-dimensional mesh to the attributes of the primitive of the first mesh description module, wherein the first mesh description module is a mesh description module corresponding to the three-dimensional mesh in the media file whose type is a G-PCC coded point cloud; According to the syntax elements in the second syntax element set, add syntax elements corresponding to each type of data included in the corresponding three-dimensional mesh to the attributes of the primitive of the second mesh description module, wherein the second mesh description module is a mesh description module corresponding to a three-dimensional mesh in a media file whose type is not a G-PCC coded point cloud; Item 12. The method according to item 11, comprising: (Item 13) adding an accessor description module corresponding to the target accessor to an accessories list of the scene description file; Item 2. The method according to item 1, wherein the target accessor is an accessor for accessing decoded data of the target media file. (Item 14) adding an accessor description module corresponding to the target accessor to the accessory list of the scene description file, adding a data type syntax element to an accessor description module corresponding to the target accessor and setting a value of the data type syntax element based on the type of data accessed by the target accessor; adding an accessor type syntax element to an accessor description module corresponding to the target accessor and setting a value of the accessor type syntax element based on the type of the target accessor; adding a data number syntax element to an accessor description module corresponding to the target accessor, and setting a value of the data number syntax element based on the number of data accessed by the target accessor; adding an MPEG time-varying accessor to an accessor description module corresponding to said target accessor; adding a buffer slice index syntax element to the MPEG time-varying accessor and setting the value of the slice index syntax element based on an index value of a buffer slice description module corresponding to a buffer slice that stores data accessed by the target accessor; and adding a time-varying syntax element to the MPEG time-varying accessor and setting the value of the time-varying syntax element based on whether the value of a syntax element in a corresponding target accessor changes over time. (Item 15) adding a buffer description module corresponding to the target buffer to a buffer list of the scene description file; 2. The method according to claim 1, wherein the target buffer is a buffer for storing decoded data of the target media file. (Item 16) adding a buffer description module corresponding to the target buffer to the buffer list of the scene description file, adding a first byte length syntax element to the buffer description module and setting a corresponding value of the first byte length syntax element based on a capacity of the target buffer; adding an MPEG ring buffer to said buffer description module; adding a segment count syntax element to the MPEG ring buffer and setting a value of the corresponding segment count syntax element based on the number of stored segments in the target buffer; adding a media index syntax element to the MPEG ring buffer and setting a value of the media index syntax element based on an index value of the target media description module; and Item 16. The method of item 15, comprising adding a second track index syntax element to the MPEG ring buffer and setting a value of the second track index syntax element based on a track index value of source data for data stored in the target buffer. (Item 17) Item 16. The method of item 15, further comprising adding a buffer slice description module corresponding to a buffer slice of the target buffer to a buffer slice list of the scene description file. (Item 18) adding a buffer slice description module corresponding to a buffer slice of the target buffer to a buffer slice list of the scene description file, adding a buffer index syntax element to a buffer slice description module corresponding to a buffer slice of the target buffer, and setting a value of the buffer index syntax element based on an index value of a buffer description module corresponding to the target buffer to which the buffer slice belongs; adding a second byte length syntax element to a buffer slice description module corresponding to a buffer slice of the target buffer, and setting a value of the second byte length syntax element based on a capacity of the buffer slice; and Item 18. The method of item 17, further comprising: adding an offset amount syntax element to a buffer slice description module corresponding to a buffer slice of the target buffer; and setting a value of the offset amount syntax element based on an offset amount corresponding to stored data of the buffer slice. (Item 19) A scene description file generation device, comprising: a memory configured to store a computer program; a processor configured to, when calling a computer program, cause the scene description file generation device to implement the scene description file generation method according to any one of items 1-18; A generator of a scene description file including: (Item 20) 1. A method for rendering a three-dimensional scene, comprising: obtaining a scene description file for a three-dimensional scene to be rendered, the scene description file including a target media file of type geometry-based point cloud compressed G-PCC coded point cloud; Obtaining description information of the target media file based on a media description module corresponding to the target media file in a media list of Dynamic Picture Experts Group MPEG media in the scene description file; sending description information of the target media file to the media access function, so that the media access function obtains the target media file based on the description information of the target media file, processes the target media file to obtain decoded data of the target media file, and writes the decoded data of the target media file to a target buffer; reading decoded data of the target media file from the target buffer; and Rendering the to-be-rendered three-dimensional scene based on the decoded data of the target media file; A method for rendering three-dimensional scenes, including (Item 21) transmitting description information of the target media file to the media access function; 21. The method of claim 20, comprising sending description information of the target media file to the media access function via a media access function application programming interface API. (Item 22) Before reading the decoded data of the target media file from the target buffer, the method further comprises: obtaining a buffer description module corresponding to the target buffer from a buffer list of the scene description file; obtaining description information of the target buffer based on a buffer description module corresponding to the target buffer; obtaining a buffer slice description module corresponding to a buffer slice of the target buffer from a buffer slice list of the scene description file; Obtaining description information of a buffer slice of the target buffer based on a buffer slice description module corresponding to the buffer slice of the target buffer; and 21. The method of claim 20, further comprising sending descriptive information of the target buffer and descriptive information of buffer slices of the target buffer to the buffer management module, such that the buffer management module creates the target buffer based on the descriptive information of the target buffer and divides buffer slices for the target buffer based on the descriptive information of buffer slices of the target buffer. (Item 23) transmitting description information of the target buffer and description information of buffer slices of the target buffer to the buffer management module; 23. The method of claim 22, comprising sending description information of the target buffer and description information of a buffer slice of the target buffer to the buffer management module via a buffer API. (Item 24) Before reading the decoded data of the target media file from the target buffer, the method further comprises: 24. The method of claim 23, further comprising sending description information of the target buffer and description information of the buffer slices of the target buffer to the media access function so that the media access function writes the decoded data of the target media file to the target buffer based on the description information of the target buffer and the description information of the buffer slices of the target buffer. (Item 25) transmitting description information of the target buffer and description information of buffer slices of the target buffer to the media access function, 25. The method of claim 24, comprising sending description information of the target buffer and description information of buffer slices of the target buffer to the media access function via a media access function API. (Item 26) transmitting description information of the target buffer and description information of buffer slices of the target buffer to the buffer management module; 23. The method of claim 22, comprising sending descriptive information of the target buffer and descriptive information of buffer slices of the target buffer to the media access function via the media access function, such that the media access function forwards the descriptive information of the target buffer and the descriptive information of buffer slices of the target buffer to the buffer management module. (Item 27) Before reading the decoded data of the target media file from the target buffer, the method further comprises: obtaining an accessor description module corresponding to a target accessor, which is an accessor for accessing decoded data of the target media file, from an accessory list of the scene description file; obtaining description information of the target accessor based on an accessor description module corresponding to the target accessor; and 23. The method of claim 22, further comprising sending description information of the target access to the media access function so that the media access function writes decoded data of the target media file to the target buffer based on the description information of the target access. (Item 28) reading decoded data of the target media file from the target buffer includes: A method according to any one of items 20-27, comprising reading data of each type of 3D mesh in the 3D scene to be rendered from the decoded data of the target media file stored in the target buffer by an accessor for accessing data of each type of 3D mesh in the 3D scene to be rendered. (Item 29) Before reading, by an accessor for accessing the data of each type of three-dimensional mesh in the three-dimensional scene to be rendered, from the decoded data of the target media file stored in the target buffer, the method includes: obtaining a mesh description module corresponding to each 3D mesh in the to-be-rendered 3D scene from the mesh list in the scene description file based on an index value of the mesh description module corresponding to the 3D mesh mounted on each node in the to-be-rendered 3D scene; and Item 29. The method of item 28, further comprising: obtaining, based on a mesh description module corresponding to each 3D mesh in the 3D scene to be rendered, data types included in each 3D mesh in the 3D scene to be rendered and accessors for accessing each type of data of each 3D mesh in the 3D scene to be rendered. (Item 30) Before obtaining a mesh description module corresponding to each 3D mesh in the to-be-rendered 3D scene from the mesh list of the scene description file based on an index value of the mesh description module corresponding to each 3D mesh mounted on each node in the to-be-rendered 3D scene, the method obtaining, from the node list of the scene description file, a node description module corresponding to each node in the three-dimensional scene to be rendered, based on an index value of the node description module corresponding to each node in the three-dimensional scene to be rendered; and obtaining description information for each node in the three-dimensional scene to be rendered based on a node description module corresponding to each node in the three-dimensional scene to be rendered, wherein the description information for any node includes an index value of a mesh description module corresponding to a three-dimensional mesh mounted on the node; 30. The method of claim 29, further comprising: (Item 31) Before obtaining a node description module corresponding to each node in the three-dimensional scene to be rendered from the node list of the scene description file based on an index value of the node description module corresponding to each node in the three-dimensional scene to be rendered, the method obtaining a scene description module corresponding to the three-dimensional scene to be rendered from a scene list in the scene description file; obtaining description information of the three-dimensional scene to be rendered based on a scene description module corresponding to the three-dimensional scene to be rendered, the description information of the three-dimensional scene to be rendered including index values ​​of node description modules corresponding to each node in the three-dimensional scene to be rendered; 31. The method of claim 30, further comprising: (Item 32) The description information of any one of the three-dimensional meshes further includes an index value of an accessor description module corresponding to an accessor for accessing each type of data of the three-dimensional mesh, and obtaining an accessor for accessing each type of data of each three-dimensional mesh in the three-dimensional scene to be rendered based on the mesh description module corresponding to each three-dimensional mesh in the three-dimensional scene to be rendered includes: acquiring, from the accessory list of the scene description file, an accessor description module corresponding to an accessor for accessing each type of data of each three-dimensional mesh in the three-dimensional scene to be rendered, based on an index value of the accessor description module corresponding to an accessor for accessing each type of data of each three-dimensional mesh in the three-dimensional scene to be rendered; obtaining description information of an accessor for accessing each type of data of each three-dimensional mesh in the three-dimensional scene to be rendered based on an accessor description module corresponding to the accessor for accessing each type of data of each three-dimensional mesh in the three-dimensional scene to be rendered; and creating an accessor for accessing each type of data of each 3D mesh in the 3D scene to be rendered based on description information of the accessor for accessing each type of data of each 3D mesh in the 3D scene to be rendered; 32. The method according to item 31, comprising: (Item 33) 1. An apparatus for rendering a three-dimensional scene, comprising: a memory configured to store a computer program; a processor configured to, when calling a computer program, cause the device for rendering a three-dimensional scene to implement the method for rendering a three-dimensional scene according to any one of items 20-32; A rendering device for three-dimensional scenes, including

Claims

1. 1. A method for generating a scene description file, comprising: determining the type of media files in the three-dimensional scene to be rendered; When the type of the target media file in the three-dimensional scene to be rendered is a geometry-based point cloud compressed G-PCC coded point cloud, generating a target description module corresponding to the target media file based on description information of the target media file; and adding said target media description module to a media list of MPEG media in a scene description file of said three-dimensional scene to be rendered; How to generate a scene description file including:

2. generating a target description module corresponding to the target media file based on description information of the target media file, 2. The method of claim 1, further comprising adding a media type syntax element to the options of the target media description module and setting the value of the media type syntax element to an encapsulation format value corresponding to a G-PCC coding point group.

3. generating a target description module corresponding to the target media file based on description information of the target media file, 2. The method of claim 1, further comprising: adding a first track index syntax element to an optional tracks array of the target media description module; and setting a value of the first track index syntax element based on an encapsulation method of the target media file.

4. setting a value of the first track index syntax element based on an encapsulation type of the target media file, If the target media file is a single-track capsule file, setting the value of the first track index syntax element to an index value of a codestream track of the target media file; 4. The method of claim 3, further comprising: if the target media file is a multi-track capsule file, setting the value of the first track index syntax element to an index value of a geometric codestream track of the target media file.

5. generating a target description module corresponding to the target media file based on description information of the target media file, 2. The method of claim 1, further comprising: adding a codec parameters syntax element to an optional track array of the target media description module; and setting a value of the codec parameters syntax element based on encoding parameters of the target media file, a codestream type of the target media file, and the ISO / IEC 23090-18G-PCC data transmission standard.

6. generating a target description module corresponding to the target media file based on description information of the target media file, The method of claim 1 , further comprising adding a uniform resource identifier syntax element to an option of the target media description module, and setting a value of the uniform resource identifier syntax element as an access address of the target media file.

7. adding a target scene description module corresponding to the three-dimensional scene to be rendered to a scene list in the scene description file; and adding index values ​​of node description modules corresponding to nodes in the scene to be rendered to a node index list of the target scene description module; The method of claim 1 further comprising:

8. adding node description modules corresponding to nodes in the scene to be rendered to a node list in the scene description file; and adding to a mesh index list of said node description module an index value of a mesh description module corresponding to a three-dimensional mesh mounted on said node; The method of claim 1 further comprising:

9. adding a mesh description module corresponding to a three-dimensional mesh in the scene to be rendered to a mesh list in the scene description file; adding syntax elements to the mesh description module corresponding to each type of data included in the three-dimensional mesh corresponding to the mesh description module; and The value of the syntax element corresponding to each type of data is set as the index value of the accessor description module corresponding to the accessor for accessing each type of data; The method of claim 1 further comprising:

10. adding syntax elements to the mesh description module corresponding to each type of data included in the three-dimensional mesh corresponding to the mesh description module, adding an extension list to a primitive of a mesh description module corresponding to a three-dimensional mesh in the target media file; adding a target extension sequence to the extension list; and adding to the target extension array a syntax element corresponding to each type of data included in the corresponding three-dimensional mesh; 10. The method of claim 9, comprising:

11. adding syntax elements to the mesh description module corresponding to each type of data included in the three-dimensional mesh corresponding to the mesh description module, 10. The method of claim 9, further comprising adding to an attribute of a primitive in the mesh description module a syntax element corresponding to each type of data included in the three-dimensional mesh corresponding to the mesh description module.

12. Adding syntax elements corresponding to each type of data included in the three-dimensional mesh corresponding to the mesh description module to the attributes of the primitive of the mesh description module includes: According to the syntax elements in the first syntax element set, add syntax elements corresponding to each type of data included in the corresponding three-dimensional mesh to the attributes of the primitive of the first mesh description module, wherein the first mesh description module is a mesh description module corresponding to the three-dimensional mesh in the media file whose type is a G-PCC coded point cloud; According to the syntax elements in the second syntax element set, syntax elements corresponding to each type of data included in the corresponding three-dimensional mesh are added to the attributes of the primitives of the second mesh description module, and the second mesh description module is a mesh description module corresponding to a three-dimensional mesh in a media file whose type is not a G-PCC coded point cloud; 12. The method of claim 11, comprising:

13. adding an accessor description module corresponding to the target accessor to an accessories list of the scene description file; The method of claim 1 , wherein the target accessor is an accessor for accessing decoded data of the target media file.

14. adding an accessor description module corresponding to the target accessor to the accessory list of the scene description file, adding a data type syntax element to an accessor description module corresponding to the target accessor and setting a value of the data type syntax element based on the type of data accessed by the target accessor; adding an accessor type syntax element to an accessor description module corresponding to the target accessor and setting a value of the accessor type syntax element based on the type of the target accessor; adding a data number syntax element to an accessor description module corresponding to the target accessor, and setting a value of the data number syntax element based on the number of data accessed by the target accessor; adding an MPEG time-varying accessor to an accessor description module corresponding to said target accessor; adding a buffer slice index syntax element to the MPEG time-varying accessor and setting the value of the slice index syntax element based on an index value of a buffer slice description module corresponding to a buffer slice that stores data accessed by the target accessor; and 14. The method of claim 13, comprising at least one of adding a time-varying syntax element to the MPEG time-varying accessor and setting a value of the time-varying syntax element based on whether a value of a syntax element in a corresponding target accessor changes over time.

15. adding a buffer description module corresponding to the target buffer to a buffer list of the scene description file; The method of claim 1 , wherein the target buffer is a buffer for storing decoded data of the target media file.

16. adding a buffer description module corresponding to the target buffer to the buffer list of the scene description file, adding a first byte length syntax element to the buffer description module and setting a corresponding value of the first byte length syntax element based on a capacity of the target buffer; adding an MPEG ring buffer to said buffer description module; adding a segment count syntax element to the MPEG ring buffer and setting a value of the corresponding segment count syntax element based on the number of stored segments in the target buffer; adding a media index syntax element to the MPEG ring buffer and setting a value of the media index syntax element based on an index value of the target media description module; and 16. The method of claim 15, further comprising: adding a second track index syntax element to the MPEG ring buffer; and setting a value of the second track index syntax element based on a track index value of a source data of the data stored in the target buffer.

17. The method of claim 15 , further comprising adding a buffer slice description module corresponding to a buffer slice of the target buffer to a buffer slice list of the scene description file.

18. adding a buffer slice description module corresponding to a buffer slice of the target buffer to a buffer slice list of the scene description file, adding a buffer index syntax element to a buffer slice description module corresponding to a buffer slice of the target buffer, and setting a value of the buffer index syntax element based on an index value of a buffer description module corresponding to the target buffer to which the buffer slice belongs; adding a second byte length syntax element to a buffer slice description module corresponding to a buffer slice of the target buffer, and setting a value of the second byte length syntax element based on a capacity of the buffer slice; and 18. The method of claim 17, further comprising: adding an offset amount syntax element to a buffer slice description module corresponding to a buffer slice of the target buffer; and setting a value of the offset amount syntax element based on an offset amount corresponding to stored data in the buffer slice.

19. A scene description file generation device, comprising: a memory configured to store a computer program; a processor configured to, when calling a computer program, cause the scene description file generation device to implement the scene description file generation method according to any one of claims 1 to 18; A generator of a scene description file including:

20. 1. A method for analyzing a scene description file, comprising: obtaining a scene description file of a 3D scene to be rendered that includes a target media file of type G-PCC coded point cloud; obtaining a target media description module corresponding to the target media file from a Dynamic Picture Experts Group MPEG media list of the scene description file; and obtaining description information of the target media file based on the target media description module; How to parse a scene description file containing:

21. obtaining description information of the target media file based on the target media description module; obtaining a name for the target media file based on a value of a media name syntax element in the target media description module; determining whether the target media file needs to be auto-played based on a value of an auto-play syntax element in the target media description module; determining whether the target media file needs to be played in a circular manner based on a value of a circular playback syntax element in the target media description module; obtaining an encapsulation format of the target media file based on a value of a media type syntax element in an option of the target media description module; obtaining an access address of the target media file based on a value of a unique address identifier syntax element in an option of the target media description module; obtaining track information for the target media file based on a value of a first track index syntax element in an optional track array of the target media description module; and 21. The method of claim 20, comprising at least one of determining the type and decoding parameters of the codestream of the target media file based on values ​​of codec parameters syntax elements in an optional track sequence of the target media description module and the ISO / IEC 23090-18G-PCC data transmission standard.

22. obtaining a target scene description module corresponding to the three-dimensional scene to be rendered from a scene list in the scene description file; and The method of claim 20, further comprising obtaining description information of the three-dimensional scene to be rendered based on the target scene description module.

23. obtaining description information of the three-dimensional scene to be rendered based on the target scene description module, 23. The method of claim 22, further comprising determining an index value of a node description module corresponding to each node in the three-dimensional scene to be rendered based on index values ​​stated in a node index list of the target scene description module.

24. After determining an index value of a node description module corresponding to each node in the 3D scene to be rendered based on the index value stated in the node index list of the target scene description module, the method: obtaining, from the node list of the scene description file, a node description module corresponding to each node in the three-dimensional scene to be rendered, based on an index value of the node description module corresponding to each node in the three-dimensional scene to be rendered; and 24. The method of claim 23, further comprising obtaining description information for each node in the three-dimensional scene to be rendered based on a node description module corresponding to each node in the three-dimensional scene to be rendered.

25. obtaining description information of each node in the three-dimensional scene to be rendered based on a node description module corresponding to each node in the three-dimensional scene to be rendered, obtaining a name for each node in the to-be-rendered three-dimensional scene based on a value of a node name syntax element in a node description module corresponding to each node in the to-be-rendered three-dimensional scene; and 25. The method of claim 24, comprising at least one of determining an index value of a mesh description module corresponding to a three-dimensional mesh mounted to each node in the three-dimensional scene to be rendered based on an index value declared in a mesh index list in a node description module corresponding to each node in the three-dimensional scene to be rendered.

26. After determining index values ​​of mesh description modules corresponding to the 3D meshes mounted on each node in the 3D scene to be rendered, the method includes: obtaining, from the mesh list of the scene description file, mesh description modules corresponding to the three-dimensional meshes mounted on each node in the three-dimensional scene to be rendered, based on index values ​​of the mesh description modules corresponding to the three-dimensional meshes mounted on each node in the three-dimensional scene to be rendered; and 26. The method of claim 25, further comprising: obtaining description information for a three-dimensional mesh mounted on each node in the three-dimensional scene to be rendered based on a mesh description module corresponding to the three-dimensional mesh mounted on each node in the three-dimensional scene to be rendered.

27. obtaining description information of a 3D mesh mounted on each node in the 3D scene to be rendered based on a mesh description module corresponding to the 3D mesh mounted on each node in the 3D scene to be rendered, obtaining a name for each three-dimensional mesh based on a mesh name syntax element in a mesh description module corresponding to each three-dimensional mesh; obtaining a data type included in each three-dimensional mesh based on a data type syntax element in a mesh description module corresponding to each three-dimensional mesh; obtaining an index value of an accessor description module corresponding to an accessor for accessing each type of data of each 3D mesh based on the value of each data type syntax element; and 27. The method of claim 26, comprising at least one of: obtaining a type of topological structure of each three-dimensional mesh based on a value of a mode syntax element in a mesh description module corresponding to each three-dimensional mesh.

28. After obtaining an index value of an accessor description module corresponding to an accessor for accessing each type of data of each 3D mesh according to the value of the syntax element of each data type, the method includes: obtaining an accessor description module corresponding to an accessor for accessing each type of data of each three-dimensional mesh from the accessory list of the scene description file based on an index value of the accessor description module corresponding to an accessor for accessing each type of data of each three-dimensional mesh; 28. The method of claim 27, further comprising: obtaining description information of an accessor for accessing each type of data of each three-dimensional mesh based on an accessor description module corresponding to the accessor for accessing each type of data of each three-dimensional mesh.

29. obtaining each buffer description module in a buffer list of said scene description file; obtaining a value of a media index syntax element in each buffer description module; determining a buffer description module whose value of the media index syntax element is the same as the index value of the target media description module as a target buffer description module corresponding to a target buffer for buffering decoded data of the target media file; and 21. The method of claim 20, further comprising obtaining description information of the target buffer based on the target buffer description module.

30. obtaining description information of the target buffer based on the target buffer description module, obtaining a capacity of the target buffer based on a value of a first byte length syntax element in the target buffer description module; determining whether the target buffer is a ring buffer based on an MPEG extension modification based on whether the target buffer description module includes an MPEG ring buffer; obtaining a number of storage segments in the MPEG ring buffer based on a value of a Number of Segments in MPEG Ring Buffer syntax element of the target buffer description module; and 30. The method of claim 29, comprising at least one of: obtaining a track index value of source data for data buffered in the MPEG ring buffer based on a value of a second track index syntax element in the MPEG ring buffer of the target buffer description module.

31. obtaining each buffer slice description module in a buffer slice list of said scene description file; obtaining a value of a buffer index syntax element in each buffer slice description module; determining a buffer slice description module whose value of the buffer index syntax element is the same as the index value of the target buffer description module as a buffer slice description module corresponding to a buffer slice of the target buffer; and 30. The method of claim 29, further comprising: obtaining description information for a buffer slice of the target buffer based on a buffer slice description module corresponding to a buffer slice of the target buffer.

32. obtaining description information of a buffer slice of the target buffer based on a buffer slice description module corresponding to the buffer slice of the target buffer, obtaining a capacity of a buffer slice of the target buffer based on a value of a second byte length syntax element in a buffer slice description module corresponding to the buffer slice of the target buffer; and 32. The method of claim 31 , comprising at least one of: obtaining an offset for a buffer slice of the target buffer based on a value of an offset syntax element in a buffer slice description module corresponding to the buffer slice of the target buffer.

33. obtaining each accessor description module in the accessory list of said scene description file; obtaining the value of the buffer slice index syntax element in each accessor description module; determining an accessor description module whose value of the buffer slice index syntax element is the same as an index value of a buffer slice description module corresponding to a buffer slice of the target buffer as an accessor description module corresponding to an accessor for accessing data in the buffer slice of the target buffer; and 32. The method of claim 31, further comprising: obtaining description information of an accessor for accessing data in a buffer slice of the target buffer based on an accessor description module corresponding to an accessor for accessing data in a buffer slice of the target buffer.

34. Obtaining description information for an accessor based on an accessor description module corresponding to the accessor includes: determining the type of data the accessor accesses based on the value of a data type syntax element in the accessor description module; determining the type of the accessor based on the value of an accessor type syntax element in the accessor description module; determining the number of data items that the accessor accesses based on the value of a number-of-data syntax element in the accessor description module; determining whether the accessor is a time-varying accessor based on whether the accessor description module includes an MPEG time-varying accessor; determining an index value of the buffer slice description module corresponding to the buffer slice that stores the data accessed by the accessor based on the value of a buffer slice index syntax element in the MPEG time-varying accessor of the accessor description module; and 34. The method of claim 28 or 33, comprising at least one of: determining whether the value of a syntax element in an accessor changes over time based on the value of a time-varying syntax element in an MPEG time-varying accessor of the accessor description module.

35. An apparatus for analyzing a scene description file, comprising: a memory configured to store a computer program; a processor configured to cause the three-dimensional scene rendering device to implement the scene description file analysis method of any one of claims 20-34 when invoking a computer program.

36. 1. A method for rendering a three-dimensional scene, comprising: obtaining a scene description file of a three-dimensional scene to be rendered, the scene description file including a target media file of type geometry-based point cloud compressed G-PCC coded point cloud; Obtaining description information of the target media file based on a media description module corresponding to the target media file in a media list of Dynamic Picture Experts Group MPEG media in the scene description file; sending description information of the target media file to the media access function, so that the media access function obtains the target media file based on the description information of the target media file, processes the target media file to obtain decoded data of the target media file, and writes the decoded data of the target media file to a target buffer; reading decoded data of the target media file from the target buffer; and Rendering the to-be-rendered 3D scene based on the decoded data of the target media file; A method for rendering a three-dimensional scene comprising:

37. transmitting description information of the target media file to the media access function; 37. The method of claim 36, comprising sending description information of the target media file to the media access function via a media access function application programming interface API.

38. Before reading the decoded data of the target media file from the target buffer, the method further comprises: obtaining a buffer description module corresponding to the target buffer from a buffer list of the scene description file; obtaining description information of the target buffer based on a buffer description module corresponding to the target buffer; obtaining a buffer slice description module corresponding to a buffer slice of the target buffer from a buffer slice list of the scene description file; Obtaining description information of a buffer slice of the target buffer based on a buffer slice description module corresponding to the buffer slice of the target buffer; and 37. The method of claim 36, further comprising sending descriptive information of the target buffer and descriptive information of buffer slices of the target buffer to the buffer management module, such that the buffer management module creates the target buffer based on the descriptive information of the target buffer and performs buffer slice division on the target buffer based on the descriptive information of buffer slices of the target buffer.

39. transmitting description information of the target buffer and description information of buffer slices of the target buffer to the buffer management module; 39. The method of claim 38, comprising sending description information of the target buffer and description information of a buffer slice of the target buffer to the buffer management module via a buffer API.

40. Before reading the decoded data of the target media file from the target buffer, the method further comprises:

40. The method of claim 39, further comprising sending description information of the target buffer and description information of buffer slices of the target buffer to the media access function so that the media access function writes decoded data of the target media file to the target buffer based on the description information of the target buffer and the description information of buffer slices of the target buffer.

41. transmitting description information of the target buffer and description information of buffer slices of the target buffer to the media access function, 41. The method of claim 40, comprising sending description information of the target buffer and description information of buffer slices of the target buffer to the media access function via a media access function API.

42. transmitting description information of the target buffer and description information of buffer slices of the target buffer to the buffer management module; 39. The method of claim 38, comprising sending descriptive information of the target buffer and descriptive information of buffer slices of the target buffer to the media access function via the media access function such that the media access function forwards the descriptive information of the target buffer and the descriptive information of buffer slices of the target buffer to the buffer management module.

43. Before reading the decoded data of the target media file from the target buffer, the method further comprises: obtaining an accessor description module corresponding to a target accessor, which is an accessor for accessing decoded data of the target media file, from an accessory list of the scene description file; obtaining description information of the target accessor based on an accessor description module corresponding to the target accessor; and 39. The method of claim 38, further comprising sending descriptive information of the target access to the media access function so that the media access function writes decoded data of the target media file to the target buffer based on the descriptive information of the target access.

44. reading decoded data of the target media file from the target buffer includes: A method according to any one of claims 36 to 43, comprising reading data of each type of three-dimensional mesh in the three-dimensional scene to be rendered from the decoded data of the target media file stored in the target buffer by an accessor for accessing data of each type of three-dimensional mesh in the three-dimensional scene to be rendered.

45. Before reading, by an accessor for accessing the data of each type of three-dimensional mesh in the three-dimensional scene to be rendered from the decoded data of the target media file stored in the target buffer, the method includes: obtaining a mesh description module corresponding to each three-dimensional mesh in the three-dimensional scene to be rendered from the mesh list of the scene description file based on an index value of the mesh description module corresponding to the three-dimensional mesh mounted on each node of the three-dimensional scene to be rendered; and 45. The method of claim 44, further comprising: obtaining, based on a mesh description module corresponding to each three-dimensional mesh in the three-dimensional scene to be rendered, a type of data included in each three-dimensional mesh in the three-dimensional scene to be rendered and an accessor for accessing each type of data for each three-dimensional mesh in the three-dimensional scene to be rendered.

46. Before obtaining a mesh description module corresponding to each three-dimensional mesh in the three-dimensional scene to be rendered from the mesh list of the scene description file based on an index value of the mesh description module corresponding to each three-dimensional mesh mounted on each node in the three-dimensional scene to be rendered, the method obtaining, from the node list of the scene description file, a node description module corresponding to each node in the three-dimensional scene to be rendered, based on an index value of the node description module corresponding to each node in the three-dimensional scene to be rendered; and obtaining description information for each node in the three-dimensional scene to be rendered based on a node description module corresponding to each node in the three-dimensional scene to be rendered, wherein the description information for any node includes an index value of a mesh description module corresponding to a three-dimensional mesh mounted on the node; 46. ​​The method of claim 45, further comprising:

47. Before obtaining a node description module corresponding to each node in the three-dimensional scene to be rendered from the node list of the scene description file based on an index value of the node description module corresponding to each node in the three-dimensional scene to be rendered, the method obtaining a scene description module corresponding to the three-dimensional scene to be rendered from a scene list in the scene description file; obtaining description information of the three-dimensional scene to be rendered based on a scene description module corresponding to the three-dimensional scene to be rendered, the description information of the three-dimensional scene to be rendered including index values ​​of node description modules corresponding to each node in the three-dimensional scene to be rendered; 47. The method of claim 46, further comprising:

48. The description information of any one of the three-dimensional meshes further includes an index value of an accessor description module corresponding to an accessor for accessing each type of data of the three-dimensional mesh, and obtaining an accessor for accessing each type of data of each three-dimensional mesh in the three-dimensional scene to be rendered based on the mesh description module corresponding to each three-dimensional mesh in the three-dimensional scene to be rendered includes: acquiring, from the accessory list of the scene description file, an accessor description module corresponding to an accessor for accessing each type of data of each three-dimensional mesh in the three-dimensional scene to be rendered, based on an index value of the accessor description module corresponding to an accessor for accessing each type of data of each three-dimensional mesh in the three-dimensional scene to be rendered; obtaining description information of an accessor for accessing each type of data of each three-dimensional mesh in the three-dimensional scene to be rendered based on an accessor description module corresponding to the accessor for accessing each type of data of each three-dimensional mesh in the three-dimensional scene to be rendered; and creating an accessor for accessing each type of data of each three-dimensional mesh in the three-dimensional scene to be rendered based on description information of the accessor for accessing each type of data of each three-dimensional mesh in the three-dimensional scene to be rendered; 48. The method of claim 47, comprising:

49. 1. An apparatus for rendering a three-dimensional scene, comprising: a memory configured to store a computer program; a processor configured to, when invoking a computer program, cause the device for rendering a three-dimensional scene to implement the method for rendering a three-dimensional scene according to any one of claims 36 to 48; 2. A three-dimensional scene rendering apparatus comprising:

Citation Information

Patent Citations

  • Three-dimensional point cloud data processing method and device, storage medium and electronic device

    CN112700550A

  • Three-dimensional content processing methods and apparatus

    WO2021258325A1

  • Tile tracks for geometry‑based point cloud data

    WO2022032161A1

  • Method and apparatus for media scene description

    WO2022150077A1