Three-dimensional scene rendering method and device
By acquiring and reconstructing facial marker information in a 3D scene, the problem of undeclared facial markers in the immersive media scene description framework was solved, enabling digital human animation and related processing, and improving immersion and interactivity.
Patent Information
- Application Number
- CN202410916908.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-09
- Publication Date
- 2026-01-09
AI Technical Summary
The existing immersive media scene description framework does not declare information related to facial markers, making it impossible to realize digital human animation and related processing based on facial markers.
The process involves obtaining the scene description file of the 3D scene to be rendered, extracting facial markers of the target digital human and description information from the media file, reconstructing the dynamic 3D model through the media access function, writing it into the buffer, and finally rendering based on the dynamic 3D model.
It enables digital human animation and related processing based on facial markers, enhancing the immersiveness and interactivity of immersive media.
Smart Images

Figure CN121304879A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Some embodiments of the present application relate to the technical field of video processing. More specifically, it relates to a rendering method and device of a three-dimensional scene. BACKGROUND
[0002] With the rapid development of immersive media related technologies, the working mode of superimposing online office is accepted by more and more people, and immersive media related technologies are further popularized in people's work and life. Compared with improving the sense of immersion by means of a large size display screen, virtual reality (VR) / augmented reality (AR) equipment has the design advantage of fitting the human physiological structure, and can make people feel a lively and real scene through binocular stereoscopic vision, so it is more likely to use VR / AR equipment as the hardware basis of immersive media in the future. Part of the VR / AR technology is user representation, or simply called avatar. When the user wears the VR / AR equipment, the user hopes that his real image or virtual image can appear in the virtual scene presented by the head-mounted equipment, and can move in the virtual scene synchronously with the user's movement in the real world, which is an important function that the avatar should have and an important link to improve the sense of immersion.
[0003] At present, the standard with the standard number of ISO / IEC 23090-14 defines the face landmarks of the default avatar model of the scene description file (SD), and points out that the face landmarks can be used to realize avatar animation and related processing. However, the related information of the face landmarks is not declared in the scene description file based on the gilt2.0 extension, so the immersive media scene description framework cannot be used to realize avatar animation and related processing based on the face landmarks at present. SUMMARY
[0004] The exemplary embodiments of the present application provide a rendering method and device of a three-dimensional scene, which are used to solve the problem that the related information of the face landmarks is not declared in the scene description file, and then the immersive media scene description framework cannot be used to realize avatar animation and related processing based on the face landmarks.
[0005] Some embodiments of the present application provide technical solutions as follows:
[0006] In a first aspect, some embodiments of the present application provide a rendering method of a three-dimensional scene, comprising:
[0007] obtaining a scene description file of a three-dimensional scene to be rendered;
[0008] The scene description file is used to obtain the description information of the facial marker points of the target digital human in the three-dimensional scene to be rendered and the description information of the first media file. The first media file is a media file that includes the index information and position coordinates of the facial marker points of the target digital human.
[0009] The description information of the facial markers of the target digital human and the description information of the first media file are sent to the media access function, so that the media access function can obtain the index information and position coordinates of the facial markers of the target digital human based on the description information of the facial markers of the target digital human and the description information of the first media file, reconstruct the dynamic three-dimensional model corresponding to the target digital human based on the index information and position coordinates of the facial markers of the target digital human, and write the dynamic three-dimensional model of the target digital human into the buffer corresponding to the target digital human;
[0010] The dynamic 3D model of the target digital human is read from the cache corresponding to the target digital human, and the 3D scene to be rendered is rendered based on the dynamic 3D model of the target digital human.
[0011] Secondly, some embodiments of this application provide a rendering apparatus for a three-dimensional scene, including:
[0012] The acquisition unit is used to acquire the scene description file of the 3D scene to be rendered;
[0013] The parsing unit is used to obtain the description information of the facial marker points of the target digital human in the three-dimensional scene to be rendered and the description information of the first media file according to the scene description file. The first media file is a media file that includes the index information and position coordinates of the facial marker points of the target digital human.
[0014] The sending unit is configured to send the description information of the facial markers of the target digital human and the description information of the first media file to the media access function, so that the media access function can obtain the index information and position coordinates of the facial markers of the target digital human based on the description information of the facial markers of the target digital human and the description information of the first media file, reconstruct the dynamic three-dimensional model corresponding to the target digital human based on the index information and position coordinates of the facial markers of the target digital human, and write the dynamic three-dimensional model of the target digital human into the buffer corresponding to the target digital human;
[0015] The rendering unit is used to read the dynamic 3D model of the target digital human from the cache corresponding to the target digital human, and to render the 3D scene to be rendered based on the dynamic 3D model of the target digital human.
[0016] Thirdly, some embodiments of this application provide an electronic device, including:
[0017] Memory, configured to store computer programs;
[0018] The processor is configured to cause the rendering apparatus of the three-dimensional scene to implement the three-dimensional scene rendering method described in the first aspect when a computer program is invoked.
[0019] Fourthly, some embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a computing device, causes the computing device to implement the three-dimensional scene rendering method described in the first aspect.
[0020] Fifthly, some embodiments of this application provide a computer program product that, when run on a computer, enables the computer to implement the rendering method for the three-dimensional scene.
[0021] In a sixth aspect, some embodiments of this application provide a chip including a processor and a memory, the memory being used to store programs or instructions executable on the processor, and the processor being used to execute the programs or instructions to cause the rendering method of the three-dimensional scene as described in the first aspect to be executed.
[0022] As can be seen from the above technical solutions, the rendering method for a three-dimensional scene provided in some embodiments of this application first obtains a scene description file of the three-dimensional scene to be rendered, then obtains the description information of the facial marker points of the target digital human in the three-dimensional scene to be rendered and the description information of the first media file according to the scene description file, and sends the description information of the facial marker points of the target digital human and the description information of the first media file to a media access function, so that the media access function obtains the index information and position coordinates of the facial marker points of the target digital human according to the description information of the facial marker points of the target digital human and the description information of the first media file, reconstructs the dynamic three-dimensional model corresponding to the target digital human according to the index information and position coordinates of the facial marker points of the target digital human, writes the dynamic three-dimensional model of the target digital human into the buffer corresponding to the target digital human, reads the dynamic three-dimensional model of the target digital human from the buffer corresponding to the target digital human, and renders the three-dimensional scene to be rendered based on the dynamic three-dimensional model of the target digital human. Since the rendering method of the three-dimensional scene provided in some embodiments of this application can obtain the description information of the facial marker points of the target digital human and the description information of the first media file including the index information and position coordinates of the facial marker points of the target digital human from the scene description file, and send the obtained information to the media access function, so that the media access function can obtain the index information and position coordinates of the facial marker points of the target digital human according to the description information of the facial marker points of the target digital human and the description information of the first media file, and reconstruct the dynamic three-dimensional model corresponding to the target digital human according to the index information and position coordinates of the facial marker points of the target digital human, some embodiments of this application can solve the problem that the immersive media scene description framework cannot realize digital human animation and related processing based on facial marker points because the scene description file does not declare the relevant information of facial marker points. Attached Figure Description
[0023] To more clearly illustrate the implementation methods in some embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.
[0024] Figure 1 A schematic diagram of the structure of an immersive media scene description framework in some embodiments of this application is shown;
[0025] Figure 2 The diagram shows a schematic representation of the structure of a scene description file in some embodiments of this application;
[0026] Figure 3The following are schematic diagrams illustrating the structure of scene description files in other embodiments of this application;
[0027] Figure 4 This illustration shows a schematic diagram of how skeletal points drive the rotation of skin apex in some embodiments of this application;
[0028] Figure 5 Schematic diagrams of animation forms based on facial blending shapes are shown in some embodiments of this application;
[0029] Figure 6 A schematic diagram showing the index values of facial marker points in some embodiments of this application is illustrated;
[0030] Figure 7 The illustration shows a schematic diagram of the application scenarios of facial markers in some embodiments of this application;
[0031] Figure 8 The following are schematic diagrams illustrating the structure of scene description files in other embodiments of this application;
[0032] Figure 9 The following are schematic diagrams illustrating the structure of scene description files in other embodiments of this application;
[0033] Figure 10 Schematic diagrams of digital human pipelines in some embodiments of this application are shown;
[0034] Figure 11 The flowcharts illustrating the steps of a scene description file generation method in some embodiments of this application are shown.
[0035] Figure 12 The flowcharts illustrating the steps of the scene description file parsing method in some embodiments of this application are shown.
[0036] Figure 13 A flowchart illustrating the steps of a three-dimensional scene rendering method in some embodiments of this application is shown;
[0037] Figure 14 A flowchart illustrating the steps of a method for processing scene data in a three-dimensional scene according to some embodiments of this application is shown.
[0038] Figure 15 The flowcharts illustrating the steps of a cache management method in some embodiments of this application are shown.
[0039] Figure 16 A schematic diagram of the structure of a rendering apparatus for a three-dimensional scene in some embodiments of this application is shown. Detailed Implementation
[0040] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.
[0041] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0042] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.
[0043] The use of phrases such as "some implementations" or "some embodiments" in the specification indicates that the described implementations or embodiments may include specific features, structures, or characteristics, but not every embodiment may necessarily include that specific feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same implementation. Additionally, when describing a specific feature, structure, or characteristic in connection with an embodiment, it is considered that implementing that feature, structure, or characteristic in connection with other implementations (whether explicitly described herein or not) is within the knowledge of those skilled in the art.
[0044] Some embodiments of this application relate to a scene description framework for immersive media. See also... Figure 1 The immersive media scene description framework shown decouples media file access and processing from media file rendering so that the display engine 11 can focus on media rendering. A Media Access Function (MAF) 12 is designed to handle media file access and processing. An Application Programming Interface (API) for the Media Access Function 12 is also designed, allowing the display engine 11 and the Media Access Function 12 to interact via the API. The display engine 11 can issue commands to the Media Access Function 12 through the API, and the Media Access Function 12 can also request commands from the display engine 11 through the API.
[0045] The general workflow of an immersive media scene description framework may include: 1) Display engine 11 obtains the scene description file provided by the immersive media service provider. 2) Display engine 11 parses the scene description file, obtains the access address of the media file, the attribute information of the media file (media type and encoding / decoding parameters, etc.), and the format requirements of the processed media file, and calls the media access function API to pass all or part of the information obtained from parsing the scene description file to media access function 12. 3) Media access function 12, based on the information passed by display engine 11, requests to download the specified media file from the media resource server or obtains the specified media file from the local machine, and establishes a corresponding pipeline for the media file. Subsequently, the media file is processed in the pipeline through decapsulation, decryption, decoding, post-processing, etc., to convert the media file from the encapsulated format to the format specified by display engine 11. 4) Media access function 12 stores the output data obtained after all processing in the specified cache. 5) Display engine 11 reads the fully processed data from the specified cache and renders the media file based on the data read from the cache.
[0046] The following section further explains the documents and functional modules involved in the scene description framework of immersive media.
[0047] I. Scene Description File
[0048] In the workflow of the scene description framework for immersive media, scene description files can be used to describe the structure of a 3D scene (its features can be described using 3D meshes), textures (such as texture maps), animations (rotation, translation), camera viewpoint position (rendering perspective), and other content.
[0049] In related technical fields, GL Transfer Format 2.0 (glTF2.0) has been identified as a candidate format for scene description files, which can meet the needs of Moving Picture Experts Group-Immersive (MPEG-I) and six degrees of freedom (6DoF) applications. (See reference...) Figure 2 As shown, Figure 2 This is a schematic diagram of the structure of the scene description file in the glTF2.0 scene description standard (ISO / IEC12113).
[0050] like Figure 2 As shown, the scene description file may include a scene description module (scene) 201. Figure 2The scene description module (scene) 201 in the scene description file shown can be used to describe the 3D scene contained in the scene description file. A scene description file may contain any number of 3D scenes, and each 3D scene is represented by a scene description module 201. The scene description modules 201 are parallel to each other, that is, the 3D scenes are parallel to each other.
[0051] like Figure 2 As shown, the scene description file may include a node description module (node) 202. Figure 2 The node description module 202 in the scene description file is the next level description module after the scene description module 201, and can be used to describe the objects contained in the 3D scene described by the scene description module 201. Each 3D scene may contain many specific objects, such as avatars, nearby 3D objects, and distant background images. The scene description file describes these specific objects through the node description module 202. Each node description module 202 can represent an object or a collection of objects. The relationship between node description modules 202 reflects the relationship between the various components in the 3D scene described by the scene description module 201. A scene described by a scene description module 201 can contain one or more objects. Multiple objects can be in a parallel or hierarchical relationship; that is, node description modules 202 can be in a parallel relationship or an inclusive / inclusive relationship. This allows multiple specific objects to be described together or separately. If a node is contained by another node, the contained node is called a child node, and "children" is used instead of "node" to represent child nodes. By flexibly combining nodes and child nodes, a hierarchical node structure can be formed, thereby expressing rich scene content.
[0052] like Figure 2 As shown, the scene description file may include a mesh description module (mash) 203. Figure 2The mesh description module 203 in the scene description file is the next level description module after the node description module 202, and can be used to describe the features of the object represented by the node description module 202. The mesh description module 203 is a collection of one or more primitives, each of which can include one attribute. The primitive's attribute defines the properties required for rendering by the graphics processing unit (GPU). Attributes can include: position (3D coordinates), normal (normal vector), tangent (tangent vector), texcoord_n (texture coordinates), color_n (color: RGB or RGBA), joints_n (attributes related to the skinning description module 214), and weights_n (attributes related to the skinning description module 214), etc. Because the mesh description module 203 contains a very large number of vertices, and each vertex contains multiple attribute information, it is inconvenient to directly store the large amount of media data contained in the media file in the mesh description module 203 of the scene description file. Instead, the scene description file specifies the access address (Uniform Resource Identifier, URI) of the media file. When data in the media file is needed, it is downloaded according to the access address of the media file, thus achieving separation between the scene description file and the media file. Therefore, under normal circumstances, the mesh description module 203 does not directly store media data, but stores the index value of the accessor description module 204 corresponding to each attribute, and points to the corresponding data in the buffer slice (bufferView) of the buffer through the accessor description module 204. In addition, the primitives of the mesh description module 203 may also contain a syntax element called mode. The mode syntax element can be used to describe the topology of the 3D mesh when the graphics processing unit (GPU) draws it, for example, mode=0 represents scattered points, mode=1 represents lines, mode=4 represents triangles, etc.
[0053] The definitions of the syntax elements in the properties (mash.primitives.attributes) of the primitives in the mesh description module 203 are shown in Table 1 below:
[0054] Table 1
[0055]
[0056] Here is an example JSON sample of a grid description module 203:
[0057]
[0058] The above example uses the attributes (mash.primitives.attributes) of the primitives of the mesh description module 203, which include the position and color syntax elements shown in Table 1 above, but exclude the normal vector syntax element (noraml), tangent vector syntax element (tangent), texture coordinate syntax element (texcoord), joint syntax element (joints_n), and weight syntax element (weights_n).
[0059] In the mesh description module 203 mentioned above, the value of "position" is 1, which points to the accessor description module with index 1, and finally points to the vertex coordinate data stored in the buffer; the value of "color 0" is 2, which points to the accessor description module with index 2, and finally points to the color data stored in the buffer.
[0060] The types of accessors pointed to by the syntax elements in the properties (mash.primitives.attributes) of the primitives in the mesh description module 203 are defined in Table 2 below:
[0061] Table 2
[0062] Accesser type Number of component channels Meaning SCALAR 1 Scalar VEC2 2 Two-dimensional vector VEC3 3 Three-dimensional vector VEC4 4 Four-dimensional vector MAT2 4 Two-dimensional matrix MAT3 9 Three-dimensional matrix MAT4 16 Four-dimensional matrix
[0063] The data types defined in the properties (mash.primitives.attributes) of the primitives in the mesh description module 203 are shown in Table 3 below:
[0064] Table 3
[0065] Type code Data type Signed or unsigned Number of bits 5120 Signed byte Signed 8 5121 Unsigned byte Unsigned 8 5122 Signed short Signed 16 5123 Unsigned short Unsigned 16 5125 Unsigned int Unsigned 32 5126 Float Signed 32
[0066] In some embodiments, scene description files and media files can be merged into a single binary file, thereby reducing the types and number of files.
[0067] like Figure 2 As shown, the scene description file may include an accessor description module (accessor) 204, a buffer slice description module (bufferView) 205, and a buffer description module (buffer) 206. Figure 2The accessor description module 204, buffer slice description module 205, and buffer description module 206 in the scene description file shown can be used to implement the mesh description module 203's layer-by-layer fine-grained indexing of the media file data. As mentioned above, the mesh description module 203 does not store specific media data, but rather stores the index values of the corresponding accessor description modules 204, and accesses the specific media data through the accessor described by the accessor description module 204 pointed to by the index value. The indexing process of the mesh description module 203 for media data may include: first, the index value declared by the syntax element in the mesh description module 203 will point to the corresponding accessor description module 204; then, the accessor description module 204 will point to the corresponding buffer slice description module 205; finally, the buffer slice description module 205 will point to the corresponding buffer description module 206. Figure 2 The buffer description module 206 in the scene description file points to the corresponding media file and contains information such as the access address and byte length of the media file. It can be used to describe the buffer that caches the media data of the media file. A buffer can be divided into one or more buffer slices. The buffer slice description module 205 is responsible for partially accessing the media data in the buffer and contains information such as the starting byte offset and byte length of the accessed data. Through the buffer slice description module 205 and the buffer description module 206, partial access to the media file data can be achieved. The accessor description module 204 is responsible for adding additional information to the portion of data defined in the buffer slice description module 205, such as data type, quantity, and numerical range. This three-layer structure enables the function of retrieving partial data from a media file, which is beneficial for accurate data retrieval and also helps to reduce the number of media files.
[0068] like Figure 2 As shown, the scene description file may include a camera description module (camera) 207. Figure 2 The camera description module 207 in the scene description file is a next-level description module below the node description module 202. It can be used to describe the viewpoint, perspective, and other visual viewing information when the user views the object described by the node description module 202. To enable the user to immerse themselves in and view the 3D scene, the node description module 202 can also point to the camera description module 207, and the camera description module 207 can describe the viewpoint, perspective, and other visual viewing information when the user views the object described by the node description module 202.
[0069] like Figure 2 As shown, the scene description file may include a lighting description module (1ight) 208. Figure 2The lighting description module (light) 208 in the scene description file is the next level description module of the node description module 202. It can be used to describe lighting-related information such as the lighting intensity, ambient light color, lighting direction, and light source position of the object described by the node description module 202.
[0070] like Figure 2 As shown, the scene description file may include a material description module (material) 209. Figure 2 The material description module 209 in the scene description file is a lower-level description module than the mesh description module 203, and can be used to describe the material information of the 3D objects described by the mesh description module 203. When describing 3D objects, simply using the mesh description module 203 to describe the geometric information of the 3D objects, or simply defining the color and / or position of the 3D objects, cannot improve the realism of the 3D objects. This requires adding more information to the surface of the 3D objects. For conventional 3D modeling techniques, this process can also be simply referred to as texture mapping or adding textures. The scene description file in the glTF2.0 scene description standard also uses this description module. The material description module 209 uses a set of general parameters to define materials to describe the material information of geometric objects appearing in the 3D scene. This material description module 209 generally uses a metallic-roughness model to describe the material of virtual objects, and the material characteristic parameters based on the metallic-roughness model adopt the widely used physically based rendering (PBR) material representation. Based on this, the material description module 209 provides a detailed description of the metal-roughness material properties of the object. The definitions of the syntax elements in the material description module 209 are shown in Table 4:
[0071] Table 4
[0072]
[0073] In some embodiments, the definitions of the syntax elements in the metal-roughness (material.PbrMetarialRoughness) of the material description module 209 are shown in Table 5 below:
[0074] Table 5
[0075]
[0076] The value of each attribute in the Metal-Roughness section of Material Description Module 209 can be defined using factors and / or textures (e.g., baseColorTexture and baseColorFactor). If no texture is given, the value of all corresponding texture components in this material model is set to 1.0. If both factors and textures are present, the factor value acts as a linear multiplier of the corresponding texture value. Texture binding is defined by the index of the texture object and, optionally, the texture coordinate index.
[0077] Here is an example of a JSON representation of a material description module 209:
[0078]
[0079] Parsing the material description module 209, the current material is named "gold" using the material name syntax element and its value ("name":"gold"). The base color value of the current material is determined to be [1.000,0.766,0.336,1.0] using the color syntax element and its value ("basecolorFactor":[1.000,0.766,0.336,1.0]) under the pbrMetallicRoughness array. The metallic value of the current material is determined to be "1.0" using the metallic syntax element and its value ("metalnessFactor":1.0) under the pbrMetallicRoughness array. Finally, the roughness value of the current material is determined to be "0.0" using the roughness syntax element and its value ("roughnessFactor":0.0) under the pbrMetallicRoughness array.
[0080] like Figure 2 As shown, the scene description file may include a texture description module (texture) 210, a sampler description module (sampler) 211, and a texture map description module (image) 212. Figure 2The texture description module 210 in the scene description file is the next level description module after the material description module 209. It can be used to describe the color of the 3D object described by the material description module 209 and other properties used in the material definition. Texture is an important aspect of giving an object a realistic appearance. Through texture, the main color of the object and other properties used in the material definition can be defined to accurately describe the appearance of the rendered object. The material itself can define multiple texture objects, which can be used as textures for virtual objects during rendering and can be used to encode different material properties. The texture description module 210 uses sampler syntax elements and texture map syntax element indexes to reference a sampler description module 211 and a texture map description module 212. The texture map description module 212 contains a Uniform Resource Identifier (URI) that links to the texture map or binary file package actually used by the texture description module 210. The sampler description module 211 can be used to describe the filtering and wrapping modes of the texture. The responsibilities and collaborative relationships of the material description module 209, texture description module 210, sampler description module 211, and texture mapping description module 212 can be summarized as follows: The material description module 209, together with the texture description module 210, defines the color and physical information of the object's surface. The sampler description module 211 defines how to apply the texture map to the object's surface. The texture description module 210 can be used to specify the sampler description module 212 and the texture mapping description module 212. The texture mapping description module 212 implements the addition of textures, while the texture mapping description module 212 uses URIs for identification and indexing, and accesses data using the accessor description module 204. The sampler description module 211 implements the specific adjustments and wrapping of the texture. The definitions of the syntax elements in the texture description module 210 are shown in Table 6 below:
[0081] Table 6
[0082]
[0083] In some embodiments, the definitions of the syntax elements in sample(texture.sample) of the texture description module 210 are shown in Table 7 below:
[0084] Table 7
[0085]
[0086] For example, the following is a JSON example of a material description module 209, a texture description module 210, a sampler description module 211, and a texture map description module 212:
[0087]
[0088]
[0089] like Figure 2 As shown, the scene description file may include an animation description module (animation) 213. Figure 2 The animation description module 213 in the scene description file is the next-level description module after the node description module 202. It can be used to describe the animation information added to the object described by the node description module 202. To prevent the object represented by the node description module 202 from being confined to a static state, animation can be added to the object described by the node description module 202. Therefore, the description level of the animation description module 213 in the scene description file is specified by the node description module 202; that is, the animation description module 213 is the next-level description module of the node description module 202. The animation description module 213 also has a corresponding relationship with the mesh description module 203. The animation description module 213 can describe animations through position movement, angle rotation, and size scaling, and can also specify the start and end times of the animation as well as the implementation method. For example, adding an animation to a mesh description module 203 representing a 3D object allows the 3D object represented by the mesh description module 203 to complete the specified animation process within a specified time window through the fusion of position movement, angle rotation, and size scaling.
[0090] like Figure 2 As shown, the scene description file may include a skin description module (skin) 214. Figure 2In the scene description file shown, skin description module 214 is the next level description module after node description module 202. It can be used to describe the motion cooperation relationship between the skeleton added to the node described by node description module 202 and the mesh representing the surface information of the object. When the node described by node description module 202 represents an object with a large degree of freedom of motion, such as a person, animal, or machine, skeletons can be filled into the object to improve the motion performance of these objects. The 3D mesh representing the surface information of the object is conceptually called skin. The description level of skin description module 214 is specified by node description module 202, that is, skin description module 214 is the next level description module of node description module 202, and skin description module 214 has a corresponding relationship with mesh description module 203. By using the movement of the skeleton to drive the movement of the mesh on the surface of the object, and combining it with biomimetic design, a relatively realistic motion effect can be achieved. For example, when a person makes a fist, the skin on the surface will stretch and cover as the internal skeleton changes. At this time, the skeleton pre-filled in the hand model and the cooperative relationship between the skeleton and the skin can be defined to achieve a realistic simulation of this action.
[0091] The scene description files in the aforementioned glTF 2.0 scene description standard only possess basic capabilities for describing 3D objects, and have limitations such as not supporting dynamic 3D immersive media, audio files, and scene updates. glTF 2.0 also declares an optional extension property under each object attribute, allowing for expansion in any part to achieve more complete functionality. This includes scene description modules (scene), node description modules (node), mesh description modules (mesh), accessor description modules (accessor), buffer description modules (buffer), animation description modules (animation), and their internally defined syntax elements, all of which include extension object properties to support functional extensions based on glTF 2.0.
[0092] Currently, mainstream immersive media includes point clouds, 3D meshes, 6DoF panoramic video, and MPEG Immersive Video (MIV). In 3D scenes, multiple types of immersive media often coexist. This requires display engines to support the encoding and decoding of various types of immersive media. Different types of display engines have emerged based on the types and number of supported codecs. Different vendors' display engines support different media types. To achieve cross-platform description of 3D scenes composed of different types of media, the Moving Picture Experts Group (MPEG) initiated the development of the MPEG scene description standard, ISO / IEC 23090-14. This standard primarily addresses the cross-platform description problem of MPEG media (including MPEG-defined codecs, MPEG file formats, and MPEG transport mechanisms) in 3D scenes.
[0093] The MPEG#128 meeting resolved to establish the MPEG Scene Description standard (MPEG-IScene Description standard) based on glTF2.0 (ISO / IEC 12113). Building upon the first version, the MPEG Scene Description standard adds extensions to address unmet needs in cross-platform 3D scene description, including interactivity, AR anchoring, user representation / avatar, haptic support, and expanded support for immersive media codecs.
[0094] The first version of the MPEG scene description standard mainly defines the following:
[0095] 1) The MPEG scene description standard defines a scene description file format that can be used to describe immersive 3D scenes. This format combines the content of the original glTF2.0 (ISO / IEC 12113) and makes a series of extensions on it.
[0096] 2) The MPEG scene description defines a scene description framework and its application programming interface (API) for inter-module collaboration. This decouples the acquisition and processing of immersive media from the media rendering process, facilitating optimizations such as adapting immersive media to different network conditions, partially acquiring immersive media files, accessing different levels of detail within immersive media, and adjusting content quality. This decoupling of immersive media acquisition and processing from immersive media rendering is crucial for achieving cross-platform 3D scene description.
[0097] 3) The MPEG scene description proposes a series of extensions based on the International Standardization Organization Base Media File Format (ISOBMFF) (ISO / IEC 14496-12) that can be used to transmit immersive media content.
[0098] Reference Figure 3 As shown, in Figure 2 Based on the scene description file in the glTF2.0 scene description standard, the extensions of the scene description file in the MPEG scene description standard include: MPEG media (MPEG_media) 301. MPEG media 301 is an independent extension that can be used to reference external media sources.
[0099] Reference Figure 3 As shown, in Figure 2 Based on the scene description file in the glTF2.0 scene description standard, the extensions to the scene description file in the MPEG scene description standard include: MPEG time-varying accessor (MPEG_accessor_timed) 302. MPEG time-varying accessor 302 is an extension of the accessor level and can be used to access time-varying media.
[0100] Reference Figure 3 As shown, in Figure 2 Based on the scene description file in the glTF2.0 scene description standard, the extensions to the scene description file in the MPEG scene description standard include: MPEG Circular Buffer (MPEG_buffer_circular) 303. The MPEG Circular Buffer is a buffer-level extension that can be used to support circular buffers.
[0101] MPEG media (MPEG_media) 301, MPEG time-varying accessor (MPEG_accessor_timed) 302, and MPEG buffer (MPEG_buffer_circular) 303 provide basic descriptions and formats of media in a scene, meeting the basic requirements for describing time-varying immersive media within a scene description framework.
[0102] Reference Figure 3 As shown, in Figure 2Based on the scene description file in the glTF2.0 scene description standard, the extensions to the scene description file in the MPEG scene description standard include: MPEG_scene_dynamic 304. MPEG_scene_dynamic 304 is a scene-level extension that can be used to support dynamic scene updates.
[0103] Reference Figure 3 As shown, in Figure 2 Based on the scene description file in the glTF2.0 scene description standard, the extensions to the scene description file in the MPEG scene description standard include: MPEG texture (MPEG_texture_video) 305. MPEG texture (MPEG_texture_video) 305 is a texture level extension that can be used to support textures in video format.
[0104] Reference Figure 3 As shown, in Figure 2 Based on the scene description file in the glTF2.0 scene description standard, the extensions to the scene description file in the MPEG scene description standard include: MPEG audio space (MPEG_audio_spatial) 306. MPEG audio space (MPEG_audio_spatial) 306 is an extension at the node level and camera level, which can be used to support spatial 3D audio.
[0105] Reference Figure 3 As shown, in Figure 2 Based on the scene description file in the glTF2.0 scene description standard, the extensions to the scene description file in the MPEG scene description standard include: MPEG viewport_recommended 307. MPEG viewport_recommended 307 is a scene-level extension that can be used to support describing recommended viewpoints in two-dimensional displays.
[0106] Reference Figure 3 As shown, in Figure 2 Based on the scene description file in the glTF2.0 scene description standard, the extensions to the scene description file in the MPEG scene description standard include: MPEG mesh linking 308. MPEG mesh linking 308 is a mesh hierarchy extension that can be used to support linking two meshes and provide mapping information.
[0107] Reference Figure 3 As shown, in Figure 2Based on the scene description file in the glTF2.0 scene description standard, the extensions to the scene description file in the MPEG scene description standard include: MPEG animation timing (MPEG_animation_timing) 309. The MPEG animation timing (MPEG_animation_timing) 309 scene-level extension can be used to support control of the animation timeline.
[0108] The following sections will elaborate on each of the above extensions:
[0109] The MPEG media in the MPEG scene description file can be used to describe the type of media file and provide necessary information about MPEG-type media files for subsequent use. The definitions of the first-level syntax elements of MPEG media are shown in Table 8 below:
[0110] Table 8
[0111]
[0112] The definitions of the syntax elements in the media list (MPEG_media.media) of MPEG media are shown in Table 9 below:
[0113] Table 9
[0114]
[0115] The definitions of the syntax elements in the alternatives list (MPEG_media.alternatives) of the media list of MPEG media are shown in Table 10 below:
[0116] Table 10
[0117]
[0118] The definitions of the syntax elements in the track array (MPEG_media.alternatives.tracks) of the media list of MPEG media are shown in Table 11 below:
[0119] Table 11
[0120]
[0121] For example, MPEG media 301 based on the scene description files in Tables 8 to 11 above can be as follows:
[0122]
[0123] Furthermore, based on ISOBMFF (ISO / IEC 14496-12), ISO / IEC 23090-14 also defines the transmission formats for the delivery of scene description files and data related to glTF 2.0 extensions. To facilitate the delivery of scene description files to clients, ISO / IEC 23090-14 defines how to encapsulate glTF files and related data as both time-invariant and time-varying data (e.g., as track samples) within ISOBMFF files. MPEG_scene_dynamic, MPEG_mesh_linking, and MPEG_animation_timing provide specific forms of time-varying data to the display engine, which should then operate accordingly based on this changing information. ISO / IEC 23090-14 also defines the format of the time-varying data for each extension and how to encapsulate it within ISOBMFF files. ISOBMFF is a container format for storing media data, supporting two data storage methods: 1) Track samples; 2) Items. The Samples of a Track storage method is suitable for data that needs to be synchronized with the timeline of the media file, such as animation keyframes or time-related metadata. A track stores one type of media data, such as video data, audio data, or metadata. A sample is the basic unit of media data, and each sample typically contains a media data fragment and timestamp information. Samples are arranged chronologically in the track and played together with other media types such as video and audio. The Items storage method is suitable for storing data that is not directly associated with the timeline. That is, time-varying media that needs updating is stored in samples of the track, while data not directly associated with the timeline is stored in the `item`. MPEF media (MPEG_media) allows referencing external media streams transmitted via protocols such as RTP / SRTP and MPEG-DASH. To allow addressing of media streams without knowing the actual protocol scheme, hostname, or port value, ISO / IEC 23090-14 defines a new URL scheme. This scheme requires a stream identifier in the query portion, but does not specify a particular type of identifier, allowing the use of the Media Stream Identification scheme (RFC5888), the labeling scheme (RFC4575), or a zero-based indexing scheme.
[0124] II. Display Engine
[0125] Reference Figure 1As shown, in the workflow of the scene description framework for immersive media, the display engine 11 can include acquiring a scene description file, parsing the acquired scene description file to obtain the composition structure and detailed information of the 3D scene to be rendered, and rendering and displaying the 3D scene to be rendered based on the information obtained from parsing the scene description file. In some embodiments of this application, the specific workflow and principles of the display engine 11 are not limited, but are defined as follows: the display engine 11 can parse the scene description file, issue instructions to the media access function 12 through the media access function API, issue instructions to the cache management module 13 through the cache API, and retrieve processed data from the cache to complete the rendering and display of the 3D scene and its objects.
[0126] III. Media Access Function API and Media Access Functions
[0127] In the workflow of the immersive media scene description framework, the display engine 11 can obtain the method for rendering the 3D scene by parsing the scene description file, and needs to pass the method for rendering the 3D scene to the media access function 12 or send instructions to the media access function 12 based on the method for rendering the 3D scene. The process of passing the method for rendering the 3D scene to the media access function 12 or sending instructions to the media access function 12 based on the method for rendering the 3D scene is implemented through the media access function API.
[0128] In some embodiments, the display engine 11 can send media access instructions or media data processing instructions to the media access function 12 via the media access function API. The instructions issued by the display engine 11 originate from the parsing results of the scene description file, and may include the address of the media file, media file attribute information (media type, codec used, etc.), and format requirements for the processed media data and other media data.
[0129] In some embodiments, the media access function 12 may also request media access instructions or media data processing instructions from the display engine 11 through the media access function API.
[0130] In the workflow of the immersive media scene description framework, the media access function 12 can receive instructions from the display engine 11 and complete the media file access and processing functions according to the instructions sent by the display engine 11. Specifically, this may include: after obtaining the media file, processing the media file to convert it into media data in the format specified by the scene description file. The processing procedures for different types of media files vary significantly. In order to achieve broad media type support and also considering the working efficiency of the media access function, multiple pipelines are designed in the media access function, and only the pipeline matching the media type is activated during processing.
[0131] The pipeline's input is media files downloaded from the server or read from local storage controls. These media files often have complex structures and cannot be directly used by the display engine 11. Therefore, the pipeline's main function is to process the data of such media files to make the data of the media files conform to the requirements of the display engine 11.
[0132] After the media access function completes the processing of the media file data, it also needs to deliver the processed data to the display engine in a standardized arrangement structure. This requires the processed data to be correctly stored in the cache, which is done by the cache management module 13. However, the cache management module needs to obtain cache management instructions from the media access function or the display engine through the cache API.
[0133] In some embodiments, the media access function can send cache management instructions to the cache management module via the cache API. These cache management instructions are sent by the display engine 11 to the media access function 12 via the media access function API.
[0134] IV. Cache API and Cache Management Module
[0135] In the workflow of the immersive media scene description framework, the media data processed by the pipeline needs to be delivered to the display engine 11 in a standardized arrangement structure. This requires the participation of the cache API and the cache management module 13. The cache API and cache management module create corresponding caches based on the format of the processed media data and are responsible for subsequent cache management, such as updates and releases. The cache management module 13 can communicate with the media access function 12 through the cache API, and it can also communicate with the display engine 11 through the cache API. The goal of communicating with the display engine 11 and / or the media access function 12 is to manage the cache. When the cache management module 13 communicates with the media access function 12, the display engine 11 needs to send the relevant cache management instructions to the media access function 12 through the media access function API. The media access function 12 then sends the relevant cache management instructions to the cache management module 13 through the cache API. When the cache management module 13 communicates with the display engine 11, the display engine 11 only needs to send the cache management description information parsed from the scene description file directly to the cache management module 13 through the cache API.
[0136] With the rapid development of immersive media technologies and the increasing acceptance of online work methods, immersive media technologies are becoming more widespread in people's work and lives. Compared to relying on large display screens to enhance immersion, Virtual Reality (VR) / Augmented Reality (AR) devices have the advantage of conforming to human physiological structure and can create vivid and realistic scenes through binocular stereoscopic vision. Therefore, VR / AR devices are more likely to serve as the hardware foundation for immersive media in the future. A part of VR / AR technology is user representation, or simply avatar. When users wear VR / AR devices, they expect their real or virtual image to appear in the virtual scene presented by the headset and to move synchronously with their movements in the virtual scene. This is an important function of avatars and a crucial element in enhancing immersion.
[0137] In the field of computer graphics (CG-based), a digital human is represented by the static geometric topology of a 3D mesh model, consisting of geometric points and a skeleton. The dynamic attributes of a digital human can be divided into two categories: limb movements, such as walking, running, jumping, turning, and picking up objects; and facial expression movements, such as smiling, frowning, and widening eyes. Limb movements have lower requirements for precision, while facial expression movements require far greater accuracy and complexity.
[0138] In the glTF 2.0-based 3D media resource format, animation is defined and implemented through syntax elements in the animation description module.
[0139] Here is an example JSON sample of an animation description module:
[0140]
[0141] The animation description module uses the node index syntax element and its value ("node":1) to indicate that the node to be animated is the node described by the second node description module in the node list. It uses the path syntax element and its value ("path":"translation") to indicate that the animation method is translation. It uses the syntax elements in "samplers" to indicate the keyframe data of the animation. Specifically, it uses the input syntax element and its value ("input":4) to indicate the input keyframe, uses the interpolation syntax element and its value ("interpolation":"LINEAR") to indicate the interpolation method of the keyframe, and uses the output syntax element and its value ("output":5) to indicate the interpolation data at the corresponding time.
[0142] The dynamic attributes of limbs are achieved through skinning to animate the body. Skinning establishes the relationship between skeletal points (joints) and skin vertices. Each skeletal point is associated with many skin vertices, and the influence weight of each skin vertex is different. Through this relationship, the overall movement of the digital human can be controlled by controlling the rotation vectors of the skeletal points. That is, in the glTF 2.0-based 3D media resource format, by setting the animation type of "path" in "target" to "rotation", and finding the rotation information data of the corresponding nodes (joints) to be animated through "samplers", the motion information of the skeleton is obtained. Then, through skinning technology, the skeleton drives the skin animation, thereby realizing the animation of the digital human.
[0143] For example, refer to Figure 4 As shown, bone point 41 is associated with skin vertices 42 to 49. When the rotation vector of bone point 41 is rotation1 = [0.0, 0.0, 0.0, 0.1], the rotation vectors rotation2 to rotation9 of skin vertices 42 to 49 can be calculated based on the association between bone point 41 and skin vertices 42 to 49. Then, skin vertices 42 to 49 are rotated according to the rotation vectors rotation2 to rotation9 of skin vertices 42 to 49 respectively.
[0144] Facial dynamic attributes refer to facial animation-related attributes such as deformation and expression changes. Because facial skin changes are subtle and detailed, animate them using the same skeletal skinning method as the body, which would be redundant and cumbersome. Therefore, CG tools and glTF 2.0 both support animation based on facial blend shapes. A blend shape refers to a deformation relative to a base shape, usually represented as vertex displacements. Multiple blend shapes can be defined for the face, each of which can be considered an "expression base." By combining different blend shapes and their corresponding weight parameters, facial skin deformation can be achieved. That is, in a glTF 2.0-based 3D media resource format, by setting the "path" animation type in the "target" to "weights," and using "samplers" to find the blend shape weights keyframe data of the mesh under the node to be animated, facial animation is achieved by combining the weights of each blend shape. For example, for a smiley face (morph target 1) and a normal face (base shape), an animation of the transition between the smiley face and the normal face can be generated by generating a blend shape.
[0145] For example, refer to Figure 5 As shown, based on the weights of facial expressions 51 and 52, a transition animation 53 between facial expressions 51 and 52 can be generated using a blend shape. Similarly, based on the weights of facial expressions 51 and 54, a transition animation 55 between facial expressions 51 and 54 can also be generated using a blend shape.
[0146] Another facial animation method supported by CG tools and glTF 2.0 is facial landmarks. Facial landmarks typically refer to key feature points in a facial image, such as the positions of the eyes, nose, and mouth. The positions of these key points can be represented by a set of coordinates, the number of which depends on the specific task. The image shows Morgan_landmarks as defined in Annex H (Annex H) of ISO / IEC 23090-14Amd2. Figure 6As shown, the Morgan_landmarks defined in Annex H of ISO / IEC 23090-14Amd2 includes 68 facial landmarks, with index numbers ranging from 0 to 67. The semantic meaning of a facial landmark can be determined based on its index number. For example, if the index number of a facial landmark is 36, its semantic meaning is the right corner of the right eye. As another example, if the index number of a facial landmark is 14, its semantic meaning is the left earlobe.
[0147] Currently, the method of reconstructing dynamic facial expressions of digital humans using facial markers is quite prominent in the relevant technical field. The overall process of the facial marker method may include the following steps (1) to (4):
[0148] (1) Predefine a neutral digital human face 3D model and predefine a set of facial markers on the model.
[0149] The neutral digital human facial 3D model can be a static 3D model, whose topology can be a 3D mesh or a 3D point cloud, etc. Facial markers typically cover spatial locations with rich variations and significant impact on facial expressions, such as the eyes, eyebrows, mouth, nose, and jaw. Currently, there is no uniform rule for the total number of facial markers, but 68 points are commonly used. Of course, 21, 29, 98, 106, and 186 points can also be used. The larger the total number of facial markers, the more accurate the facial markers will be, but the corresponding data volume and computational load will also be greater. Therefore, in practical applications, the number of facial markers can be set according to the accuracy requirements of the facial 3D model and the performance of the relevant equipment.
[0150] (2) Obtain the input data at the target time and obtain a set of motion vectors of facial markers based on the input data.
[0151] The types of input data and the methods of processing it are highly diverse. For example, when the input data is depth video captured by an RGBD camera, the motion vectors of facial markers can be obtained by using computer vision and neural network methods to locate the facial markers. For instance, using a facial marker extraction network, the spatial coordinates of the facial markers can be obtained, and then the difference between these coordinates and those from previous moments can be used to obtain the motion vectors of the facial markers at the target moment. When the input data is text representing the emotions of a digital human, the motion vectors of the facial markers can be obtained by using a facial expression blending model. This involves first using a weighted superposition of basic expressions to obtain a realistic complex expression, and then locating the facial markers and calculating the motion vectors on this complex face.
[0152] (3) Using the motion vectors of this set of facial markers, obtain the motion vectors of all vertices in the digital human face 3D model.
[0153] When considering the amount of data and the number of points, facial markers are equivalent to sparse data, while all vertices of the static 3D model of a digital human face are equivalent to dense data. The mapping of motion vectors from facial markers to all vertices of the facial model is a mapping from sparse data to dense data. This step can use computer vision and neural network methods, such as using motion diffusion networks, to map the motion vectors of sparse data to the motion vectors of dense data.
[0154] (4) Superimpose the motion vectors of all vertices onto the neutral 3D model of the digital human face to obtain the data of the digital human's dynamic facial expression at the target time.
[0155] The process of superimposing vertex motion vectors onto a neutral 3D digital facial model varies depending on the topology of the static 3D model. When the topology is a 3D point cloud, the motion vector of each point is directly summed with the spatial coordinates of the corresponding point in the 3D point cloud, and the resulting spatial coordinates are the superimposed result. Additionally, the attribute information attached to the 3D point cloud needs to be mapped; that is, the spatial coordinates of the points in the point cloud change, but the attribute information remains unchanged. When the topology is a 3D mesh, the motion vector of each point is directly summed with the spatial coordinates of the corresponding vertex in the 3D mesh, and the resulting spatial coordinates are the superimposed result. Furthermore, the attribute information attached to the 3D mesh needs to be mapped one-to-one; the attribute information remains unchanged, and the connection relationships between vertices remain unchanged. That is, vertices that were previously connected still have the same connection relationships after coordinate transformation.
[0156] Through the above steps (1) to (4), the data of the digital human dynamic facial expression at the target time can be obtained. By continuously repeating the above process, the data of the digital human dynamic facial expression at each time can be obtained. By using the data at different times to drive the dynamic neutral digital human facial 3D model, the reconstruction and mapping of the digital human dynamic facial expression can be realized.
[0157] In some embodiments, after step (4) above, the dynamic facial expressions of the neutral digital human can also be mapped onto the stylized digital human face to make the visual experience more realistic.
[0158] In the scene description framework of some embodiments of this application, the 3D scene service provider needs to provide a scene description file to the display engine, which then parses the scene description file. This scene description file includes, but is not limited to: description information of the entire 3D scene, description information of the digital human in the 3D scene, and information such as the location and index of facial marker points of the digital human in the 3D scene.
[0159] Furthermore, facial landmarks are used for digital human animation and related processing, and hold an important position in multiple fields. For example... Figure 7 As shown, Figure 7 This is a schematic diagram illustrating a scenario for implementing digital human animation and related processing based on facial markers. For example... Figure 7 As shown, applications of digital human animation and related processing based on facial markers can include: 71. Face Tracking: Accurate feature localization helps extract facial descriptive features, improving recognition accuracy. They can also be used for face alignment, ensuring that face images are standardized during preprocessing. 72. Face Retargeting: Key feature points of the original and target faces, such as the positions of eyes, nose, and mouth, are located using facial marker algorithms. Using the located facial markers, the deformation mapping relationship between the source and target faces can be calculated. After determining the deformation mapping relationship, the texture (i.e., pixel information) of the source face can be mapped onto the geometry of the target face. 73. Face Alignment: Face alignment is a technique for locating the coordinates of key facial features on a face. Its goal is to accurately align and standardize the positions of facial features in a face image. It can also enable downstream tasks such as "makeup" 731, "adding props" 732, and "expression recognition and modification" 733. 74. Face Animation: Establishes the association between facial markers and the vertices of the 3D mesh model. Based on the new positions of the facial markers, a deformation field is calculated. This deformation field describes the mapping relationship from the original 3D mesh model to the deformed 3D mesh model. Therefore, the deformation of the 3D mesh model is obtained from the change in the facial markers. 75. Others: Other application scenarios for digital human animation and related processing based on facial markers.
[0160] Reference Figure 8 As shown, Figure 8 This is a schematic diagram illustrating the structure of scene description files in some embodiments of this application. For example... Figure 8As shown, the scene description file includes, but is not limited to, the following modules: MPEG media (MPEG_media) 801, scene description module (scene) 802, node description module (node) 803, mesh description module (mesh) 804, accessor description module (accessor) 805, buffer slice description module (bufferView) 806, buffer description module (buffer) 807, camera description module (camera) 808, material description module (material) 809, texture description module (texture) 810, sampler description module (sampler) 811, texture mapping description module (image) 812, skin description module (skin) 813, and animation description module (animation) 814. The node description module (node) 803 may include: a digital human node array ("MPEG_node_avatar":{}) 8031, which may include: an active identifier syntax element ("isAvatar") 80311, a digital human type syntax element ("type") 80312, and a mapping list ("mappings":[]) 80313. The mapping list ("mappings":[]) 80313 may include: a part name syntax element ("path") 803131 and a node index syntax element ("node") 803132. The following sections will discuss... Figure 8 The various parts of the scenario description file will be explained.
[0161] like Figure 8 As shown, the scene description file may include MPEG media (MPEG_media) 801. Figure 8 The MPEG media (MPEG_media) 801 in the scene description file shown can be used to describe the type of media file and provide necessary information about MPEG-type media files for subsequent use. Media files that include facial marker data can be described in the MPEG media.
[0162] like Figure 8 As shown, the scene description file may include a scene description module (scene) 802. Figure 8 The scene description module (scene) 802 in the scene description file shown can be used to describe the 3D scenes contained in the scene description file. A scene description file may contain any number of 3D scenes, and each 3D scene is represented by a separate scene description module.
[0163] like Figure 8As shown, the scene description file may include a node description module (node) 803. The node description module (node) 803 may include: a digital human node array ("MPEG_node_avatar":{}) 8031.
[0164] In some embodiments of this application, a digital human in a three-dimensional scene will be represented by a node, described by a corresponding node description module, and multiple levels of child nodes will be attached under the node. The specific parts of the digital human's body will be described layer by layer by the node description module corresponding to the child nodes. Through the hierarchical structure, the deepest unit in the upper body can be each finger joint. For example: Using the buttocks as a node to represent a digital human (layer 1), the node representing the digital human includes three child nodes: spine, left thigh, and right thigh (layer 2); selecting the child node representing the spine and continuing to extend it deeper, the child node representing the spine includes a child node representing the chest (layer 3); selecting the child node representing the chest and continuing to extend it deeper, the child node representing the chest includes a child node representing the upper chest (layer 4); selecting the child node representing the upper chest and continuing to extend it deeper, the child node representing the upper chest includes three child nodes representing the left shoulder, right shoulder, and neck (layer 5); selecting the child node representing the neck and continuing to extend it deeper, it includes a child node representing the head (layer 6).
[0165] In some embodiments, the hierarchical standard names for digital human body parts may be as follows:
[0166]
[0167]
[0168]
[0169] In the node and child node hierarchy describing the specific parts of the digital human's body, each node and child node may have one or more 3D meshes attached to it, serving as 3D models of these body parts. Furthermore, using the displacement, rotation, scaling, and other parameters of the nodes and child nodes, these 3D models of body parts can be realistically pieced together to form a complete 3D model of the digital human.
[0170] Whether a node description module includes the extension "MPEG_node_avatar":{}" indicates whether the node corresponding to that node is used to represent a digital human. Specifically, if a node description module in a scene description file includes the extension "MPEG_node_avatar":{}", then it is determined that the node corresponding to that node description module represents a digital human.
[0171] It should be noted that in the scene description file, a digital human is represented by a node, and multiple levels of child nodes can be attached to this node. The specific parts of the digital human's body are described layer by layer by the node description module corresponding to the child nodes. In order to avoid information redundancy, in some embodiments of this application, only the digital human node array ("MPEG_node_avatar":{}) is added to the node description module corresponding to the top-level node representing the digital human, and the digital human node array is not added to the node description module corresponding to the child nodes attached to it.
[0172] In some embodiments, when a node represents a digital human, an array of digital human nodes ("MPEG_node_avatar":{}) can be added to the extension list ("extensions":{}) of the node description module corresponding to that node.
[0173] like Figure 8 As shown, the digital human node array ("MPEG_node_avatar":{})8031 in the node description module 803 may include: an active identifier syntax element ("isAvatar")80311, a digital human type syntax element ("type")80312, and a mapping list ("mappings":[])80313.
[0174] The active identifier syntax element ("isAvatar") in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) can be used to indicate whether the digital human represented by the node corresponding to the node description module (node) is active (whether the digital human needs to be rendered during the rendering process, and whether the display engine, media access function, etc. in the scene description box need to process the relevant data and information from the digital human).
[0175] In some embodiments, the method of indicating whether the digital person represented by the node corresponding to the node description module is active by using the active identifier syntax element ("isAvatar") may include: determining that the digital person represented by the node corresponding to the node description module is active when the value of the active identifier syntax element ("isAvatar") is a first value, and determining that the digital person represented by the node corresponding to the node description module is inactive when the value of the active identifier syntax element ("isAvatar") is a second value.
[0176] In some embodiments, the first value and the second value can be true and false, respectively.
[0177] In other embodiments, the first value and the second value may be 1 and 0, respectively.
[0178] In some embodiments, the data type of the value of the active identifier syntax element ("isAvatar") can be boolean.
[0179] The digital human type syntax element ("type") in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) can be used to indicate the representation scheme of the digital human represented by the node corresponding to the node description module (node).
[0180] In some embodiments, the digital human type syntax element ("type") uses a Uniform Resource Name (URN) to describe the representation scheme of the digital human. The representation scheme of a digital human can be understood as the technical solution / architecture to which the digital human belongs. This technical solution / architecture defines the media format / type and other information of all body components of the digital human. Using this information, media access functions can establish corresponding pipelines for the reconstruction / driving of the corresponding digital human body components, or for the reconstruction / driving of the entire digital human body. For example, when the digital human represented by the node corresponding to the node description module (node) is an MPEG reference digital human, the value of the digital human type syntax element ("type") can be set to the Uniform Resource Name of the MPEG reference digital human. The URN of the MPEG reference digital human is urn:mpeg:sd:2023:avatar. Therefore, when the digital human represented by the node corresponding to the node description module (node) is an MPEG reference digital human, the digital human type syntax element and its value are: "type":"urn:mpeg:sd:2023:avatar".
[0181] In some embodiments, the data type of the value of the digital human type syntax element ("type") can be string.
[0182] like Figure 8As shown, the mapping list ("mappings":[]) 80313 may include: a part name syntax element ("path") 803131 and a node index syntax element ("node") 803132. The mapping list ("mappings":[]) in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) describes the mapping between digital human child nodes and the hierarchical standard names of digital human body parts. The data type of the mapping list ("mappings":[]) is an array containing two syntax elements: a part name syntax element ("path") and a node index syntax element ("node"). The part name syntax element ("path") can be used to describe the standard name of the digital human body part corresponding to the node / child node of the node description module, such as " / full_body / upper_body / head / face / eye_right", etc. The node index syntax element ("node") can be used to index the nodes that describe the body parts identified by the part name syntax element ("path"), so that the body part name corresponds to the node.
[0183] In summary, the syntax elements in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) are shown in Table 12 below:
[0184] Table 12
[0185]
[0186] The syntax elements in the mapping list ("mappings":[]) of the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) are shown in Table 13 below:
[0187] Table 13
[0188]
[0189] In some embodiments, the syntax elements in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) may not include the active identifier syntax element ("isAvatar"). When the syntax elements in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) do not include the active identifier syntax element ("isAvatar"), the syntax elements in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) are shown in Table 14 below:
[0190] Table 14
[0191]
[0192] Figure 9 The scene description file includes the definition of a digital human node array ("MPEG_node_avatar":{}) and provides a related reference model. The node description module, including the digital human node array and the reference model, provides a standardized way to define the geometry and semantics of the digital human's geometric description. Furthermore, it can also include definitions of interaction between the digital human and the scene. However, it currently does not describe how to animate and encode the digital human; the animation of the digital human is based on the animation description module in the scene description file. The animation description module indexes the node description module through internal syntax elements to define combinations of local translation, rotation, and scaling transformations for the corresponding nodes, transforming or defining the weights of the node's morph targets.
[0193] Currently, relevant technical standards define the facial landmarks of the default digital human model in scene description files, and indicate that these landmarks can be used to implement digital human animation and related processing. However, scene description files do not declare information related to facial landmarks, therefore, scene description frames cannot currently be used to implement digital human animation and related processing based on facial landmarks. In view of this, some embodiments of this application further propose the following technical solutions:
[0194] Add a face marker array ("landmarks":{}) to the digital human node array ("MPEG_node_avatar":{}) in the node description module (node), and define the level type of the face markers, the accessor for accessing the location coordinates, the accessor for accessing the index information, etc. through the contents of the face marker array.
[0195] In some embodiments, the data type of the contents of the facial marker array ("landmarks":{}) is a data object.
[0196] In some embodiments, the use of the facial marker array ("landmarks":{}) is not mandatory.
[0197] In some embodiments, the array of facial markers ("landmarks":{}) has no default value.
[0198] Based on the above scheme, the syntax elements in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) are shown in Table 15 or Table 16 below:
[0199] Table 15
[0200]
[0201] Table 16
[0202]
[0203] In some embodiments, the face marker array ("landmarks":{}) includes a level type syntax element ("level_type"), and the number and type of face markers are declared through the level type syntax element ("level_type").
[0204] In some embodiments, the data type of the value of the level type syntax element ("level_type") is integer.
[0205] In some embodiments, the use of the level type syntax element ("level_type") is optional.
[0206] In some embodiments, the default value of the level type syntax element ("level_type") is 0.
[0207] In some embodiments, the face marker array ("landmarks":{}) further includes an index syntax element ("indices"), and declares the index value of the accessor description module corresponding to the accessor used to access the index information of the face markers through the index syntax element ("indices"), so that the scene description framework can access the index information of the face markers according to the declaration of the index syntax element, and obtain the semantic information of the face markers according to the index information of the face markers.
[0208] In some embodiments, the data type of the value of the index syntax element ("indices") is integer.
[0209] In some embodiments, the use of the index syntax element ("indices") is mandatory.
[0210] In some embodiments, the index syntax element ("indices") has no default value.
[0211] In some embodiments, the face marker array ("landmarks":{}) further includes a position coordinate syntax element ("position"), and declares the index value of the accessor description module corresponding to the accessor used to access the index information of the face markers through the position coordinate syntax element ("position"), so that the scene description framework can access the position coordinates of the face markers according to the declaration of the position coordinate syntax element.
[0212] In some embodiments, the data type of the value of the position syntax element ("position") is integer.
[0213] In some embodiments, the use of the position syntax element ("position") is mandatory.
[0214] In some embodiments, the position syntax element ("position") has no default value.
[0215] Based on the above scheme, the syntax elements in the face marker array ("landmarks":{}) of the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) may include one or more of the elements in Table 17 below:
[0216] Table 17
[0217]
[0218] In some embodiments, the correspondence between the value of the level type syntax element ("level_type") and the number and type of facial markers can be shown in Table 18 below:
[0219] Table 18
[0220]
[0221] Reference Figure 9 As shown, Figure 8 This is a schematic diagram illustrating the structure of a scene description file that supports the definition of facial marker points, provided for some embodiments of this application. Figure 8 Based on the scenario description file shown. Figure 8 The scene description file shown further includes:
[0222] The facial marker array ("landmarks":{})80314 is set in the digital human node array ("MPEG_node_avatar":{})8031 in the node description module (node) 803. The facial marker array ("landmarks":{})80314 may include: a level_type syntax element ("level_type")803141, an index syntax element ("indices")803142, and a position coordinate syntax element ("position")803143.
[0223] The description method of the mesh description module (mesh) 804 in the scene description file can include: for nodes representing 3D objects, the scene description file will use the data format of a 3D mesh, that is, use the mesh description module (mesh) 804 for description. The mesh description module contains various specific information, such as 3D coordinates and color information, all of which belong to the syntax elements in the attribute (mesh.primitives.attribute) of the mesh description module 804 primitive. The most basic mesh description module includes the 3D coordinates of vertices, the connection relationships between vertices, the color information of vertices, and the 3D coordinates that require texture values. At the mesh description module level, the required 3D coordinate set, texture coordinates (Texcoord), needs to be enabled so that the haptic values obtained from the texture can be attached to the corresponding 3D coordinates of the 3D object. Texture coordinates (Texcoord) and color data, etc., belong to the attribute (mesh.primitives.attribute) of the mesh description module primitive.
[0224] In some embodiments, the description methods of the accessor description module 805, the buffer slice description module 806, and the buffer description module 807 in the scene description file may include: a part used to describe the output data of the media access function, a part used to index the input data of the media access function (such as a media file including data related to facial marker points), and there is an interleaving relationship between the content used to describe the output data of the media access function and the content used to index the input data of the media access function.
[0225] like Figure 8As shown in Figure 9, the scene description file may include an accessor description module (accessor) 805, a buffer slice description module (bufferView) 806, and a buffer description module (buffer) 807. The method for describing the output data of the media access function through these modules may include: a mesh description module (mesh) 804 representing the dynamic 3D model of a digital human face in the described 3D mesh points to a specific accessor description module (accessor) 805; this accessor description module (accessor) 805 then points to the corresponding buffer slice description module (bufferView) 806; and this buffer slice description module (bufferView) 806 then points to the corresponding buffer description module (buffer) 807. This description method enables the storage and indexing of digital human facial expression model data. Specifically, the buffer described by the buffer description module (buffer) 807 can be directly read by the display engine, and the stored data is data that can be directly used for rendering. In some embodiments of this application, the data stored in the buffer includes the index information and position coordinates of digital human facial markers, as well as a dynamic 3D model of the digital human. The buffer slice description module (bufferView) 806 is responsible for slicing the data in the buffer, which can be achieved using two parameters: the starting byte offset (byteOffset) and the byte length (byteLength). The accessor description module (accessor) is responsible for adding additional information to the data in the buffer slice, such as the data type, the quantity of a certain type of data, and the numerical range of a certain type of data. The mesh description module (mesh) 804, used to describe the 3D mesh representing the digital human face, will point to the accessor description module (accessor) 805 to retrieve the dynamic 3D model of the digital human to be rendered. The facial marker array, which can be used to describe the descriptive information of the digital human facial markers, will point to the accessor description module (accessor) 805 to retrieve the index information and position coordinates of the digital human facial markers.
[0226] Since the data related to digital facial markers is a time-varying medium, the accessor used to access this data needs to be capable of accessing time-varying media. Therefore, the accessor description module (accessor) 805 describing the accessor used to access the data related to digital facial markers should include an MPEG ring buffer ("MPEG_accessor_timed":{}) extension, and use the MPEG ring buffer to transform the described accessor into a time-varying accessor. When the accessor description module (accessor) 805 includes an MPEG ring buffer ("MPEG_accessor_timed":{}), the accessor description module (accessor) 805 contains two buffer slice index syntax elements ("bufferView"), one of which is outside the MPEG ring buffer, and the other is inside the MPEG ring buffer. The buffer slice index syntax element ("bufferView") outside the MPEG circular buffer will point to the relevant data of digital human facial markers or the dynamic 3D model of the digital human. The buffer slice index syntax element ("bufferView") inside the MPEG circular buffer will point to the header of the time-varying accessor. The header of the time-varying accessor, also known as the parameters of the time-varying accessor, describes the indexing method, data type, and data quantity of the time-varying data stored in the time-varying accessor. These parameters change over time and are therefore appended to the time-varying accessor as a header. The header may include: the buffer slice index syntax element ("bufferView") outside the MPEG circular buffer, the data type syntax element ("componentType"), and the value of the accessor type syntax element ("type") at different times.
[0227] Since the data related to digital facial markers is a time-varying medium, the buffer used to cache this data needs to be capable of caching time-varying media. Therefore, the buffer description module (buffer) 807, which describes the buffer used to cache the data related to digital facial markers, should include an MPEG circular buffer extension ("MPEG_buffer_circular":{}) and transform the buffer into a circular buffer using the MPEG circular buffer extension. The MPEG circular buffer ("MPEG_buffer_circular":{}) can contain syntax elements such as the media index syntax element ("media"), the track index syntax element ("tracks"), and the number of stages syntax element ("count"). The media index syntax element ("media") and its value will point to the media file described in the MPEG media. The track index syntax element ("tracks") and its value can be used to describe the track information of the source data of the data cached in the buffer. The number of stages syntax element ("count") and its value can be used to specify the number of storage stages in the circular buffer.
[0228] Through the above two extensions, different stages in the ring buffer store data of the time-varying medium at different times, and the different storage stages are separated in a slice manner by cache slicing. The time-varying accessor will access data from different cache slices (storage stages) at different times.
[0229] In some embodiments, the method of describing the input data of a media access function through an accessor description module (accessor) 805, a buffer slice description module (bufferView) 806, and a buffer description module (buffer) 807 may include: the input data of the media access function may be described as a media file in MPEG media (MPEG_media), and the media file declared in the MPEG media (MPEG_media) may be indexed by an MPEG circular buffer ("MPEG_buffer_circular":{}) in the buffer description module (buffer) 807, thereby delivering the input data of the media access function to the media access function.
[0230] In some embodiments, the method of describing the input data of a media access function through an accessor description module (accessor) 805, a buffer slice description module (bufferView) 806, and a buffer description module (buffer) 807 may include: instead of using a media file to carry the input data of the media access function, a scene description file is used to carry the input data of the media access function. The input data of the media access function will be delivered by the display engine to the media access function through the media access function API after the display engine completes the parsing of the scene description file. One implementation of directly describing the input data of the media access function using a scene description file may include: describing the input data of the media access function in the buffer description module (buffer) 807 of the scene description file.
[0231] In some embodiments, the method of describing the input data of a media access function through an accessor description module (accessor) 805, a buffer slice description module (bufferView) 806, and a buffer description module (buffer) 807 may include: describing a portion of the input data of the media access function as a media file in MPEG media (MPEG_media), and indexing the media file declared in MPEG media (MPEG_media) through an MPEG circular buffer ("MPEG_buffer_circular":{}) in the buffer description module (buffer) 807; and directly describing another portion of the input data of the media access function using a scene description file.
[0232] like Figure 8 As shown in Figure 9, the scene description file may include a camera description module (camera) 808. In some embodiments, the description method of the camera description module (camera) 808 in the scene description file may include: defining viewing-related visual information such as viewpoint and viewing angle of the node description module (node) 803 through the camera description module (camera) 808.
[0233] like Figure 8As shown in Figure 9, the scene description file may include a material description module 809, a texture description module 810, a sampler description module 811, and a texture mapping description module 812. In some embodiments, the description method of the material description module 809 and texture description module 810 in the scene description file, which supports obtaining the tactile values of tactile material properties of a reference type, may include: describing additional information about the surface of a 3D object through the material description module 809, texture description module 810, sampler description module 811, and texture mapping description module 812. The collaborative relationship between the material description module 809, texture description module 810, sampler description module 811, and texture mapping description module 812 may include: the material description module 809 and texture description module 810 jointly define the color and physical information of the object surface. The texture description module 810 specifies the sampler description module 811 and the image description module 812. The sampler description module 811 defines how to map the texture map onto the object surface, implementing specific adjustments and wrapping of the texture. The image description module 812 uses ULRs to identify and index the texture map.
[0234] like Figure 8 As shown in Figure 9, the scene description file may include a skin description module (skin) 813. In some embodiments, the description method of the skin description module (skin) 813 in the scene description file may include: defining the motion and deformation relationship between the 3D mesh mounted by the node description module (node) 813 and the corresponding skeleton through the skin description module (skin) 813.
[0235] like Figure 10 As shown in Figure 9, the scene description file may include an animation description module 814. In some embodiments, the description method of the animation description module 814 in the scene description file may include: defining the animation added by the node description module 803 through the animation description module 814.
[0236] In some embodiments, the animation description module 814 can describe the animation added to the node description module 803 by one or more of position movement, angle rotation, and size scaling.
[0237] In some embodiments, the animation description module 814 may also indicate at least one of the start time, end time, and implementation method of the animation added to the node description module 803.
[0238] That is, in the scene description file provided in some embodiments of this application, animations can also be added to nodes representing objects in a 3D object. The animation description module (animation) 814 describes the animations added to the nodes in three ways: position movement, angle rotation, and size scaling. It can also specify the start and end times of the animation and the implementation method of the animation.
[0239] For example, the scenario description file provided in the above embodiments will be described below with reference to a specific scenario description file.
[0240]
[0241]
[0242]
[0243]
[0244] The curly braces between lines 1 and 150 in the example above contain the main content of the scene description file provided in the above embodiment. The scene description file may include: digital asset description module (asset), extensionUsed list (extensionUsed), MPEG media (MPEG_media), scene declaration (scene), scene list (scenes), node list (nodes), mesh list (meshes), accessor list (accessors), bufferview list (bufferViews), and buffer list (buffers). The content of each part and the information contained in each part from the perspective of parsing are explained below.
[0245] 1. Digital Asset Description Module (asset): The digital asset description module is represented by lines 2-4. The "version": "2.0" in line 3 of the digital asset description module indicates that the scene description file is written based on glTF version 2.0, which is also the reference version for the scene description standard. From a parsing perspective, the display engine can determine which parser to use to parse the scene description file based on the digital asset description module.
[0246] 2. Extension Used: Lines 6-11 show the extension used list. Since the extension used list includes four extensions: MPEG media, MPEG buffer circular, MPEG accessor timed, and MPEG node avatar, it can be determined that the scene description file uses these four MPEG extensions. From a parsing perspective, the display engine can infer in advance from the content of the extension used list that the subsequent parsing involves these four MPEG extensions: MPEG media, MPEG buffer circular, MPEG accessor timed, and MPEG node avatar.
[0247] 3. MPEG Media (MPEG_media): Lines 13-33 define the MPEG media. The MPEG media declares the media files contained in the 3D scene. Line 19, "mimeType":"text / plain", indicates the file type of the container file corresponding to the first media file, and line 20, "uri":"texts / indices.txt", indicates the access address of the first media file. Line 27, "mimeType":"text / plain", indicates the file type of the container file corresponding to the second media file, and line 28, "uri":"texts / indices.txt", indicates the access address of the second media file. From a parsing perspective, the display engine can determine the existence of two media files in the 3D scene to be rendered by parsing the MPEG media, and learn the methods for accessing and parsing these media files.
[0248] 4. Scene Declaration: The scene declaration is on line 35. Since a scene description file can theoretically include multiple 3D scenes, the above scene description file first indicates through "scene":0 on line 35 that the scene described by the scene description file is the first 3D scene in the scene list, that is, the scene described by the scene description module enclosed in the curly braces on lines 38 to 41.
[0249] 5. Scene List (scenes): The scene list is on lines 37-41. The scene list contains only one curly brace, indicating that the scene description file contains one scene description module, and also that the scene description file contains only one 3D scene. Within the curly brace (scene description module), "nodes":[0] on line 39 indicates that the 3D scene contains only one node with index 0. From a parsing perspective, this scene list clarifies that the entire scene description framework should select the first 3D scene in the scene list for subsequent processing and rendering, clarifies the overall structure of the 3D scene, and points to the more detailed next-level node description module.
[0250] 6. Node List (nodes): The node list is located in lines 43-68. The node list contains only one curly brace, indicating that it includes only one node description module (node). This 3D scene has only one node, and this node is the same node as the node with an index value of 0 contained in the scene description module. The two are associated through an index. In this unique node description module, line 45, "mesh":"0", indicates that the 3D mesh attached to the node is the 3D mesh described by the first mesh description module in the mesh list, which corresponds to the mesh description module of the next layer. Line 47, "MPEG_node_avatar", indicates that the node described by this node description module represents a digital human. Line 48, "isAvatar":True, indicates that the digital human represented by the node described by this node description module is active. Line 49, "type":"urn:mpeg:sd:2023:avatar", indicates that the representation scheme of the digital human represented by the node described by this node description module is the MPEG reference digital human represented by urn:mpeg:sd:2023:avatar. Line 52, "path":"full_body / upper_body / arm_left", indicates that the hierarchical standard name of the body part of the digital human in this mapping relationship is the left arm. Line 53, "node":0", indicates that the index value of the node description module corresponding to this mapping relationship is 1. The path in line 56, "path":"full_body / upper_body / arm_right", indicates that the hierarchical standard name of the digital human body part in this mapping relationship is the right arm. The node in line 57, "node":2", indicates that the index value of the node description module corresponding to the node in this mapping relationship is 2. Based on the level type syntax element and its value "level_type":1 in the face marker array ("landmarks":{}), the type and number of face markers can be determined to be the 68 face markers provided in MOGAN. The index syntax element and its value "indices":0 indicate that the index information of the face markers can be accessed through the accessor corresponding to the first accessor description module in the accessor list ("accessors":[]). The position coordinate syntax element and its value "position":1 indicate that the position coordinates of the face markers can be accessed through the accessor corresponding to the second accessor description module in the accessor list ("accessors":[]).From a parsing perspective, this node list indicates that the unique node in the 3D scene is attached to a 3D mesh, and that this 3D mesh is the 3D mesh described by the first mesh description module in the mesh list. This node represents a digital human, and the representation scheme of this digital human is the MPEG Reference Digital Human. The facial markers of this digital human are the 68 facial markers provided in MOGAN. The index information of the digital human's facial markers can be accessed through the accessor corresponding to the first accessor description module in the accessor list, and the position coordinates of the digital human's facial markers can be accessed through the accessor corresponding to the second accessor description module in the accessor list.
[0251] 7. Mesh List: The mesh list is located on lines 70-81. In the mesh description module, line 72, "primitives," indicates that the 3D mesh has primitives. Lines 75, "attributes," and 77, "mode," indicate the presence of attributes and modes within these primitives. Line 75, "position": 0, indicates that the 3D mesh has geometric coordinate data, and the accessor describing the accessor for accessing these coordinates is the first accessor in the accessor list. Furthermore, line 77, "mode": 4, confirms that the 3D mesh topology is a triangular mesh. From an analytical perspective, this description module determines the actual data types and topological type of the 3D mesh.
[0252] 8. Buffer List: The buffer list is located on lines 132-149. The buffer list contains two curly braces, indicating that the scene description file includes two buffer description modules, and the display of this 3D scene requires the use of two buffers. Both buffer description modules use the MPEG circular buffer extension ("MPEG_buffer_circular":{}), indicating that the buffers described by these two buffer description modules are circular buffers obtained using the MPEG extension. According to line 137 of the MPEG circular buffer ("MPEG_buffer_circular":{}), "media:0" indicates that the data source in the first circular buffer is the first media file declared within the MPEG media (MPEG_media), and according to line 145 of the MPEG circular buffer ("MPEG_buffer_circular":{}), "media:1" indicates that the data source in the second circular buffer is the second media file declared within the MPEG media (MPEG_media). From a parsing perspective, the buffer list maps declared media files in MPEG media (MPEG_media) to buffers, or in other words, allows buffers to reference previously declared but unused media files. It's important to note that the media files referenced here are unprocessed encapsulated files. These media files need to be processed by media access functions to extract directly usable rendering data, such as the 3D coordinates mentioned in the mesh description module (mesh) of the mesh list.
[0253] 9. Buffer Slice List (bufferViews): The buffer slice list is located on lines 113-130. Each buffer slice description module contains four parallel curly braces, indicating that the data of the declared media files in the MPEG media is divided into four buffer slices. The first curly brace (the first buffer slice description module) first points to the buffer description module with index 0, i.e., the first buffer description module in the buffer list. Then, the data slice range of the corresponding buffer slice is limited to bytes 0-272 by the two parameters: byte offset (0) and byte length (272). The information provided by the other buffer slice description modules is similar to that provided by the first buffer slice description module, and will not be described in detail here to avoid redundancy.
[0254] 10. Accessors List: The accessors list is located on lines 84-111. The accessors list contains two curly braces, indicating that this scene description file should only include two accessor description modules. The display of this 3D scene requires access to media data through these two accessors. Furthermore, both accessor description modules contain MPEG time-varying accessors ("MPEG_accessor_timed":{}), indicating that both accessors point to MPEG-defined time-varying media. Within the first set of curly braces (the first accessor description module), line 86's "bufferView":0 indicates that the accessor corresponding to this accessor description module should retrieve data from the cache corresponding to the first cache slice description module in the cache slice list ("bufferViews":[]). Line 88's "componentType":5125 indicates that the data type of the accessor corresponding to this accessor description module is an unsigned integer. Line 89's "type":"SCALAR" indicates that the accessor type of the accessor corresponding to this accessor description module is a string. Line 90's "Count":68 indicates that the number of data items the accessor corresponding to this accessor description module needs to access is 68. Line 94's "bufferView":2 indicates that the header information of the accessor corresponding to this accessor description module should be retrieved from the cache corresponding to the third cache slice description module in the cache slice list ("bufferViews":[]). Within the second set of curly braces (the second accessor description module), line 99's "bufferView..." The line `:1` indicates that the accessor corresponding to this accessor description module should obtain data from the cache corresponding to the first cache slice description module in the cache slice list ("bufferViews":[]). Line 101, `componentType":5126`, indicates that the data type of the accessor corresponding to this accessor description module is floating-point. Line 102, `type":"VEC3"`, indicates that the accessor type of the accessor corresponding to this accessor description module is a three-dimensional vector. Line 103, `count":3`, indicates that the number of data items the accessor corresponding to this accessor description module needs to access is 3. Line 107, `bufferView":3`, indicates that the header information of the accessor corresponding to this accessor description module should be obtained from the cache corresponding to the fourth cache slice description module in the cache slice list ("bufferViews":[]). From a parsing perspective, this accessor list completes the definition of the data required for rendering. For example, data types missing in the cache slice description modules are defined in the corresponding accessor description modules in the accessor list.
[0255] II. Display Engine
[0256] In the workflow of the immersive media scene description framework, the display engine's functions may include: parsing the scene description file to obtain the method for rendering the 3D scene; sending media access instructions or media data processing instructions to the media access function via the media access function API; sending cache management instructions to the cache management module via the cache API; retrieving processed data from the cache, and completing the rendering and display of the 3D scene and objects in the 3D scene based on the read data. Therefore, when it is necessary to use facial markers to implement digital human animation and related processing, the display engine's functions may include: 1. being able to parse the scene description file that defines the relevant information of the digital human's facial markers, which may include: parsing the content of the scene description file to determine that the media file containing the relevant data of the digital human's facial markers is timed media, obtaining the media file access address stored in the MPEG circular buffer ("MPEG_buffer_circular":{}), the media file access address pointing to the media file containing the relevant data of the digital human's facial markers, and obtaining the format information of the relevant data of the facial markers from the MPEG circular buffer time-varying accessor ("MPEG_accessor_timed":{}). 2. It can transmit media access commands or media data processing commands through the media access function API and media access functions; these commands originate from the parsing results of a scene description file containing information related to facial marker points. 3. It can send cache management commands to the cache management module through the cache API. 4. It can retrieve processed dynamic 3D models of digital humans from the cache and render and display the 3D scene and the digital human within it based on the retrieved data.
[0257] III. Media Access Function API
[0258] In the workflow of the scene description framework for immersive media, the display engine can obtain the method for rendering the 3D scene by parsing the scene description file. It needs to pass the method for rendering the 3D scene to the media access function or send instructions to the media access function based on the method for rendering the 3D scene. The process of passing the method for rendering the 3D scene to the media access function or sending instructions to the media access function based on the method for rendering the 3D scene is implemented through the media access function API.
[0259] In some embodiments, the display engine can send media access instructions or media data processing instructions to the media access function via the Media Access Function API. These instructions originate from the parsing results of a scene description file containing information about facial markers. The media access instructions or media data processing instructions may include: the index of the media file, the URL of the media file, the type of the media file, the codec used by the media file, and format requirements for the processed digital human's dynamic 3D model and other media data.
[0260] In some embodiments, the media access function may also proactively request media access instructions or media data processing instructions from the display engine through the media access function API.
[0261] IV. Media Access Functions
[0262] In the workflow of the immersive media scene description framework, after receiving media access instructions or media data processing instructions issued by the display engine through the Media Access Function API, the media access function executes these instructions. For example: acquiring the data media file used to reconstruct the digital human's dynamic facial expressions; establishing appropriate pipelines for the reconstructed dynamic 3D model media file of the digital human; and writing the processed dynamic 3D model of the digital human, formatted according to the scene description file, into an appropriate cache. (See reference...) Figure 10 As shown, Figure 10 This is a schematic diagram of an Avatar Media Pipeline provided for some embodiments of this application. For example... Figure 11 As shown, the digital human media pipeline 100 may include: an Avatar LandmarksDecoder 101 and an Avatar Reconstruction module 102. The workflow of the digital human media pipeline 100 includes: after receiving the data format of digital human facial markers sent by the display engine 11 via the media access function API, firstly, initializing the digital human media pipeline 100 to create a pipeline that can process media files corresponding to the facial markers; secondly, obtaining the data stream (Landmarks stream) of the media files corresponding to the facial markers from the cloud server; performing decoding and reconstruction operations through the digital human facial markers decoder 101 and the digital human reconstruction module 102 in the created pipeline; converting the media files corresponding to the facial markers into the required format; and writing the converted data into a cache through the cache management module 13 so that the display engine 11 can read the data and complete the rendering and display of the 3D scene and the digital human in the 3D scene based on the read data.
[0263] V. Supports caching API for defining facial markers
[0264] After the media access function obtains the decoded data of the media file, which includes the index information and position coordinates of facial markers, it needs to write the decoded data of the media file into a cache with a standardized arrangement structure in order to retrieve the index information and position coordinates of the facial markers. After the media access function generates the dynamic 3D model of the digital human through the pipeline, it also needs to deliver the dynamic 3D model of the digital human to the display engine with a standardized arrangement structure. All of these require that the relevant data be correctly stored in the cache. This work is done by the cache management module, but the cache management module needs to obtain cache management instructions from the media access function or the display engine through the cache API.
[0265] In some embodiments, the media access function can send cache management instructions to the cache management module via the cache API. These cache management instructions are those sent by the display engine to the media access function via the media access function API.
[0266] In some embodiments, the display engine can send cache management instructions to the cache management module via the cache API.
[0267] In other words, the cache management module can communicate with the media access function and the display engine through the cache API, and the purpose of communicating with either the media access function or the display engine is to manage the cache. When the cache management module communicates with the media access function through the cache API, the display engine needs to send the cache management instructions to the media access function through the media access function API, and the media access function then sends the cache management instructions to the cache management module through the cache API. When the cache management module communicates with the display engine through the cache API, the display engine only needs to generate cache management instructions based on the cache management information parsed from the scene description file and send them to the cache management module through the cache API.
[0268] In some embodiments, cache management instructions may include one or more of the following: instructions to create a cache, instructions to update a cache, and instructions to release a cache.
[0269] VI. Cache Management Module
[0270] In the workflow of the immersive media scene description framework, after the media access function obtains the decoding data of the media file containing the index information and position coordinates of the facial markers, it needs to write the decoding data of the media file into the cache in a standardized arrangement structure so that the index information and position coordinates of the facial markers can be retrieved. After the media access function completes the reconstruction of the dynamic 3D model of the digital human through the pipeline, the dynamic 3D model of the digital human needs to be delivered to the display engine in a standardized arrangement structure. This requires the dynamic 3D model of the digital human to be correctly stored in the cache. These tasks are handled by the cache management module.
[0271] The cache management module implements management operations such as cache creation, updating, and release. Operation instructions are received through the cache API. Cache management rules are recorded in the scene description file, parsed by the display engine, and ultimately issued to the cache management module by the display engine or media access function. After media files are processed by the media access function, they need to be stored in an appropriate cache before being retrieved by the media access function or display engine. The role of cache management is to effectively manage these caches, ensuring they match the format of the processed media data without disrupting it. The specific design method of the media management module should refer to the design of the display engine and media access function.
[0272] VII. Media files containing data related to digital facial markers
[0273] Based on a digital human representation format that supports digital human models, digital human facial marker data can be extracted from a user's video capture by, for example, training a neural network (facial keypoint detection algorithm). This data can then be transmitted to the receiver as ISOBMFF track samples in a transmission format compatible with scene description files.
[0274] As mentioned above, ISOBMFF's Samples of a Track format is suitable for data that needs to be synchronized with the timeline in the media file; ISOBMFF's Items format can be used to store data that is not directly associated with the timeline, while the data related to digital facial markers needs to change and be updated dynamically and needs to be synchronized with the timeline. Therefore, in some embodiments of this application, ISOBMFF's Samples of a Track format is selected to store and transmit media files that include data related to digital facial markers.
[0275] When using the ISOBMFF Samples of a Track format to store and transmit media files containing data related to digital human facial markers, the ISOBMFF file can include two parts: a facial marker sample header (Landmarks sample entry) and a facial marker sample content (Landmarks sample format). The contents of the Landmarks sample entry and Landmarks sample format are described in detail below.
[0276] The Landmarks sample entry section defines the head information of the facial marker sample, including the number and type of facial markers. The Landmarks sample entry can be considered the metadata of the facial marker sample. The syntax of the Landmarks sample entry is as follows:
[0277]
[0278] Among them, unsigned int(3) level_type specifies the number and type of facial landmarks used, as shown in Table 18 above. Different integer values correspond to different numbers and types of facial landmarks; bit(5) reserved is a reserved bit that can be used for future expansion. It should be set to 0 when there is no expansion.
[0279] The Landmarks sample format section mainly defines the index information and location coordinates of digital human face markers.
[0280] In some embodiments, the syntax of the Landmarks sample format can be as follows:
[0281]
[0282] Among them, unsigned int(32)sample_index specifies the index value of the facial marker sample, which can be used to distinguish different samples; unsigned int(32)landmarks_count specifies the number of facial markers contained in the facial marker sample; unsigned int(32)position_index[i] specifies the index information of the i-th facial marker in the facial marker sample, which can be used to uniquely identify or reference a specific facial marker; float(32)[3]position_value[i] specifies the floating-point three-dimensional position coordinates (x,y,z) of the i-th facial marker in the facial marker sample.
[0283] In other embodiments, the syntax of the Landmarks sample format may be as follows:
[0284]
[0285] Among them, unsigned int(32)sample_index specifies the identifier of the facial marker sample, which can be used to distinguish different samples; unsigned int(32)landmarks_count specifies the number of facial markers contained in the facial marker sample; float(32)delta_x[i] specifies the floating-point position change of the i-th facial marker in the facial marker sample in the x-axis direction relative to the previous frame; float(32)delta_y[i] specifies the floating-point position change of the i-th facial marker in the facial marker sample in the y-axis direction relative to the previous frame; and float(32)delta_z[i] specifies the floating-point position change of the i-th facial marker in the facial marker sample in the z-axis direction relative to the previous frame.
[0286] In some embodiments, the precision values in LandmarksSampleHeader are subject to the following conditions:
[0287] If the value of the Immutable syntax element in the MPEG time-varying accessor (MPEG_accessor_timed) of the accessor description module corresponding to the accessor is True, and there is no accessor slice syntax element (bufferView) in the MPEG time-varying accessor, then the value of the count syntax element (count) in the accessor description module is the number of face markers specified by unsigned int(3)level_type in the Landmarks sampleentry; otherwise, it is obtained from the buffer slice (timed accessor information header) indicated by the accessor slice syntax element (bufferView) in the MPEG time-varying accessor.
[0288] In some embodiments, the LandmarksSample is processed as follows:
[0289] 1. The sample_index is represented in the corresponding frame of the circular buffer.
[0290] 2. The number of data that the accessor needs to access is represented in the time accessor header field of the corresponding frame in the circular buffer.
[0291] 3. For each facial marker, position_index[i] should be represented in the corresponding frame of the circular buffer.
[0292] 4. For each facial marker, position_value[i] should be represented in the corresponding frame of the circular buffer, containing the specific position data of the facial marker.
[0293] Some embodiments of this application provide a method for generating a scene description file, see below. Figure 12 As shown, the method for generating this scene description file may include the following steps:
[0294] S111. Obtain the description information of the facial marker points of the target digital human in the 3D scene to be rendered.
[0295] In some embodiments of this application, the description information of the facial landmarks of the target digital human may include at least one of the following: the level and number of the facial landmarks of the target digital human, the index value of the accessor description module corresponding to the accessor for accessing the index information of the facial landmarks of the target digital human, and the index value of the accessor description module corresponding to the accessor for accessing the position coordinates of the facial landmarks of the target digital human.
[0296] In some embodiments, the level and number of facial markers for the target digital human can be any one of the following 1-4:
[0297] 1. The facial markers for the target digital human are the 68 facial markers provided by MOGAN;
[0298] 2. The facial markers for the target digital human are 21 facial markers provided by the AFLW dataset;
[0299] 3. The facial markers for the target digital human are the 29 facial markers provided by the LFPW dataset;
[0300] 4. The facial markers for the target digital human are 98 facial markers provided by the WFLW dataset.
[0301] In some embodiments, the position coordinates of the facial markers of the target digital human can be the absolute position coordinates of the facial markers of the target digital human in three-dimensional space.
[0302] In some embodiments, the position coordinates of the facial markers of the target digital human can be the relative position coordinates of the facial markers in three-dimensional space. For example, the position coordinates of the facial markers of the target digital human can be the offset between the position coordinates of the facial markers of the target digital human in the current frame and the position coordinates of the facial markers of the target digital human in the previous frame.
[0303] S112. Generate a facial marker array ("landmarks":{}) based on the description information of the facial markers.
[0304] In some embodiments, the description information of the facial markers of the target digital human may include: the number and type of facial markers of the target digital human. Step S112 (generating a facial marker array based on the description information of the facial markers) includes:
[0305] Add a level type syntax element ("level_type") to the facial marker array ("landmarks":{}), and set the value of the level type syntax element according to the number and type of facial markers of the target digital human.
[0306] In some embodiments, setting the value of the level_type syntax element according to the level type of the facial markers of the target digital human may include: when the facial markers of the target digital human are 68 facial markers provided by MOGAN, the level_type syntax element and its value are set to "level_type":0; when the facial markers of the target digital human are 21 facial markers provided by AFLW, the level_type syntax element and its value are set to "level_type":1; when the facial markers of the target digital human are 29 facial markers provided by LFPW, the level_type syntax element and its value are set to "level_type":2; when the facial markers of the target digital human are 98 facial markers provided by WFLW, the level_type syntax element and its value are set to "level_type":3.
[0307] It should also be noted that when the face marker array ("landmarks":{}) does not include the level type syntax element ("level_type") and its value, the default face markers are the 68 face markers provided by MOGAN. That is, unless otherwise specified, the default level type syntax element and its value in the scene description file are "level_type":0.
[0308] In some embodiments, the description information of the facial marker points of the target digital human may include: the index value of the accessor description module corresponding to the accessor used to access the index information of the facial marker points of the target digital human. Step S112 (generating a facial marker point array based on the description information of the facial marker points) includes:
[0309] Add an index syntax element ("indices") to the face marker array ("landmarks":{}) and set the value of the index syntax element to the index value of the accessor description module corresponding to the first accessor in the accessor list of the scene description file.
[0310] The first accessor is an accessor used to access the index information of the facial marker points of the target digital human.
[0311] For example: if the accessor used to access the index information of the facial markers of the target digital human is the accessor corresponding to the third accessor description module in the accessor list, then the index syntax element and its value are set to: "indices": 2.
[0312] In some embodiments, the description information of the facial marker points of the target digital human includes: the index value of the accessor description module corresponding to the accessor used to access the position coordinates of the facial marker points of the target digital human. Step S112 (generating a facial marker point array based on the description information of the facial marker points) includes:
[0313] Add a position coordinate syntax element ("position") to the face marker point array, and set the value of the position coordinate syntax element to the index value of the accessor description module corresponding to the second accessor in the accessor list of the scene description file.
[0314] The second accessor is used to access the position coordinates of the facial marker points of the target digital human.
[0315] For example: if the accessor used to access the position coordinates of the facial marker points of the target digital human is the accessor corresponding to the fourth accessor description module in the accessor list, then the index syntax element and its value are set to: "position":3.
[0316] For example, the target digital human's facial markers are 68 facial markers provided by MOGAN. The accessor used to access the index information of the target digital human's facial markers is the accessor corresponding to the first accessor description module in the accessor list, and the accessor used to access the position coordinates of the target digital human's facial markers is the accessor corresponding to the second accessor description module in the accessor list. Then, the facial marker array ("landmarks":{}) generated according to the description information of the facial markers can be as follows:
[0317]
[0318]
[0319] In some embodiments, the face marker array ("landmarks":{}) may omit the level type syntax element and its value, thereby generating the face marker array ("landmarks":{}) based on the description information of the face markers as follows:
[0320]
[0321] S113. Add the array of facial marker points to the digital human node array ("MPEG_node_avatar":{}) of the first node description module of the scene description file.
[0322] The first node description module is the node description module corresponding to the node in the node list ("nodes":[]) of the scene description file that represents the target digital human.
[0323] According to the second revision of the scene description standard, a digital human in a three-dimensional scene will be represented by a node and described by a corresponding node description module. Therefore, the three-dimensional scene to be rendered includes the node representing the target digital human, and the scene description file of the three-dimensional scene to be rendered includes the node description module corresponding to the node representing the target digital human.
[0324] The scene description file generation method provided in some embodiments of this application first obtains the description information of facial marker points of the target digital human in the 3D scene to be rendered, then generates a facial marker point array based on the description information of the facial marker points, and adds the facial marker point array to the digital human node array of the first node description module corresponding to the node representing the target digital human in the node list of the scene description file. Since the scene description file generation method provided in some embodiments of this application can generate a facial marker point array based on the description information of the facial marker points and add it to the scene description file, parsing the facial marker point array in the scene description file can obtain the description information of the target digital human's facial marker points, and then realize the animation and related processing of the target digital human based on the description information of the target digital human's facial marker points. Therefore, some embodiments of this application can solve the problem that the scene description file does not declare facial marker point information, thus preventing the immersive media scene description framework from using facial marker points to realize digital human animation and related processing.
[0325] The method for generating scene description files provided in some embodiments of this application further includes:
[0326] Add a digital human type syntax element ("type") to the digital human node array ("MPEG_node_avatar":{}) of the first node description module, and set the value of the digital human type syntax element to the Uniform Resource Name of the representation scheme of the target digital human.
[0327] For example, the target digital human is represented by the MPEG Reference Digital Human. Since the unified resource name of the MPEG Reference Digital Human is "mpeg:sd:2023:avatar", a digital human type syntax element ("type") is added to the digital human node array ("MPEG_node_avatar":{}) of the first node description module, and the value of the digital human type syntax element is set to "urn:mpeg:sd:2023:avatar".
[0328] The method for generating scene description files provided in some embodiments of this application further includes:
[0329] Based on the hierarchical standard names of the digital human body parts represented by each digital human component of the target digital human and the node description modules in the node list corresponding to each digital human component of the target digital human, a mapping list ("mappings":[]) corresponding to the target digital human is generated; and the mapping list ("mappings":[]) is added to the digital human node array ("MPEG_node_avatar":{}) of the first node description module.
[0330] In some embodiments, a mapping list corresponding to the target digital human is generated based on the hierarchical standard names of the digital human body parts represented by each digital human component of the target digital human and the node description modules in the node list corresponding to each digital human component of the target digital human, including the following steps ① to ③:
[0331] Step ①: Create a mapping array corresponding to each digital human component of the target digital human in the mapping list ("mappings":[]) corresponding to the target digital human.
[0332] Step 2: Add a part name syntax element ("path") to the mapping array corresponding to each digital human component of the target digital human, and set the value of the part name syntax element ("path") to the hierarchical standard name of the digital human body part represented by the corresponding digital human component.
[0333] For example, the left arm of a digital human represented by a certain digital human component will have its part name syntax element and value set to "path":"full_body / upper_body / arm_left" in the mapping array corresponding to that digital human component.
[0334] Step 3: Add node index syntax elements ("node") to the mapping array corresponding to each digital human component of the target digital human, and set the value of the node index syntax element to the index value of the corresponding node description module.
[0335] For example, if the index value of the node description module corresponding to a node of a certain digital human component is 1, then the node index syntax element (node) added to the mapping array corresponding to the digital human component and its value will be set to "node":1.
[0336] For example, if the index value of the node description module corresponding to a certain component of the target digital human is 2, and the digital human part corresponding to this component is the face, then the mapping array corresponding to this digital human component can be as follows:
[0337]
[0338] The method for generating scene description files provided in some embodiments of this application further includes:
[0339] Add an active identifier syntax element ("isAvatar") to the digital human node array ("MPEG_node_avatar":{}), and set the value of the active identifier syntax element according to whether the target digital human is an active digital human.
[0340] In some embodiments of this application, whether a digital human is an active digital human refers to whether the digital human needs to be rendered during the 3D scene rendering process, and whether the display engine, media access function, etc. in the scene description framework need to process the relevant data and information of the digital human.
[0341] In some embodiments, setting the value of the active identifier syntax element ("isAvatar") based on whether the target digital person is an active digital person may include: if the target digital person is an active digital person, then setting the active identifier syntax element and its value to "isAvatar":true or "isAvatar":1; if the target digital person is an inactive digital person, then setting the active identifier syntax element and its value to "isAvatar":false or "isAvatar":0.
[0342] For example, the target digital human is an active digital human, which includes three digital human components: digital human component A, digital human component B, and digital human component C. The representation scheme of the digital human is: MPEG reference digital human. Then, the digital human node array ("MPEG_node_avatar":{}) according to the first node description module can be as follows:
[0343]
[0344] The method for generating scene description files according to some embodiments of this application further includes the following steps a and b:
[0345] Step a: Generate the target media description module corresponding to the target media file based on the description information of the target media file.
[0346] The target media file can be any media file in the 3D scene to be rendered.
[0347] For example, the target media file may include the index information and location coordinates of the target digital human face markers.
[0348] Step b: Add the target media description module to the media list ("media":{}) of the Moving Picture Experts Group MPEG media ("MPEG_media":{}) in the scene description file.
[0349] In some embodiments, step a (generating a target media description module corresponding to the target media file based on the description information of the target media file) may include at least one of the following steps a1 to a5:
[0350] Step a1: Add a media name syntax element (name) to the target media description module, and set the value of the media name syntax element according to the name of the target media file.
[0351] For example, if the name of the target media file is "landmarks", then add a media name syntax element (name) to the target media description module, and set the media name syntax element and its value to "name":"landmarks".
[0352] Step a2: Add an autoplay syntax element to the target media description module, and set the value of the autoplay syntax element according to whether the target media file needs to be autoplayed.
[0353] For example, if the target media file needs to play automatically, add the syntax element "autoplay" to the target media description module, and set the autoplay syntax element and its value to "autoplay":true or "autoplay":1".
[0354] For example, if the target media file does not need to play automatically, then add the syntax element "autoplay" to the target media description module, and set the autoplay syntax element and its value to "autoplay":false or "autoplay":0.
[0355] Step a3: In the target media description module, loop the playback syntax element (loop), and set the value of the loop playback syntax element according to whether the target media file needs to be looped.
[0356] For example, if the target media file needs to be played in a loop, then add the syntax element "loop" to the target media description module, and set the loop playback syntax element and its value to "loop":true or "autoplay":1".
[0357] For example, if the target media file does not need to be played in a loop, then add the syntax element "loop" to the target media description module, and set the loop playback syntax element and its value to "loop": false or "autoplay": 0.
[0358] Step a4: Add an alternatives list ("alternatives":[]) to the target media description module.
[0359] Step a5: Generate alternative description modules corresponding to each alternative version of the target media file based on the description information of each alternative version of the target media file, and add the alternative description modules corresponding to each alternative version of the target media file to the alternative list ("alternatives":[]).
[0360] In some embodiments, the step a5 above, which generates the optional description module corresponding to each optional version of the target media file based on the description information of each optional version of the target media file, may include at least one of the following steps a51 to a55:
[0361] Step a51: Add a media type syntax element ("mimeType") to the first optional description module corresponding to the first optional version, and set the value of the media type syntax element according to the encapsulation format of the first optional version.
[0362] Wherein, the first optional version is any optional version of the target media file. That is, the optional description module can be generated for each optional version of the target media file through the embodiments of this application.
[0363] Step a52: Add a Uniform Resource Identifier (URI) syntax element ("uri") to the first optional description module, and set the value of the Uniform Resource Identifier syntax element according to the first optional version of the Uniform Resource Identifier (URI).
[0364] For example, if the Uniform Resource Identifier (URI) of the first optional version of the target media file is "www.example.com / avatarface / index", then a Uniform Resource Identifier syntax element ("uri") is added to the first optional description module, and the Uniform Resource Identifier syntax element and its value are set to "uri":"www.example.com / avatarface / index".
[0365] Step a53: Add a track array (tracks[]) to the first optional description module.
[0366] Step a54: Add a track index syntax element ("track") to the track array (tracks[]) and set the value of the track index syntax element according to the track information of the first optional version.
[0367] Step a55: Add codecs to the track array (tracks[]) and set the value of the codecs according to the codec type of the bitstream of the first optional version.
[0368] For example, if the target media file is named "landmarks", and the target media needs to loop and autoplay, including two optional versions, one with a Uniform Resource Identifier of "www.example.com / landmarks / index=0" and the other with a URI of "www.example.com / landmarks / index=1", then the target media description module corresponding to the target media generated according to the above-described embodiment can be as follows:
[0369]
[0370]
[0371] It should be noted that "=TypeValue1", "trackIndex=1", "CodecsValue1", "=TypeValue2", "trackIndex=2", and "CodecsValue2" in the above examples are representative rather than specific values. The values of "mimeType", "track", and "codecs" need to be set according to the optional versions of the corresponding media file and relevant standards. Since this embodiment does not limit the types of the various optional versions of the target media file, the specific values of "mimeType", "track", and "codecs" are not limited in this embodiment.
[0372] The method for generating scene description files provided in some embodiments of this application further includes the following steps c and d:
[0373] Step c: Generate a target scene description module corresponding to the 3D scene to be rendered based on the description information of the 3D scene to be rendered.
[0374] Step d: Add the target scene description module to the scene list ("scenes":[]) of the scene description file.
[0375] The step of generating a target scene description module corresponding to the three-dimensional scene to be rendered based on the description information of the three-dimensional scene to be rendered may include: adding a node index list ("nodes":[]) to the target scene description module, and adding the index value of the node description module corresponding to each top-level node in the three-dimensional scene to be rendered to the node index list.
[0376] It should be noted that, in this embodiment of the application, adding the index value of the node description module corresponding to each top-level node in the 3D scene to be rendered to the node index list means adding the index value of the node description module corresponding to each root node in the 3D scene to be rendered to the node index list, rather than the index values of the node description modules corresponding to all nodes (excluding the index values of the node description modules corresponding to child nodes).
[0377] For example: the 3D scene to be rendered includes two nodes, one of which is a child node of the other, and the index value of the node description module corresponding to the parent node is 0. Then, the target scene description module corresponding to the 3D scene to be rendered added to the scene description file can be as follows:
[0378]
[0379] In the example above, the 3D scene to be rendered includes two nodes, and the index value of the node description module corresponding to the top-level node of the two nodes is 0. Therefore, the index value 0 is added to the node list (nodes) of the target scene description module corresponding to the 3D scene to be rendered.
[0380] The method for generating scene description files provided in some embodiments of this application further includes the following steps e and f:
[0381] Step e: Generate the target node description module corresponding to the target node based on the description information of the target node.
[0382] The target node is any node in the 3D scene to be rendered.
[0383] Step f: Add the target node description module to the node list ("nodes":[]) of the scene description file.
[0384] In some embodiments, step c (generating the target node description module corresponding to the target node based on the description information of the target node) may include at least one of the following steps c1 to c4:
[0385] Step e1: Add a node name syntax element ("name") to the target node description module, and set the value of the node name syntax element according to the name of the target node.
[0386] For example, if the name of the target node is "avatarThorax", then add a node name syntax element ("name") to the target node description module, and set the node name syntax element and its value in the target node description module to "name":"avatarThorax".
[0387] Step e2: Add a child node index list ("children":[]) to the target node description module, and add the index value of the node description module corresponding to each child node attached to the target node to the child node index list.
[0388] For example: If the target node has two child nodes, one child node has an index value of 1 in its node description module, and the other child node has an index value of 2 in its node description module, then the child node index list of the target node description module can be as follows:
[0389]
[0390] Step e3: Add a mesh index syntax element ("mesh") to the target node description module, and set the value of the mesh index syntax element ("mesh") to the index value of the mesh description module corresponding to the 3D mesh mounted on the target node.
[0391] For example: if the target node is mounted with a three-dimensional mesh and the index value of the mesh description module corresponding to the three-dimensional mesh is 0, then add a mesh index syntax element ("mesh") to the target media description module and set the mesh index syntax element and its value in the target node description module to "mesh":0.
[0392] Step e4: Add a position offset syntax element ("translation") to the target node description module, and set the value of the position offset syntax element according to the spatial position offset of the target node relative to its parent node.
[0393] For example, a node in the 3D scene to be rendered is named "avatarThorax". This node includes two child nodes, and the index values of the node description modules corresponding to the two child nodes are 1 and 2, respectively. The index value of the mesh description module corresponding to the 3D mesh attached to this node is 0, and the spatial offset of this node relative to its parent node is 0. Then, the node description module generated based on the description information of this node can be as follows:
[0394]
[0395] For example, a node in the 3D scene to be rendered is named "avatarNeck". This node includes a child node, and the index value of the node description module corresponding to the child node is 2. The index value of the mesh description module corresponding to the 3D mesh attached to the child node is 1. The offset of the child node relative to its parent node is (0.0, 10.0, 20.0). Then, the node description module generated based on the description information of the child node can be as follows:
[0396]
[0397] The method for generating scene description files provided in some embodiments of this application further includes the following steps g and h:
[0398] Step g: Generate a target mesh description module corresponding to the target 3D mesh based on the description information of the target 3D mesh.
[0399] The target 3D mesh can be any 3D mesh in the 3D scene to be rendered.
[0400] Since the 3D mesh in the scene description file is the next level below the node, and the 3D mesh is attached to the node, the target 3D mesh can also be described as the 3D mesh attached to any node in the 3D scene to be rendered.
[0401] Step h adds the target mesh description module to the mesh list ("meshes":[]) of the scene description file.
[0402] In some embodiments, step g (generating a target mesh description module corresponding to the target three-dimensional mesh based on the description information of the target three-dimensional mesh) may include the following steps g1 to g3:
[0403] Step g1: Add a mesh name syntax element ("name") to the target mesh description module, and set the value of the mesh name syntax element according to the name of the target 3D mesh.
[0404] For example, if the name of the target 3D mesh is "avatarFace", then add a mesh name syntax element (name) to the target mesh description module, and set the mesh name syntax element and its value in the target mesh description module to "name":"avatarFace".
[0405] Step g2: Add a position syntax element ("position") to the attribute ("attributes":{}) of the primitive ("primitives":[]) of the target mesh description module, and set the value of the position syntax element according to the index value of the accessor description module corresponding to the accessor used to access the dynamic 3D model of the digital human component corresponding to the target 3D mesh.
[0406] Step g3: Add a mode syntax element ("mode") to the primitives ("primitives":[]) of the target mesh description module, and set the value of the mode syntax element according to the topology type of the target 3D mesh.
[0407] For example, if a 3D mesh is named "avatarFace", the accessor description module corresponding to the accessor used to access the dynamic 3D model of the digital human component corresponding to this 3D mesh has an index value of 0, and the topology type of this 3D mesh is triangular facets, then the mesh description module corresponding to this 3D mesh can be as follows:
[0408]
[0409] The method for generating scene description files in some embodiments of this application further includes the following steps i and j:
[0410] Step i: Generate the target accessor description module corresponding to the target accessor based on the description information of the target accessor.
[0411] The target accessor is any accessor used to render the 3D scene to be rendered.
[0412] Step j: Add the target accessor description module to the accessor list ("accessors":[]) of the scene description file.
[0413] In some embodiments, step i (generating a target accessor description module corresponding to the target accessor based on the target accessor's description information) may include at least one of the following steps i1 to i8:
[0414] Step i1: Add a data type syntax element ("componentType") to the target accessor description module, and set the value of the data type syntax element according to the data type accessed by the target accessor.
[0415] For example, if the data type accessed by a certain accessor is 5126, then the data type syntax element and its value in the accessor description module corresponding to that accessor are set to: "componentType":5126.
[0416] Step i2: Add an accessor type syntax element ("type") to the target accessor description module, and set the value of the accessor type syntax element according to the type of the target accessor.
[0417] For example, if an accessor accesses a two-dimensional vector, then the accessor type syntax element ("type") and its value in the accessor description module corresponding to that accessor are set to: "type":"VEC2".
[0418] Step i3: Add a data count syntax element ("count") to the target accessor description module, and set the value of the data count syntax element according to the number of data accessed by the target accessor.
[0419] For example, if an accessor accesses 1000 data items, then the data count syntax element ("count") and its value in the accessor description module corresponding to that accessor are set to: "count":1000.
[0420] Step i4: Add a first cache slice index syntax element ("bufferView") to the target accessor description module, and set the value of the first cache slice index syntax element according to the index value of the cache slice description module corresponding to the cache slice used to cache the data accessed by the target accessor.
[0421] For example: If the cache slice description module corresponding to the cache slice that caches the data accessed by a certain accessor is the fourth cache slice description module (index value 3) in the cache slice list (bufferViews) of the scene description file, then the target cache slice index syntax element and its value in the accessor description module corresponding to the accessor are set to: "bufferView":3.
[0422] Step i5: Add an MPEG time-varying accessor ("MPEG_accessor_timed":{}) to the extension list ("extensions":{}) of the target accessor description module.
[0423] Step i6: Add a second buffer slice index syntax element ("bufferView") to the MPEG time-varying accessor ("MPEG_accessor_timed":{}), and set the value of the second buffer slice index syntax element according to the index value of the buffer slice description module corresponding to the buffer slice used to cache the time-varying parameters of the target accessor.
[0424] For example, if the cache slice description module corresponding to the cache slice of a time-varying parameter of a certain accessor is the second cache slice description module in the cache slice list (bufferViews) of the scene description file, then the second cache slice index syntax element and its value in the accessor description module corresponding to that accessor are set to: "bufferView":1.
[0425] Step i7: Add a time-varying syntax element ("immutable") to the MPEG time-varying accessor ("MPEG_accessor_timed":{}), and set the value of the time-varying syntax element according to whether the value of the syntax element in the target accessor changes over time.
[0426] In some embodiments, setting the value of the time-varying syntax element based on whether the value of the syntax element in the target accessor changes over time may include: when the value of the syntax element in the target accessor does not change over time, setting the time-varying syntax element and its value in the MPEG time-varying accessor of the target accessor description module to "immutable":true or "immutable":1; when the value of the syntax element in the target accessor changes over time, setting the time-varying syntax element and its value in the MPEG time-varying accessor of the target accessor description module to "immutable":false or "immutable":0.
[0427] Step i8: Add an accessor name syntax element ("name") to the target accessor description module, and set the value of the accessor name syntax element according to the name of the target accessor.
[0428] For example, if an accessor accesses data of type 5126, the accessor's name is "avatarThorax", the accessor's type is VEC3, the number of data accessed by the accessor is 1828, the index value of the cache slice description module corresponding to the cache slice caching the data accessed by the accessor is 0, and the values of the syntax elements within the accessor change over time, the index value of the cache slice description module corresponding to the cache slice caching the time-varying parameters of the accessor is 2, then the accessor description module corresponding to this accessor can be as follows:
[0429]
[0430] The method for generating scene description files in some embodiments of this application further includes the following steps k and l:
[0431] Step k: Generate the target cache description module corresponding to the target cache based on the description information of the target cache.
[0432] The target buffer is any buffer used to render the 3D scene to be rendered.
[0433] Step 1: Add the target buffer description module to the buffer list ("buffers":[]) of the scene description file.
[0434] In some embodiments, step k (generating a target cache description module corresponding to the target cache based on the target cache description information) may include at least one of the following steps k1 to k7:
[0435] Step k1: Add a first byte length syntax element (byteLength) to the target buffer description module, and set the value of the first byte length syntax element according to the capacity of the target buffer.
[0436] The target buffer is any buffer used to render the 3D scene to be rendered.
[0437] For example, when the capacity of a certain cache is 15000 bytes, the first byte length syntax element in the cache description module corresponding to the cache is set to: "byteLenth":15000.
[0438] Step k2: Add an MPEG circular buffer ("MPEG_buffer_circular":{}) to the target buffer description module.
[0439] Step k3: Add a stage count syntax element ("count") to the MPEG circular buffer ("MPEG_buffer_circular":{}), and set the value of the stage count syntax element according to the number of stages stored in the target buffer.
[0440] For example, if a buffer has 5 storage stages, then the stage count syntax element and its value in the MPEG circular buffer of the buffer description module corresponding to that buffer are set to: "count":5.
[0441] Step k4: Add a media index syntax element ("media") to the MPEG circular buffer, and set the value of the media index syntax element according to the index value of the media description module corresponding to the media file to which the source data of the data cached by the target buffer belongs.
[0442] For example: if the index value of the media description module corresponding to the media file to which the source data of the data cached by a certain buffer belongs is 0, then the media index syntax element and its value in the MPEG circular buffer of the buffer description module corresponding to that buffer are set to "media":0.
[0443] Step k5: Add track index syntax elements (tracks) to the MPEG circular buffer, and set the value of the track index syntax elements according to the track index value of the source data of the data cached by the target buffer.
[0444] Step k6: Add a cache name syntax element ("name") to the cache description module corresponding to the target cache, and set the cache name syntax element according to the name of the target cache.
[0445] Step k7: Add a Uniform Resource Identifier (URI) syntax element to the target cache description module, and set the value of the URI syntax element according to at least a portion of the data used to reconstruct the dynamic 3D model of the corresponding digital human component.
[0446] That is, the data of the dynamic 3D model used to reconstruct the corresponding digital human component is directly added to the scene description file.
[0447] For example, a buffer named "avatarFace" has a capacity of 108940 bytes, has 3 storage stages, and the index value of the media description module corresponding to the media file to which the source data of the buffered data belongs is 1. Then, the buffer description module corresponding to this buffer added to the buffer list of the scene description file can be as follows:
[0448]
[0449] In some embodiments, the method for generating the scene description file may further include the following steps m and n:
[0450] Step m: Generate a target cache slice description module corresponding to the target cache slice based on the description information of the target cache slice.
[0451] The target cache slice is any cache slice of the cache used to implement the rendering of the 3D scene to be rendered.
[0452] Step n: Add the target cache slice description module to the cache slice list ("bufferViews":[]) of the scene description file.
[0453] In some embodiments, step m (generating a target cache slice description module corresponding to the target cache slice based on the description information of the target cache slice) may include at least one of the following steps m1 to m4:
[0454] Step m1: Add a buffer index syntax element ("buffer") to the target buffer slice description module, and set the value of the buffer index syntax element according to the index value of the buffer description module corresponding to the buffer to which the target buffer slice belongs.
[0455] For example, if the index value of the cache description module corresponding to a certain cache is 2, then the cache index syntax element and its value in the cache slice description module corresponding to the cache slice of that cache are set to: "buffer":2.
[0456] Step m2: Add a second byte length syntax element ("byteLength") to the cache slice description module corresponding to the target cache slice, and set the value of the second byte length syntax element ("byteLength") according to the capacity of the target cache slice.
[0457] Step m3: Add an offset syntax element ("byteOffset") to the target cache slice description module, and set the value of the offset syntax element according to the offset of the data cached by the target cache slice.
[0458] Step m4: Add a cache slice name syntax element ("name") to the target cache slice description module, and set the value of the cache slice name syntax element according to the name of the target cache slice.
[0459] For example, if the index value of the cache description module corresponding to a certain cache is 0, the capacity of the cache is 43972, and the cache includes three cache slices, the first cache slice is named "avatarThorax1" with a capacity of 21936 and an offset of 0, the second cache slice is named "avatarThorax2" with a capacity of 21936 and an offset of 21936, and the third cache slice is named "avatarThoraxHead" with a capacity of 100 and an offset of 43872, then the cache slice description module part corresponding to the cache slice of this cache in the scene description file can be as follows:
[0460]
[0461]
[0462] In some embodiments, the method for generating the scene description file may further include:
[0463] Add a digital asset description module (asset) to the scene description file, add a version syntax element (version) to the digital asset description module, and set the value of the version syntax element according to the version information of the scene description file.
[0464] For example, when the scene description file is written based on glTF version 2.0, the value of the version syntax element is set to 2.0.
[0465] For example, the digital asset description module added to the scene description file can be as follows:
[0466]
[0467] In some embodiments, the method for generating the scene description file may further include:
[0468] Add an extensionsUsed list to the scene description file, and add the top-level extension item of the MPEG to glTF2.0 version scene description file used by the scene description file to the extensionsUsed list.
[0469] For example, if the MPEG extensions used in the scene description file include: MPEG media, MPEG buffer circular, MPEG accessor timed, and MPEG node avatar, then the list of extensions used in the scene description file can be as follows:
[0470]
[0471] In some embodiments, the method for generating the scene description file may further include:
[0472] Add a scene declaration to the scene description file and set the value of the scene declaration to the index value of the scene description module corresponding to the scene to be rendered.
[0473] For example, if the index value of the scene description module corresponding to the scene to be rendered is 0, then adding a scene declaration to the scene description file can be done as follows:
[0474]
[0475] Some embodiments of this application also provide a method for parsing a scene description file, see below. Figure 13As shown, the method for parsing this scene description file may include the following steps S121 to S123:
[0476] S121. Obtain the first node description module from the node list ("nodes":[]) of the scene description file of the 3D scene to be rendered.
[0477] The first node description module is the node description module corresponding to the node representing the target digital human.
[0478] In some embodiments, the first node description module obtained from the scene description file may be as follows:
[0479]
[0480]
[0481] In some embodiments, the module for obtaining the target node description from the scene description file may also be as follows:
[0482]
[0483] S122. Obtain the face marker array ("landmarks":{}) from the digital human node array ("MPEG_node_avatar":{}) of the first node description module.
[0484] Following the examples above, in some embodiments, the facial marker array ("landmarks":{}) obtained from the digital human node array ("MPEG_node_avatar":{}) of the first node description module can be as follows:
[0485]
[0486] Following the examples above, in some embodiments, the facial marker array ("landmarks":{}) obtained from the digital human node array ("MPEG_node_avatar":{}) of the first node description module can also be as follows:
[0487]
[0488] S123. Obtain the description information of the facial landmarks (landmarks) of the target digital human according to the facial landmark array ("landmarks":{}).
[0489] In this embodiment of the application, the description information of the facial landmarks of the target digital human may include at least one of the following: the level and number of the facial landmarks of the target digital human, the index value of the accessor description module corresponding to the accessor for accessing the index information of the facial landmarks of the target digital human, and the index value of the accessor description module corresponding to the accessor for accessing the position coordinates of the facial landmarks of the target digital human.
[0490] In some embodiments, the facial marker array ("landmarks":{}) may include: a level type syntax element ("level_type") and the value of the level type syntax element. Step S123 above (obtaining description information of the facial markers of the target digital human based on the facial marker array) may include:
[0491] Based on the value of the level type syntax element ("level_type"), obtain the number and type of facial markers of the target digital human.
[0492] In some embodiments, obtaining the level type of the facial markers of the target digital human based on the value of the level type syntax element may include: obtaining the number and type of the facial markers of the target digital human based on the value of the level type syntax element and the mapping relationship shown in Table 18.
[0493] Based on the mapping relationship shown in Table 18 above, when the level type syntax element and its value are "level_type":0, the facial markers of the target digital human can be determined to be 68 facial markers provided by MOGAN; when the level type syntax element and its value are "level_type":1, the facial markers of the target digital human can be determined to be 21 facial markers provided by the AFLW dataset; when the level type syntax element and its value are "level_type":2, the facial markers of the target digital human can be determined to be 29 facial markers provided by the LFPW dataset; and when the level type syntax element and its value are "level_type":3, the facial markers of the target digital human can be determined to be 98 facial markers provided by the WFLW dataset.
[0494] In some embodiments, the facial marker array ("landmarks":{}) may include: index syntax elements ("indices") and the values of the index syntax elements. Step S123 above (obtaining descriptive information of the facial markers of the target digital human based on the facial marker array) may include:
[0495] Based on the value of the index syntax element ("indices"), obtain the index value of the accessor description module corresponding to the first accessor in the accessor list ("accessors":[]) of the scene description file.
[0496] The first accessor is an accessor used to access the index information of the facial marker points of the target digital human.
[0497] For example, the index syntax element ("indices") and its value "indices":2 can determine that the accessor description module corresponding to the first accessor is the third accessor description module in the accessor list ("accessors":[]).
[0498] In some embodiments, the facial marker array ("landmarks":{}) may include: a position coordinate syntax element ("position") and the value of the position coordinate syntax element. Step S123 above (obtaining description information of the facial markers of the target digital human based on the facial marker array) may include:
[0499] Based on the value of the position coordinate syntax element ("position"), obtain the index value of the accessor description module corresponding to the second accessor in the accessor list ("accessors":[]) of the scene description file.
[0500] The second accessor is used to access the position coordinates of the facial marker points of the target digital human.
[0501] For example, when the position coordinate syntax element ("position") and its value is "position":1, it can be determined that the accessor description module corresponding to the second accessor is the second accessor description module in the accessor list ("accessors":[]).
[0502] In some embodiments, the position coordinates of the facial markers of the target digital human obtained by accessing the second accessor are the absolute position coordinates of the facial markers of the target digital human in three-dimensional space.
[0503] In some embodiments, the position coordinates of the facial marker points of the target digital human obtained by accessing the second accessor are the relative position coordinates of the facial marker points of the target digital human in three-dimensional space. For example, the position coordinates of the facial marker points of the target digital human can be the offset between the position coordinates of the facial marker points of the target digital human in the current frame and the position coordinates of the facial marker points of the target digital human in the previous frame.
[0504] For example, the facial marker array ("landmarks":{}) obtained from the digital human node array ("MPEG_node_avatar":{}) of the first node description module is as follows:
[0505]
[0506] Alternatively, the facial marker array ("landmarks":{}) obtained from the digital human node array ("MPEG_node_avatar":{}) of the first node description module is as follows:
[0507]
[0508] Therefore, the facial marker description information that can be obtained by parsing the facial marker array ("landmarks":{}) may include: the facial markers of the target digital human are 68 facial markers provided by MOGAN; the index information of the target digital human can be accessed through the accessor corresponding to the first accessor description module in the accessor list ("accessors":[]); and the position coordinates of the target digital human can be accessed through the accessor corresponding to the second accessor description module in the accessor list ("accessors":[]).
[0509] The scene description file parsing method of some embodiments of this application first obtains the first node description module corresponding to the node representing the target digital human from the node list of the scene description file of the 3D scene to be rendered. Then, it obtains the facial marker point array from the digital human node array of the first node description module, and obtains the description information of the facial marker points of the target digital human based on the facial marker point array. Since the embodiments of this application can obtain the node description module corresponding to the target digital human from the scene description file, obtain the facial marker point array corresponding to the target digital human from the node description module corresponding to the target digital human, and obtain the description information of the facial marker points of the target digital human based on the facial marker point array, the embodiments of this application can parse the scene description file to obtain the description information of the facial marker points of the target digital human, and realize the animation and related processing of the target digital human based on the description information of the facial marker points of the target digital human. Therefore, the embodiments of this application can solve the problem that the scene description file does not declare the relevant information of the facial marker points, which leads to the immersive media scene description framework being unable to realize digital human animation and related processing based on the facial marker points.
[0510] The method for parsing scene description files in some embodiments of this application may further include:
[0511] Obtain the digital human type syntax element ("type") and its value from the digital human node array ("MPEG_node_avatar":{}) of the first node description module; obtain the representation scheme of the target digital human based on the value of the digital human type syntax element ("type").
[0512] For example, when the digital human type syntax element and its value are "type":"urn:mpeg:sd:2023:avatar", the representation scheme of the target digital human can be determined as MPEG reference digital human based on the value of the digital human type syntax element (type).
[0513] The method for parsing scene description files in some embodiments of this application may further include the following steps (1) and (2):
[0514] Step (1): Obtain the mapping list ("mappings":[]) corresponding to the target number from the digital human node array ("MPEG_node_avatar":{}) of the first node description module.
[0515] For example, obtaining the mapping list ("mappings":[]) corresponding to the target digit from the digit human node array ("MPEG_node_avatar":{}) of the first node description module can be as follows:
[0516]
[0517]
[0518] Step 2: Based on the mapping list ("mappings":[]) corresponding to the target digital human, obtain the digital human body parts represented by each digital human component of the target digital human and the node description modules in the node list corresponding to each digital human component of the target digital human.
[0519] In some embodiments, obtaining the digital human body parts represented by each digital human component of the target digital human and the node description module in the node list corresponding to each digital human component of the target digital human based on the mapping list ("mappings":[]) corresponding to the target digital human may include the following steps 1) to 3):
[0520] Step 1) Obtain the mapping array corresponding to each digital human component of the target digital human from the mapping list ("mappings":[]) corresponding to the target digital human.
[0521] For example, the mapping array corresponding to one digital human component of the target digital human can be as follows:
[0522]
[0523] Step 2) Obtain the digital human body parts represented by each digital human component of the target digital human based on the value of the part name syntax element ("path") in the mapping array corresponding to each digital human component of the target digital human.
[0524] For example, if the part name syntax element in the mapping array corresponding to a digital human component of the target digital human is "path":"full_body / upper_body / head / face", then the digital human body part represented by the digital human component can be determined to be the face based on the value of the part name syntax element in the mapping array corresponding to the digital human component.
[0525] Step 3) Based on the value of the node index syntax element ("node") in the mapping array corresponding to each digital human component of the target digital human, obtain the node description module in the node list corresponding to each digital human component of the target digital human.
[0526] For example, if the node index syntax element in the mapping array corresponding to a digital human component of the target digital human is "node":2, then the node description module corresponding to the digital human component can be determined to be the third node description module in the node list based on the value of the index syntax element in the mapping array corresponding to the digital human component.
[0527] The method for parsing scene description files in some embodiments of this application may further include:
[0528] The value of the active identifier syntax element (isAvatar) in the digital human node array is used to determine whether the target digital human is an active digital human.
[0529] In some embodiments, determining whether the target digital person is an active digital person based on the value of the active identifier syntax element in the digital person node array of the target node description module may include: if the value of the active identifier syntax element in the digital person node array of the target node description module is 1 or true, then the target digital person is determined to be an active digital person; if the value of the active identifier syntax element in the digital person node array of the target node description module is 0 or false, then the target digital person is determined to be an inactive digital person.
[0530] For example, the node description module corresponding to the first node of the target digital human is shown below:
[0531]
[0532]
[0533] Therefore, according to "mesh":0 in line n+01, we know that the first node mounts the 3D mesh as the 3D mesh described by the first mesh description module in the mesh list; according to "isAvatar":True in line n+04, we know that the target digital human is an active digital human; according to "type":"urn:mpeg:sd:2023:avatar" in line n+05, we know that the representation scheme of the digital human represented by the node described by this node description module is MPEG reference digital human; according to "path":"full_body / upper_body / arm_left" in line n+08, we know that the target digital human includes a digital human component representing the left arm of the digital human; according to "node":1 in line n+09, we know that the digital human component representing the left arm is the node corresponding to the second node description module in the node list; according to line n+12... The path "full_body / upper_body / arm_left" indicates that the target digital human includes a digital human component representing the right arm. Based on line n+13's "node":1, the digital human component representing the right arm corresponds to the third node description module in the node list. Line n+17's "level_type":0 indicates that the target digital human's facial markers are the 68 facial markers provided by MOGAN. Line n+18's "indices":0 indicates that the accessor used to access the index information of the target digital human's facial markers is the accessor corresponding to the first accessor description module in the accessor list. Line n+19's "position":0 indicates that the accessor used to access the position coordinates of the target digital human's facial markers is the accessor corresponding to the second accessor description module in the accessor list.
[0534] The method for parsing scene description files in some embodiments of this application may further include the following steps 1 and 2:
[0535] Step 1: Obtain the target media description module from the media list ("media":[]) of the MPEG media ("MPEG_media":{}) in the scene description file.
[0536] The target media description module is any media description module in the media list of the MPEG media of the scene description file.
[0537] Step 2: Obtain the description information of the target media file corresponding to the target media description module according to the target media description module.
[0538] In some embodiments, obtaining the description information of the target media file according to the target media description module may include at least one of the following steps 21 to 24:
[0539] Step 21: Obtain the name of the target media file based on the value of the media name syntax element ("name") in the target media description module.
[0540] For example, if the media name syntax element and its value in the target media description module are "name":"avatarThorax", then the name of the target media file can be determined to be: avatarThorax.
[0541] Step 22: Determine whether the target media file needs to be automatically played based on the value of the autoplay syntax element ("autoplay") in the target media description module.
[0542] In some embodiments, determining whether the target media file needs to be automatically played based on the value of the autoplay syntax element ("autoplay") in the target media description module may include: if the autoplay syntax element ("autoplay") in the target media description module and its value are "autoplay":true or "autoplay":1, then the target media file needs to be automatically played; and if the autoplay syntax element ("autoplay") in the target media description module and its value are "autoplay":false or "autoplay":0, then the target media file does not need to be automatically played.
[0543] Step 23: Determine whether the target media file needs to be played in a loop based on the value of the loop playback syntax element ("loop") in the target media description module.
[0544] In some embodiments, determining whether the target media file needs to be looped based on the value of the loop playback syntax element ("loop") in the target media description module may include: if the loop playback syntax element ("loop") in the target media description module and its value are "loop":true or "loop":1, then the target media file needs to be looped; and if the loop playback syntax element (loop) in the target media description module and its value are "loop":false or "loop":0, then the target media file does not need to be looped.
[0545] Step 24: Obtain the description information of each optional version of the target media file according to each optional description module in the optional list ("alternatives":[]) of the target media description module.
[0546] In some embodiments, step 24 above (obtaining description information of each optional version of the target media file according to each optional description module in the optional list of the target media description module) may include at least one of the following steps 241 to 244:
[0547] Step 241: Obtain the encapsulation format of the first optional version corresponding to the first optional description module based on the value of the media type syntax element ("mimeType") in the first optional description module.
[0548] Wherein, the first optional description module can be any optional description module in the optional list. Since the first optional description module can be any optional description module in the optional list, the description information of each optional version of the target media file can be obtained through the embodiments of this application.
[0549] For example, if the media type syntax element ("mimeType") in the optional description module corresponding to a certain optional version of the target media file has the value "mimeType":"application / mp4", then it can be determined that the container format of that optional version of the target media file is MP4.
[0550] Step 242: Obtain the first optional version of the Uniform Resource Identifier (URI) based on the value of the Uniform Resource Identifier syntax element ("uri") in the first optional description module.
[0551] For example, if the Uniform Resource Identifier syntax element and its value in a certain alternative description module in the alternative list ("alternatives":[]) of the target media description module is "uri":"https: / / www.example.com / avatarthorax", then the Uniform Resource Identifier of the alternative version corresponding to that alternative description module can be determined to be "https: / / www.example.com / avatarthorax".
[0552] Step 243: Obtain the track information of the first optional version based on the value of the first track index syntax element ("track") in the track array ("tracks":[]) of the first optional description module.
[0553] Step 244: Obtain the decoder type of the first optional version of the bitstream based on the value of the codecs element in the track array ("tracks":[]) of the first optional description module.
[0554] The scenario description file parsing method in some embodiments of this application may include the following steps 3 and 4:
[0555] Step 3: Obtain the target scene description module corresponding to the 3D scene to be rendered from the scene list ("scenes":[]) of the scene description file.
[0556] Step 4: Obtain the description information of the 3D scene to be rendered according to the target scene description module.
[0557] In some embodiments, step 4 (obtaining the description information of the three-dimensional scene to be rendered according to the target scene description module) may include: determining the index value of the node description module corresponding to each top-level node in the three-dimensional scene to be rendered according to the index value declared in the node index list ("nodes":[]) of the target scene description module.
[0558] For example, if the node index list of the target scene description module and its declared index value is "nodes":[0], then it can be determined that the three-dimensional scene to be rendered includes only one top-level node, and the node description module corresponding to the top-level node is the first node description module in the node list of the scene description file.
[0559] For example, if the node index list of the target scene description module and its declared index value is "nodes":[0,2], then it can be determined that the three-dimensional scene to be rendered includes two top-level nodes, and the node description modules corresponding to the two top-level nodes are the first node description module and the third node description module in the node list of the scene description file, respectively.
[0560] The scenario description file parsing method in some embodiments of this application may include the following steps 5 and 6:
[0561] Step 5: Obtain the target node description module from the node list ("nodes":[]) of the scene description file.
[0562] The target node description module can be any node description module in the node list.
[0563] Step 6: Obtain the description information of the target node corresponding to the target node description module according to the target node description module.
[0564] In some embodiments, step 4 above (obtaining the target node description module from the node list of the scene description file) may include at least one of the following steps 61 to 64:
[0565] Step 61: Obtain the name of the target node based on the value of the node name syntax element ("name") in the target node description module.
[0566] For example, if the node name syntax element in the target node description module has the value "name":"avatarThorax", then the name of the target node can be obtained as avatarThorax based on the value of the node name syntax element in the target node description module.
[0567] Step 62: Based on the index values declared in the child node index list ("children":[]) of the target node description module, obtain the index values of the node description modules corresponding to each child node attached to the target node.
[0568] For example: if the child node index list in the target node description module and its declared index value is "children":[1,2], then according to the declared index value in the child node index list of the target node description module, it can be determined that the target node has two child nodes, and the two child nodes attached to the target node are the nodes corresponding to the second and third node description modules in the node list of the scene description file.
[0569] Step 63: Based on the index value declared in the mesh index syntax element ("mesh") of the target node description module, obtain the index value of the mesh description module corresponding to each 3D mesh mounted on the target node.
[0570] For example, if the mesh index syntax element and its value in the target node description module are "mesh":0, then the 3D mesh corresponding to the first mesh description module in the mesh list ("meshes":[]) of the scene description file can be obtained according to the index value declared in the mesh index syntax element of the target node description module.
[0571] Step 64: Obtain the spatial position offset of the target node relative to its parent node based on the value of the position offset syntax element ("translation") in the target node description module.
[0572] For example, if the position offset syntax element in a node description module has the value "translation":[0.0,0.0,20.0], then the spatial position offset of the node relative to its parent node can be obtained as [0.0,0.0,20.0] based on the value of the position offset syntax element in the node description module.
[0573] It should be noted that the spatial position offset of a node relative to its parent node is directly obtained from the value of the position offset syntax element in the node description module. This spatial position offset is not necessarily the spatial position offset of the node relative to the node representing the target digital person. The spatial position offset of the node relative to the node representing the target digital person is the sum of the spatial offsets of each node on the path from the node to the node representing the target digital person.
[0574] For example: Node A is a child node of the node representing the target digital human, and node B is a child node of node A. The position offset syntax element in the node description module corresponding to node A is "translation":[0.0,10.0,25.0], and the position offset syntax element in the node description module corresponding to node B is "translation":[0.0,10.0,20.0]. Then, we can first obtain the offset of node A relative to the node representing the target digital human as [0.0,10.0,25.0], and the offset of node B relative to node A as [0.0,00.0,20.0]. Then, we can sum the position offsets to obtain the offset of node B relative to the node representing the target digital human as [0.0,10.0,45.0].
[0575] The method for parsing scene description files in some embodiments of this application may further include:
[0576] Step 7: Obtain the target mesh description module from the mesh list ("meshes":[]) of the scene description file.
[0577] The target mesh description module is any mesh description module in the mesh list ("meshes":[]).
[0578] Step 8: Obtain the description information of the target 3D mesh corresponding to the target mesh description module according to the target mesh description module.
[0579] In some embodiments, step 8 (obtaining the description information of the target 3D mesh corresponding to the target mesh description module according to the target mesh description module) may include at least one of the following steps 81 to 83:
[0580] Step 81: Obtain the name of the target 3D mesh based on the value of the mesh name syntax element ("name") in the target mesh description module.
[0581] For example, if the mesh name syntax element in a mesh description module in the scene description file has the value "name":"avatarNeck", then the name of the 3D mesh corresponding to that mesh description module can be obtained as "avatarNeck" based on the value of the mesh name syntax element in that mesh description module.
[0582] Step 82: Based on the value of the position syntax element ("position") in the attribute ("attributes":{}) of the primitive ("primitives"[]) of the target mesh description module, obtain the index value of the accessor description module corresponding to the accessor used to access the dynamic 3D model of the digital human component corresponding to the target 3D mesh.
[0583] For example, if the position syntax element in the properties of a primitive of a mesh description module in the scene description file has a value of "position":2, then based on the value of the position syntax element in the properties of the primitive of the mesh description module, the accessor used to access the dynamic 3D model of the digital human component corresponding to the 3D mesh of the mesh description module can be determined to be the accessor corresponding to the third accessor description module in the scene description file.
[0584] Step 83: Obtain the type of topology of the target 3D mesh based on the value of the mode syntax element ("mode") in the primitives ("primitives"[]) of the target mesh description module.
[0585] For example, if the pattern syntax element in the primitive of a mesh description module in the scene description file has the value "mode":0, then the topology type of the 3D mesh corresponding to the mesh description module can be obtained as scattered points based on the value of the pattern syntax element in the primitive of the mesh description module.
[0586] In some embodiments, the method for parsing scene description files provided in some embodiments of this application may further include the following steps 9 and 10:
[0587] Step 9: Obtain the target accessor description module from the accessor list ("accessors":[]) of the scene description file.
[0588] The target accessor description module is any accessor description module in the accessor list ("accessors":[]) of the scene description file.
[0589] Step 10: Obtain the description information of the target accessor corresponding to the target accessor description module according to the target accessor description module.
[0590] In some embodiments, step 10 (obtaining the description information of the target accessor corresponding to the target accessor description module according to the target accessor description module) may include at least one of the following steps 101 to 108:
[0591] Step 101: Obtain the data type of the data accessed by the target accessor based on the value of the data type syntax element ("componentType") in the target accessor description module.
[0592] For example, if the data type syntax element and its value in a certain accessor description module are "componentType": 5126, then it can be determined that the data type accessed by the accessor corresponding to the accessor description module is a 32-bit floating-point number (float).
[0593] Step 102: Determine the type of the target accessor based on the value of the accessor type syntax element ("type") in the target accessor description module.
[0594] For example, if the accessor type syntax element in an accessor description module has the value "type":"VEC2", then the type of the accessor corresponding to that accessor description module is a two-dimensional vector.
[0595] Step 103: Determine the number of data accessed by the target accessor based on the value of the data count syntax element ("count") in the target accessor description module.
[0596] For example, if the data count syntax element and its value in a certain accessor description module are "count": 1828, then it can be determined that the number of data accessed by the accessor corresponding to that accessor description module is 1828.
[0597] Step 104: Determine the index value of the cache slice description module corresponding to the cache slice that caches the data accessed by the target accessor based on the value of the first cache slice index syntax element ("bufferViews") in the target accessor description module.
[0598] For example, if the first cache slice index syntax element of a certain accessor description module and its value is "bufferView":1, then it can be determined that the data accessed by the accessor corresponding to the accessor description module is cached in the cache slice corresponding to the second cache slice description module in the cache slice list ("bufferViews":[]).
[0599] Step 105: Determine whether the target accessor is a time-varying accessor based on MPEG extensions based on whether the target accessor description module contains an MPEG time-varying accessor ("MPEG_accessor_timed":{}).
[0600] In some embodiments, determining whether the target accessor is a time-varying accessor based on MPEG extensions based on whether the target accessor description module contains an MPEG time-varying accessor ("MPEG_accessor_timed":{}) may include: if the target accessor description module contains an MPEG time-varying accessor, then the target accessor is determined to be a time-varying accessor based on MPEG extensions; and if the target accessor description module does not contain an MPEG time-varying accessor, then the target accessor is determined not to be a time-varying accessor based on MPEG extensions.
[0601] Step 106: Determine the cache slice description module corresponding to the cache slice that caches the time-varying parameters of the target accessor based on the value of the second cache slice index syntax element ("bufferView") in the MPEG time-varying accessor ("MPEG_accessor_timed":{}) of the target accessor description module.
[0602] For example, if the second buffer slice index syntax element in the MPEG time-varying accessor of a certain accessor description module is "bufferView":3, then it can be determined that the time-varying parameters of the accessor corresponding to the accessor description module are cached in the buffer slice corresponding to the fourth buffer slice description module in the buffer slice list ("bufferViews":[]).
[0603] Step 107: Determine whether the value of the syntax element in the target accessor changes over time based on the value of the time-varying syntax element ("immutable") in the MPEG time-varying accessor ("MPEG_accessor_timed":{}) of the target accessor description module.
[0604] In some embodiments, determining whether the value of a syntax element in the target accessor changes over time based on the value of the time-varying syntax element ("immutable") in the MPEG time-varying accessor of the target accessor description module may include: if the time-varying syntax element in the MPEG time-varying accessor of the target accessor description module and its value are "immutable":true or "immutable":1, then it is determined that the value of the syntax element in the target accessor does not change over time; and if the time-varying syntax element in the MPEG time-varying accessor of the target accessor description module and its value are "immutable":false or "immutable":0, then it is determined that the value of the syntax element in the target accessor changes over time.
[0605] Step 108: Determine the name of the target accessor based on the value of the accessor name syntax element ("name") in the target accessor description module.
[0606] In some embodiments of this application, the method for parsing scene description files may further include the following steps 11 and 12:
[0607] Step 11: Obtain the target buffer description module from the buffer list ("buffers":[]) of the scene description file.
[0608] The target buffer description module is any buffer description module in the buffer list ("buffers":[]).
[0609] Step 12: Obtain the description information of the target cache corresponding to the target cache description module according to the target cache description module.
[0610] In some embodiments, step 12 (obtaining the description information of the target cache corresponding to the target cache description module according to the target cache description module) may include at least one of the following steps 121 to 127:
[0611] Step 121: Obtain the name of the target cache based on the value of the cache name syntax element ("name") in the target cache description module.
[0612] Step 122: Obtain the capacity of the target cache corresponding to the target cache description module based on the value of the first byte length syntax element ("byteLength") in the target cache description module.
[0613] The target buffer description module is any buffer description module obtained from the buffer list in the scene description file.
[0614] For example, if the first byte length syntax element in a certain buffer description module has the value "byteLength":43972, then the capacity of the buffer corresponding to that buffer description module is 43972 bytes.
[0615] Step 123: Determine whether the target buffer is a circular buffer based on MPEG extensions based on whether the target buffer description module contains an MPEG circular buffer ("MPEG_buffer_circular":{}).
[0616] In some embodiments, determining whether the target buffer is an MPEG-extended circular buffer based on whether the target buffer description module contains an MPEG circular buffer may include: if the target buffer description module contains an MPEG circular buffer, determining that the target buffer is an MPEG-extended circular buffer; and if the target buffer description module does not contain an MPEG circular buffer, determining that the target buffer is not an MPEG-extended circular buffer.
[0617] Step 124: Obtain the number of storage stages in the target buffer based on the value of the stage number syntax element ("count") in the MPEG circular buffer ("MPEG_buffer_circular":{}) of the target buffer description module.
[0618] For example, if the syntax element for the number of stages in the MPEG circular buffer of a certain buffer description module is "count":3, then it can be determined that the buffer corresponding to this buffer description module includes 3 storage stages.
[0619] Step 125: Based on the value of the media index syntax element ("media") in the MPEG circular buffer ("MPEG_buffer_circular":{}) of the target buffer description module, obtain the index value of the media description module corresponding to the media file to which the source data of the data cached by the target buffer belongs.
[0620] Step 126: Obtain the track index value of the source data of the data cached by the target buffer based on the value of the second track index syntax element ("tracks") in the MPEG circular buffer ("MPEG_buffer_circular":{}) of the target buffer description module.
[0621] Step 127: Obtain the data for reconstructing the dynamic 3D model of the corresponding digital human component based on the value of the Uniform Resource Identifier syntax element ("uri") in the target cache description module.
[0622] In some embodiments, the method for parsing scene description files provided in the above embodiments may further include the following steps 13 and 14:
[0623] Step 13: Obtain the target cache slice description module from the cache slice list ("bufferViews":[]) of the scene description file.
[0624] The target cache slice description module can be any cache slice description module in the cache slice list.
[0625] Step 14: Obtain the description information of the target cache slice corresponding to the target cache slice description module according to the target cache slice description module.
[0626] In some embodiments, step 14 above (obtaining the description information of the target cache slice corresponding to the target cache slice description module according to the target cache slice description module) may include at least one of the following steps 141 to 143:
[0627] Step 141: Obtain the capacity of the target cache slice corresponding to the target cache slice description module based on the value of the second byte length syntax element ("byteLength") in the target cache slice description module.
[0628] For example, if the second byte length syntax element in a cache slice description module has the value "byteLength":21936, then the capacity of the cache slice corresponding to that cache slice description module is 21936 bytes.
[0629] Step 142: Obtain the offset of the data cached by the target cache slice based on the value of the offset syntax element ("byteOffset") in the target cache slice description module.
[0630] For example, if the offset syntax element and its value in a cache slice description module are "byteOffset":0, then it can be determined that the offset of the data cached by the cache slice corresponding to the cache description module is 0.
[0631] Step 143: Obtain the name of the target cache slice based on the value of the cache slice name syntax element ("name") in the target cache slice description module.
[0632] In some embodiments, the method for parsing the scene description file may further include: determining the version number of the scene description file based on the version syntax element ("version") and its value in the digital asset description module ("asset":{}) of the scene description file.
[0633] For example, the digital assets in the scene description file are shown below:
[0634]
[0635] Therefore, it can be determined that the scene description file is written based on glTF version 2.0, and the version of the scene description file is the reference version of the scene description standard.
[0636] In some embodiments, the method for parsing the scene description file may further include: obtaining the extensions used by the scene description file according to the extensions used list ("extensionsUsed":[]) in the scene description file.
[0637] For example, the extension usage list ("extensionsUsed":[]) of the scene description file is shown below:
[0638]
[0639]
[0640] Then, the extended items used by the scene description file can be obtained, including: MPEG media, MPEG buffer circular, MPEG accessor timed, and MPEG node avatar.
[0641] Some embodiments of this application also provide a method for rendering a three-dimensional scene, wherein the execution entity of the three-dimensional scene rendering method is the display engine in the immersive media scene description framework, referring to... Figure 12 As shown, the rendering method for this 3D scene may include the following steps S131 to S134:
[0642] S131. Obtain the scene description file of the 3D scene to be rendered.
[0643] In some embodiments, obtaining a scene description file of a 3D scene to be rendered may include: sending a request message to a media resource server for requesting a scene description file of the 3D scene to be rendered, and receiving a request response from the media resource server carrying a scene description file of the 3D scene to be rendered.
[0644] In other embodiments, obtaining the scene description file of the 3D scene to be rendered may include: reading the scene description file of the 3D scene to be rendered from a specified storage space.
[0645] S132. Obtain the description information of the facial marker points of the target digital human in the three-dimensional scene to be rendered and the description information of the first media file according to the scene description file.
[0646] The first media file is a media file that includes the index information and location coordinates of the facial marker points of the target digital human.
[0647] In some embodiments, step S132 (obtaining description information of facial marker points of the target digital human in the three-dimensional scene to be rendered according to the scene description file) may include the following steps 1321 to 1323:
[0648] Step 1321: Obtain the first node description module from the node list ("nodes":[]) of the scene description file of the 3D scene to be rendered.
[0649] The first node description module is the node description module corresponding to the node representing the target digital human.
[0650] Step 1322: Obtain the face marker array ("landmarks":{}) from the digital human node array ("MPEG_node_avatar":{}) of the first node description module.
[0651] Step 1323: Obtain the description information of the facial markers of the target digital human based on the facial marker array ("landmarks":{}).
[0652] The implementation method of step 1323 above (obtaining the description information of the facial marker points of the target digital human based on the facial marker point array) can refer to Figure 14 The implementation of step S123 in the illustrated embodiment will not be described in detail here to avoid redundancy.
[0653] The description information of the facial markers of the target digital human obtained from the facial marker array ("landmarks":{}) may include at least one of the following: the type and number of facial markers of the target digital human, the index value of the accessor description module corresponding to the first accessor used to access the index information of the facial markers of the target digital human, and the index value of the accessor description module corresponding to the second accessor used to access the position coordinates of the facial markers of the target digital human.
[0654] In some embodiments, step S132 (obtaining the description information of the first media file according to the scene description file) may include: obtaining the media description module corresponding to the first media file from the media list ("media":[]) of the MPEG media ("MPEG_media":{}) of the scene description file; and obtaining the description information of the first media file according to the media description module corresponding to the first media file.
[0655] The implementation method for obtaining the description information of the first media file according to the media description module corresponding to the first media file can refer to the implementation method of step 2 in the above embodiment. To avoid redundancy, it will not be described in detail here.
[0656] S133. Send the description information of the facial markers of the target digital human and the description information of the first media file to the media access function.
[0657] After the display engine sends the description information of the facial markers of the target digital human and the description information of the first media file to the media access function, the media access function can obtain the index information and position coordinates of the facial markers of the target digital human based on the description information of the facial markers of the target digital human and the description information of the first media file, reconstruct the dynamic 3D model corresponding to the target digital human based on the index information and position coordinates of the facial markers of the target digital human, and write the dynamic 3D model of the target digital human into the buffer corresponding to the target digital human.
[0658] In some embodiments, the target digital human is composed of multiple digital human components. Reconstructing a dynamic 3D model of the target digital human based on the index information and position coordinates of the facial marker points of the target digital human, and writing the dynamic 3D model of the target digital human into a corresponding buffer, may include: reconstructing a dynamic 3D model of each digital human component of the target digital human based on the index information and position coordinates of the facial marker points of the target digital human, and writing each digital human component of the target digital human into a corresponding buffer.
[0659] In some embodiments, sending description information of the facial markers of the target digital human and description information of the first media file to the media access function may include: sending description information of the facial markers of the target digital human and description information of the first media file to the media access function via the media access function API.
[0660] S134. Read the dynamic 3D model of the target digital human from the cache corresponding to the target digital human, and render the 3D scene to be rendered based on the dynamic 3D model of the target digital human.
[0661] In some embodiments, step S134 (reading the dynamic 3D model of the target digital human from the cache corresponding to the target digital human) may include the following steps 1341 and 1342:
[0662] Step 1341: Obtain the description information of the accessor used to access the dynamic 3D model of the target digital human based on the scene description file.
[0663] In some embodiments, obtaining description information of the accessor for accessing the dynamic 3D model of the target digital human based on the scene description file may include: obtaining the mesh description module corresponding to each digital human component of the target digital human from the mesh list ("meshes":[]) of the scene description file; obtaining the index value declared by the position coordinate syntax element ("position") in the mesh description module corresponding to each digital human component of the target digital human; obtaining the accessor description module corresponding to each digital human component of the target digital human from the accessor list ("accessors":[]) of the scene description file based on the index value declared by the position coordinate syntax element in the mesh description module corresponding to each digital human component of the target digital human; and obtaining the description information of the accessor for accessing the dynamic 3D model corresponding to each digital human component of the target digital human based on the accessor description module corresponding to each digital human component of the target digital human.
[0664] In some embodiments, the accessor's description information may include at least one of the following:
[0665] The accessor's name, the data type of the data accessed by the accessor, the accessor's type, the amount of data accessed by the accessor, whether the accessor is a time-varying accessor based on MPEG extensions, the index value of the cache slice description module corresponding to the cache slice storing the data accessed by the accessor, the index value of the cache slice description module corresponding to the cache slice storing the time-varying parameters of the accessor, and whether the accessor parameters change over time.
[0666] Step 1342: Based on the description information of the accessor used to access the dynamic 3D model corresponding to each digital human component of the target digital human, read the dynamic 3D model corresponding to each digital human component of the target digital human from the cache corresponding to each digital human component of the target digital human.
[0667] The rendering method for a 3D scene provided in some embodiments of this application first obtains a scene description file of the 3D scene to be rendered, then obtains description information of facial marker points of a target digital human in the 3D scene to be rendered and description information of a first media file based on the scene description file, and sends the description information of the facial marker points of the target digital human and the description information of the first media file to a media access function, so that the media access function obtains the index information and position coordinates of the facial marker points of the target digital human based on the description information of the facial marker points of the target digital human and the description information of the first media file, reconstructs the dynamic 3D model corresponding to the target digital human based on the index information and position coordinates of the facial marker points of the target digital human, writes the dynamic 3D model of the target digital human into a buffer corresponding to the target digital human, reads the dynamic 3D model of the target digital human from the buffer corresponding to the target digital human, and renders the 3D scene to be rendered based on the dynamic 3D model of the target digital human. Since the rendering method of some embodiments of this application can obtain the description information of the facial marker points of the target digital human and the description information of the first media file including the index information and position coordinates of the facial marker points of the target digital human from the scene description file, and send the obtained information to the media access function, the media access function can obtain the index information and position coordinates of the facial marker points of the target digital human based on the description information of the facial marker points of the target digital human and the description information of the first media file, and reconstruct the dynamic three-dimensional model corresponding to the target digital human based on the index information and position coordinates of the facial marker points of the target digital human. Therefore, some embodiments of this application can solve the problem that the immersive media scene description framework cannot realize digital human animation and related processing based on facial marker points because the scene description file does not declare the relevant information of facial marker points.
[0668] After obtaining the decoded data of the first media file (which may include the index information and position coordinates of the facial markers of the target digital human) based on the description information of the first media file, the media access function needs to first write the decoded data of the first media file into the buffer corresponding to the first media file, and then retrieve the index information and position coordinates of the facial markers of the target digital human from the buffer corresponding to the first media file. Therefore, the rendering method of the three-dimensional scene provided in some embodiments of this application may further include: obtaining the description information of the first buffer and the description information of the buffer slices of the first buffer based on the scene description file, and sending the description information of the first buffer and the description information of the buffer slices of the first buffer to the media access function, so that the media access function writes the decoded data of the first media file into the first buffer. The first buffer is a buffer used to store the decoded data of the first media file.
[0669] In some embodiments, the description information of the first cache may include at least one of the following:
[0670] The name of the first buffer, the capacity of the first buffer, whether the first buffer is a ring buffer based on MPEG extension, the number of storage links in the first buffer, the index value of the media description module corresponding to the media file (first media file) to which the source data of the data cached by the first buffer belongs, and the track index value of the source data of the data cached by the buffer.
[0671] In some embodiments, the description information of the cache slice of the first cache may include at least one of the following:
[0672] The name of the cache slice, the index value of the cache description module corresponding to the cache to which the cache slice belongs (the first cache), the capacity of the cache slice of the cache, and the offset of the data cached by the cache slice of the cache.
[0673] In some embodiments, sending the description information of the first buffer and the description information of the buffer slice of the first buffer to the media access function may include:
[0674] The description information of the first buffer and the description information of the buffer slice of the first buffer are sent to the media access function through the media access function API.
[0675] As described above, after the media access function writes the decoded data of the first media file into the buffer corresponding to the first media file, it also needs to read the index information of the facial marker points of the target digital human from the buffer corresponding to the first media file. Therefore, the rendering method of the three-dimensional scene may further include: obtaining the accessor description module corresponding to the first accessor from the accessor list ("accessors":[]) of the scene description file, obtaining the description information of the first accessor according to the accessor description module corresponding to the first accessor, and sending the description information of the first accessor to the media access function.
[0676] After sending the description information of the first accessor to the media access function, the media access function can access the index information of the facial marker points of the target digital human from the first buffer according to the description information of the first accessor.
[0677] As described above, after the media access function writes the decoded data of the first media file into the buffer corresponding to the first media file, it also needs to read the position coordinates of the facial marker points of the target digital human from the buffer corresponding to the first media file. Therefore, the method may further include: obtaining the accessor description module corresponding to the second accessor from the accessor list of the scene description file, obtaining the description information of the second accessor according to the accessor description module corresponding to the second accessor, and sending the description information of the second accessor to the media access function.
[0678] After sending the description information of the second accessor to the media access function, the media access function can access the position coordinates of the facial marker points of the target digital human from the first buffer according to the description information of the second accessor.
[0679] In some embodiments, the accessor's description information may include at least one of the following:
[0680] The accessor's name, the data type of the data accessed by the accessor, the accessor's type, the amount of data accessed by the accessor, whether the accessor is a time-varying accessor based on MPEG extensions, the index value of the cache slice description module corresponding to the cache slice storing the data accessed by the accessor, the index value of the cache slice description module corresponding to the cache slice storing the time-varying parameters of the accessor, and whether the accessor parameters change over time.
[0681] Since the media access function writes the decoded data of the first media file into the first buffer corresponding to the first media file, in some embodiments, the method may further include: sending the description information of the first buffer and the description information of the cache slice of the first buffer to the cache management module, so that the cache management module creates a buffer for caching the decoded data of the first media file according to the description information of the first buffer and the description information of the cache slice of the first buffer.
[0682] In some embodiments, sending the description information of the first cache and the description information of the cache slice of the first cache to the cache management module may include: sending the description information of the first cache and the description information of the cache slice of the first cache to the cache management module through the cache API.
[0683] Some embodiments of this application also provide a method for processing scene data of a three-dimensional scene. The execution subject of this method is a media access function in an immersive media scene description framework. (Refer to...) Figure 15 As shown, the method for processing scene data in this 3D scene may include the following steps:
[0684] S141. Receive description information of facial marker points of the target digital human in the three-dimensional scene to be rendered and description information of the first media file sent by the display engine.
[0685] The first media file is a media file that includes the index information and location coordinates of the facial marker points of the target digital human.
[0686] In some embodiments, receiving description information of facial marker points of the target digital human in the three-dimensional scene to be rendered and description information of the first media file sent by the display engine may include: receiving description information of facial marker points of the target digital human in the three-dimensional scene to be rendered and description information of the first media file sent by the display engine through a media access function API.
[0687] In some embodiments, the descriptive information of the facial marker points of the target digital human may include:
[0688] The type and number of facial markers of the target digital human, the index value of the accessor description module corresponding to the first accessor used to access the index information of the facial markers of the target digital human, and the index value of the accessor description module corresponding to the second accessor used to access the position coordinates of the facial markers of the target digital human.
[0689] In some embodiments, the description information of the first media file may include: the Uniform Resource Identifier (URI) of the first media file, the track information of the first media file, the encapsulation type of the first media file, and the encoding / decoding parameters of the first media file.
[0690] S142. Construct a dynamic three-dimensional model of the target digital human based on the description information of the facial markers of the target digital human and the description information of the first media file.
[0691] In some embodiments, the target digital human is composed of multiple digital human components. Constructing a dynamic three-dimensional model of the target digital human may include: constructing dynamic three-dimensional models of each digital human component of the target digital human.
[0692] In some embodiments, step S142 (constructing a dynamic three-dimensional model of the target digital human based on the description information of the facial markers of the target digital human and the description information of the first media file) may include: creating a pipeline corresponding to the target digital human based on the description information of the facial markers of the target digital human and the description information of the first media file, and constructing a dynamic three-dimensional model of the target digital human based on the pipeline corresponding to the target digital human.
[0693] In some embodiments, step S142 (constructing a dynamic three-dimensional model of the target digital human based on the description information of the facial markers of the target digital human and the description information of the first media file) may include: obtaining the decoded data of the first media file based on the description information of the first media file; reading the index information and position coordinates of the facial markers of the target digital human from the decoded data of the first media file based on the description information of the facial markers of the target digital human; and constructing a dynamic three-dimensional model of the target digital human based on the index information and position coordinates of the facial markers of the target digital human.
[0694] In some embodiments, the description information of the first media file may include: a Uniform Resource Identifier (URI) of the first media file, track information of the first media file, encapsulation type of the first media file, and encoding / decoding parameters of the first media file. Obtaining the decoded data of the first media file based on its description information may include: obtaining the first media file based on its URI; decapsulating the first media file according to its encapsulation type and track information to obtain the bitstream of each track; and decoding the bitstream of each track according to its encoding / decoding parameters to obtain the decoded data of the first media file.
[0695] That is, the first media file can be described as an MPEG type media file in the media list ("media":[]) of MPEG media ("MPEG_media":{}), and the first media file declared in MPEG media ("MPEG_media":{}) is indexed by the media file index syntax element and its value ("media":0) in the MPEG circular buffer ("MPEG_buffer_circular":{}) of the buffer description module (buffer) corresponding to the first buffer. In this way, the input data of the media access function is delivered to the media access function.
[0696] In some embodiments, obtaining the first media file based on the URI of the first media file may include: sending a media resource request to a media resource server based on the URI of the first media file; and receiving a media resource response carrying the first media file sent by the media server.
[0697] In some embodiments, obtaining the first media file based on the URI of the first media file may include: reading the first media file from a preset storage space based on the URI of the first media file.
[0698] That is, the first media file can be stored on a cloud server or in local storage space.
[0699] In some embodiments, reading the index information and position coordinates of the facial markers of the target digital human from the decoded data of the first media file based on the description information of the facial markers of the target digital human may include:
[0700] Based on the index value of the accessor description module corresponding to the first accessor used to access the index information of the facial markers of the target digital human in the description information of the facial markers of the target digital human, the description information of the first accessor is obtained, and the index information of the facial markers of the target digital human is read from the decoded data of the first media file based on the description information of the first accessor.
[0701] Based on the index value of the accessor description module corresponding to the second accessor used to access the position coordinates of the facial markers of the target digital human in the description information of the facial markers of the target digital human, the description information of the second accessor is obtained, and the position coordinates of the facial markers of the target digital human are read from the decoded data of the first media file based on the description information of the second accessor.
[0702] In some embodiments, the accessor's description information may include at least one of the following:
[0703] The accessor's name, the data type of the data accessed by the accessor, the accessor's type, the amount of data accessed by the accessor, whether the accessor is a time-varying accessor based on MPEG extensions, the index value of the cache slice description module corresponding to the cache slice storing the data accessed by the accessor, the index value of the cache slice description module corresponding to the cache slice storing the time-varying parameters of the accessor, and whether the accessor parameters change over time.
[0704] S143. Write the dynamic 3D model of the target digital human into the cache corresponding to the target digital human.
[0705] After the media access function writes the dynamic 3D model of the target digital human into the cache corresponding to the target digital human, the display engine can read the dynamic 3D model corresponding to the target digital human from the cache corresponding to the target digital human, and render the 3D scene to be rendered based on the dynamic 3D model corresponding to the target digital human.
[0706] In some embodiments, step S143 (writing the dynamic 3D model of the target digital human into the cache corresponding to the target digital human) may include: receiving description information of the accessor corresponding to the target digital human, description information of the cache corresponding to the target digital human, and description information of the cache slice of the cache corresponding to the target digital human sent by the display engine; and writing the dynamic 3D model of the target digital human into the cache corresponding to the target digital human according to the description information of the accessor corresponding to the target digital human, description information of the cache corresponding to the target digital human, and description information of the cache slice of the cache corresponding to the target digital human.
[0707] In some embodiments, the method for processing the media file may further include: receiving description information of a first buffer and description information of a cache slice of the first buffer sent by the display engine; after obtaining the decoded data of the first media file according to the description information of the first media file, the method may further include: writing the decoded data of the first media file into the first buffer according to the description information of the first buffer and the description information of the cache slice of the first buffer. Wherein, the first buffer is a buffer used to store the decoded data of the first media file;
[0708] In some embodiments, the description information of the cache may include at least one of the following:
[0709] The buffer name, buffer capacity, whether the buffer is an MPEG-based circular buffer, the number of storage stages in the buffer, the index value of the media description module corresponding to the media file to which the source data of the buffered data belongs, and the track index value of the source data of the buffered data.
[0710] In some embodiments, the description information of a cache slice may include at least one of the following:
[0711] The name of the cache slice, the index value of the cache description module corresponding to the cache slice, the capacity of the cache slice of the cache, and the offset of the data cached by the cache slice of the cache.
[0712] In some embodiments, receiving the description information of the first cache and the description information of the cache slice of the first cache sent by the display engine may include:
[0713] The system receives the description information of the first cache and the description information of the cache slice of the first cache sent by the display engine through the media access function API.
[0714] The method for processing scene data of a three-dimensional scene provided in some embodiments of this application may further include: sending the description information of the first cache and the description information of the cache slice of the first cache to a cache management module, so that the cache management module can create the first cache according to the description information of the first cache and the description information of the cache slice of the first cache.
[0715] In some embodiments, sending the description information of the first cache and the description information of the cache slice of the first cache to the cache management module may include: sending the description information of the first cache and the description information of the cache slice of the first cache to the cache management module through the cache API.
[0716] The scene data processing method for a three-dimensional scene provided in some embodiments of this application, after receiving description information of facial marker points of a target digital human in a three-dimensional scene to be rendered and description information of a first media file sent by a display engine, constructs a dynamic three-dimensional model of the target digital human based on the description information of the facial marker points of the target digital human and the description information of the first media file; and writes the dynamic three-dimensional model of the target digital human into a cache corresponding to the target digital human, so that the display engine reads the dynamic three-dimensional model of the target digital human from the cache of the target digital human, and renders the three-dimensional scene to be rendered based on the dynamic three-dimensional model of the target digital human. Since the first media file is a media file that includes the index information and position coordinates of the facial marker points of the target digital human, and the display engine sends the description information of the facial marker points of the target digital human in the three-dimensional scene to be rendered to the media access function, the media access function can obtain the relevant information of the facial marker points of the target digital human based on the description information of the facial marker points of the target digital human in the three-dimensional scene to be rendered sent by the display engine and the description information of the first media file. Then, it can describe the animation and related processing of the digital human based on the relevant information of the facial marker points of the target digital human. Therefore, the embodiments of this application can solve the problem that the immersive media scene description framework cannot be used to realize digital human animation and related processing based on facial marker points because the scene description file does not declare the relevant information of facial marker points.
[0717] Some embodiments of this application also provide a cache management method, wherein the execution subject of the cache management method is the cache management module in the immersive media scene description framework, refer to... Figure 16 As shown, the cache management method may include the following steps S151 and S152:
[0718] S151. Receive the description information of the first cache and the description information of the cache slice of the first cache.
[0719] The first buffer is used to cache the first media file, which is a media file that includes the index information and position coordinates of the facial marker points of the target digital human in the three-dimensional scene to be rendered.
[0720] In some embodiments, the description information of the first cache may include at least one of the following:
[0721] The name of the first buffer, the capacity of the first buffer, whether the first buffer is a ring buffer based on MPEG extension, the number of storage links in the first buffer, the index value of the media description module corresponding to the first media file, and the track index value of the source data of the data cached by the first buffer.
[0722] In some embodiments, the description information of a cache slice may include at least one of the following:
[0723] The name of the cache slice, the index value of the cache description module corresponding to the cache, the capacity of the cache slice, and the offset of the data cached by the cache slice.
[0724] In some embodiments, step S151 (receiving the description information of the first cache and the description information of the cache slice of the first cache) may include: receiving the description information of the first cache and the description information of the cache slice of the first cache sent by the display engine.
[0725] In some embodiments, receiving the description information of the first cache and the description information of the cache slice of the first cache may include: receiving the description information of the first cache and the description information of the cache slice of the first cache sent by the display engine through the cache management API.
[0726] In some embodiments, step S151 (receiving the description information of the first buffer and the description information of the cache slice of the first buffer) may include: receiving the description information of the first buffer and the description information of the cache slice of the first buffer sent by the media access function.
[0727] In some embodiments, receiving the description information of the first cache and the description information of the cache slice of the first cache sent by the media access function may include: receiving the description information of the first cache and the description information of the cache slice of the first cache sent by the media access function through the cache management API.
[0728] S152. Based on the description information of the first buffer and the description information of the cache slice of the first buffer, create the first buffer so that the media access function writes the decoded data of the first media file into the first buffer, and reads the index information and position coordinates of the facial markers of the target digital human from the first buffer based on the description information of the facial markers of the target digital human.
[0729] The cache management method provided in some embodiments of this application further includes:
[0730] The system receives description information of a first accessor and a second accessor sent by the display engine and / or the media access function. The first accessor is an accessor for accessing the index information of the facial marker points of the target digital human, and the second accessor is an accessor for accessing the position coordinates of the facial marker points of the target digital human. The system performs cache management on the first cache according to the description information of the first accessor and the description information of the second accessor.
[0731] In some embodiments, the accessor's description information may include at least one of the following:
[0732] The index value of the cache slice description module corresponding to the cache slice of the data accessed by the storage accessor, the data type of the data accessed by the accessor, the type of the target accessor, the amount of data accessed by the accessor, whether the accessor is a time-varying accessor based on MPEG extensions, the index value of the cache slice description module corresponding to the cache slice of the time-varying parameters of the storage accessor, and whether the parameters of the accessor change over time.
[0733] The cache management method provided in some embodiments of this application may further include: receiving a cache management instruction, the cache management instruction being used to instruct the release of the first cache or the updating of cached data in the first cache; and releasing the first cache or updating the cached data in the first cache in response to the cache management instruction.
[0734] The cache management method provided in some embodiments of this application, upon receiving the description information of the first cache and the description information of the cache slices of the first cache, can create the first cache based on the received description information of the first cache and the description information of the cache slices of the first cache, thereby enabling the media access function to write the decoded data of the first media file into the first cache. Since the first media file is a media file including the index information and position coordinates of the facial marker points of the target digital human in the 3D scene to be rendered, the media access function can write the decoded data of the first media file including the index information and position coordinates of the facial marker points of the target digital human into the first cache, and can read the index information and position coordinates of the facial marker points of the target digital human from the first cache based on the description information of the facial marker points of the target digital human. Therefore, the cache management method provided in some embodiments of this application can solve the problem that the immersive media scene description framework cannot be used to implement digital human animation and related processing based on facial marker points because the scene description file does not declare relevant information about facial marker points.
[0735] In some embodiments of this application, a rendering apparatus for a three-dimensional scene is provided, with reference to... As shown, the rendering device 1600 for this 3D scene includes:
[0736] Acquisition unit 161 is used to acquire the scene description file of the 3D scene to be rendered;
[0737] The parsing unit 162 is used to obtain the description information of the facial marker points of the target digital human in the three-dimensional scene to be rendered and the description information of the first media file according to the scene description file. The first media file is a media file that includes the index information and position coordinates of the facial marker points of the target digital human.
[0738] The sending unit 163 is used to send the description information of the facial markers of the target digital human and the description information of the first media file to the media access function, so that the media access function can obtain the index information and position coordinates of the facial markers of the target digital human based on the description information of the facial markers of the target digital human and the description information of the first media file, reconstruct the dynamic three-dimensional model corresponding to the target digital human based on the index information and position coordinates of the facial markers of the target digital human, and write the dynamic three-dimensional model of the target digital human into the buffer corresponding to the target digital human;
[0739] The rendering unit 164 is used to read the dynamic 3D model of the target digital human from the cache corresponding to the target digital human, and to render the 3D scene to be rendered based on the dynamic 3D model of the target digital human.
[0740] The rendering device for the three-dimensional scene provided in the above embodiments can execute the rendering method for the three-dimensional scene provided in the above embodiments. Its principle and technical effect are similar, and to avoid redundancy, it will not be described in detail here.
[0741] Some embodiments of this application provide an electronic device, which includes:
[0742] Memory, configured to store computer programs;
[0743] The processor is configured to cause the rendering apparatus of the three-dimensional scene to implement the three-dimensional scene rendering method described in any of the above embodiments when a computer program is invoked.
[0744] In some embodiments, this application provides a computer-readable storage medium storing a computer program that, when executed by a computing device, causes the computing device to implement the three-dimensional scene rendering method described in any of the above embodiments.
[0745] In some embodiments, this application provides a computer program product that, when run on a computer, enables the computer to implement the three-dimensional scene rendering method described in any of the above embodiments.
[0746] Some embodiments of this application provide a chip including a processor and a memory. The memory stores programs or instructions executable on the processor, and the processor executes the programs or instructions to perform the rendering method for a three-dimensional scene described in any of the above embodiments. The processor can be a logic circuit, an integrated circuit, a general-purpose processor, etc.
[0747] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0748] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.
Claims
1. A method for rendering a three-dimensional scene, characterized in that, include: Obtain the scene description file of the 3D scene to be rendered; The scene description file is used to obtain the description information of the facial marker points of the target digital human in the three-dimensional scene to be rendered and the description information of the first media file. The first media file is a media file that includes the index information and position coordinates of the facial marker points of the target digital human. The description information of the facial markers of the target digital human and the description information of the first media file are sent to the media access function, so that the media access function can obtain the index information and position coordinates of the facial markers of the target digital human based on the description information of the facial markers of the target digital human and the description information of the first media file, reconstruct the dynamic three-dimensional model corresponding to the target digital human based on the index information and position coordinates of the facial markers of the target digital human, and write the dynamic three-dimensional model of the target digital human into the buffer corresponding to the target digital human; The dynamic 3D model of the target digital human is read from the cache corresponding to the target digital human, and the 3D scene to be rendered is rendered based on the dynamic 3D model of the target digital human.
2. The method according to claim 1, characterized in that, The step of obtaining the description information of the facial marker points of the target digital human in the 3D scene to be rendered according to the scene description file includes: The first node description module is obtained from the node list of the scene description file of the 3D scene to be rendered. The first node description module is the node description module corresponding to the node representing the target digital human. Obtain the facial marker point array from the digital human node array of the first node description module; Descriptive information of the facial markers of the target digital human is obtained based on the facial marker array.
3. The method according to claim 2, characterized in that, The step of obtaining the description information of the facial marker points of the target digital human based on the facial marker point array includes: The type and number of facial markers of the target digital person are obtained based on the value of the level type syntax element in the facial marker array.
4. The method according to claim 2, characterized in that, The step of obtaining the description information of the facial marker points of the target digital human based on the facial marker point array includes: The index value of the accessor description module corresponding to the first accessor is obtained based on the value of the index syntax element in the facial marker array; the first accessor is an accessor used to access the index information of the facial markers of the target digital human.
5. The method according to claim 4, characterized in that, The method further includes: Obtain the accessor description module corresponding to the first accessor from the accessor list in the scene description file; The description information of the first accessor is obtained from the accessor description module corresponding to the first accessor. Send the description information of the first accessor to the media access function.
6. The method according to claim 2, characterized in that, The step of obtaining the description information of the facial marker points of the target digital human based on the facial marker point array includes: The index value of the accessor description module corresponding to the second accessor is obtained based on the value of the position coordinate syntax element in the facial marker array; the second accessor is an accessor used to access the position coordinates of the facial markers of the target digital human.
7. The method according to claim 6, characterized in that, The method further includes: Obtain the accessor description module corresponding to the second accessor from the accessor list in the scene description file; The description information of the second accessor is obtained from the accessor description module corresponding to the second accessor. Send the description information of the second accessor to the media access function.
8. The method according to claim 1, characterized in that, Obtaining the description information of the first media file based on the scene description file includes: Obtain the media description module corresponding to the first media file from the media list of the MPEG media in the scene description file; The description information of the first media file is obtained from the media description module corresponding to the first media file.
9. The method according to claim 1, characterized in that, Sending the description information of the facial markers of the target digital human and the description information of the first media file to the media access function includes: The description information of the facial markers of the target digital human and the description information of the first media file are sent to the media access function through the Media Access Function Application Programming Interface (API).
10. The method according to claim 1, characterized in that, The method further includes: The description information of the first buffer and the description information of the buffer slice of the first buffer are obtained according to the scene description file. The first buffer is a buffer used to store the decoding data of the first media file. The media access function sends the description information of the first buffer and the description information of the buffer slice of the first buffer to the media access function so that the media access function writes the decoded data of the first media file into the first buffer.
11. The method according to claim 10, characterized in that, After obtaining the description information of the first cache and the description information of the cache slices of the first cache according to the scenario description file, the method further includes: The cache management module sends the description information of the first cache and the description information of the cache slice of the first cache to the cache management module, so that the cache management module can create a cache for caching the decoded data of the first media file based on the description information of the first cache and the description information of the cache slice of the first cache.
12. The method according to claim 11, characterized in that, The step of sending the description information of the cache corresponding to each digital human component of the target digital human and the description information of the cache slice of the cache corresponding to each digital human component of the target digital human to the cache management module includes: The cache management module is sent with description information of the cache corresponding to each digital human component of the target digital human and description information of the cache slices of the cache corresponding to each digital human component of the target digital human via the cache API.
13. A rendering device for a three-dimensional scene, characterized in that, include: The acquisition unit is used to acquire the scene description file of the 3D scene to be rendered; The parsing unit is used to obtain the description information of the facial marker points of the target digital human in the three-dimensional scene to be rendered and the description information of the first media file according to the scene description file. The first media file is a media file that includes the index information and position coordinates of the facial marker points of the target digital human. The sending unit is configured to send the description information of the facial markers of the target digital human and the description information of the first media file to the media access function, so that the media access function can obtain the index information and position coordinates of the facial markers of the target digital human based on the description information of the facial markers of the target digital human and the description information of the first media file, reconstruct the dynamic three-dimensional model corresponding to the target digital human based on the index information and position coordinates of the facial markers of the target digital human, and write the dynamic three-dimensional model of the target digital human into the buffer corresponding to the target digital human; The rendering unit is used to read the dynamic 3D model of the target digital human from the cache corresponding to the target digital human, and to render the 3D scene to be rendered based on the dynamic 3D model of the target digital human.
14. An electronic device, characterized in that, include: Memory, configured to store computer programs; The processor is configured to cause the electronic device to implement the rendering method of the three-dimensional scene according to any one of claims 1-12 when a computer program is invoked.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a computing device, causes the computing device to implement the rendering method for a three-dimensional scene as described in any one of claims 1-12.
16. A chip, characterized in that, The chip includes a processor and a memory, the memory being used to store programs or instructions that can run on the processor, and the processor being used to execute the programs or instructions to cause the rendering method of the three-dimensional scene as described in any one of claims 1-12 to be executed.