Avatar mesh encoding in video-based dynamic mesh scheme

By integrating avatar-specific parameters with video-based dynamic mesh compression, the method addresses synchronization issues in current encoding schemes, enhancing data transmission and rendering efficiency for complex animated objects.

WO2025176451A1PCT designated stage Publication Date: 2025-08-28INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/052753
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-23
Filing Date
2025-02-04
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Current video-based dynamic mesh encoding schemes for complex animated objects, such as human models and mechanical devices, require dedicated decoders and renderers, leading to synchronization issues and inefficiencies in data transmission and rendering.

Method used

A method that integrates avatar-specific parameters with video-based dynamic mesh compression, encoding a 3D mesh as a template in a video stream and using metadata to reflect changes in animation, allowing for common decoders and renderers to handle both formats efficiently.

Benefits of technology

This approach reduces data payload and improves transmission efficiency by focusing on static meshes with dynamic animation parameters, enabling seamless integration and rendering of complex animated objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025052753_28082025_PF_FP_ABST
    Figure EP2025052753_28082025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and apparatus to encode, transmit and decode a dynamic mesh sequence comprising objects which can be represented as avatars are proposed. The complex animated objects comprise a 3D mesh and avatar specific parameters for animating and rendering them. The 3D meshes of the avatars are encoded as template meshes, for example as a static mesh in a video-based dynamic mesh compression format. The avatar specific parameters are encoded as metadata associated with frames of the video. The avatars may be inserted, located and transformed, in a 3D scene that is encoded as a dynamic mesh sequence in frames of the video.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AVATAR MESH ENCODING IN VIDEO-BASED DYNAMIC MESH SCHEME

[0002] 1. Technical Field

[0003] The present principles generally relate to the domain of encoding, transmitting, decoding and rendering dynamic mesh sequences comprising complex animated object, like models of human beings, animals or articulated mechanical devices. In particular, the present principles relate to formatting a part of the dynamic mesh as a video-based sequence and a part of the dynamic mesh as avatars.

[0004] 2. Background

[0005] The present section is intended to introduce the reader to various aspects of art, which may be related to various aspects of the present principles that are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present principles. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art.

[0006] A dynamic mesh sequence is a sequence of 3D scenes in which objects are represented as a set of 3D meshes. Encoding a dynamic mesh sequence requires a large amount of data for storage and bandwidth transmission. Video-based schemes to encode dynamic mesh sequences have been set up to reduce the required amount of data. For example, current contributions in MPEG for the signalling of mesh streams allows the encoding of general scenes including complex animated objects (like models of human beings, animals or articulated mechanical devices) using Video-based Dynamic Mesh Compression (V-DMC) specification. In such solutions, dynamic meshes are encoded at each frame of the video-based format. Video-based schemes require dedicated decoders and Tenderers that can re-build the 3D meshes of each frame of the dynamic mesh sequence.

[0007] For the specific case of complex animated objects, representation formats as avatars also exist. In such formats, the complex animated object (CAO) is encoded as one rigged mesh and with a set of parameters. The animation of the rigged mesh is described with semantic parameters. For an animation of a 3D CAO mesh, such a format requires less data than a videobased encoding of the same animation of the same 3D CAO mesh. However, animated 3D CAO meshes cannot be streamed and avatar-based solutions requires dedicated decoders and renderers.

[0008] A solution integrating the two kinds of decoders and renders is very complex as a lot of technical issue would arise from a necessary bridge between the two applications and would introduce synchronicity issues. So, there is a lack of a solution that integrates the two approaches in a single format associated with common decoders and renderers.

[0009] 3. Summary

[0010] The following presents a simplified summary of the present principles to provide a basic understanding of some aspects of the present principles. This summary is not an extensive overview of the present principles. It is not intended to identify key or critical elements of the present principles. The following summary merely presents some aspects of the present principles in a simplified form as a prelude to the more detailed description provided below.

[0011] The present principles relate to a method for encoding a video-based dynamic mesh comprising at least a complex animated object which can be represented as an avatar. The method comprises obtaining a representation of an avatar. This representation comprises a three-dimensional (3D) mesh and avatar specific parameters for animating and rendering the avatar. The 3D mesh of the avatar is encoded as a template mesh in a frame of a video stream, for example as a static mesh according to a video-based dynamic mesh compression format. This may be the first frame of the video. The avatar specific parameters are encoded as first metadata associated with frames of the video stream to reflect changes in the animating and rendering of the avatar. In an embodiment, the at least one avatar is inserted in a 3D scene. In this embodiment, the 3D scene is encoded as a dynamic mesh sequence in frames of the video stream, and second metadata associated with frames of the video stream are encoded for locating and transforming the avatar in the 3D scene. The avatar specific parameters may be encoded as a type of parameters and at least a matrix of values.

[0012] The present principles also relate to a device comprising a processor and a memory associated with the processor that is configured to implement the method above.

[0013] The present principles also relate to a method for decoding a dynamic mesh sequence from a video stream. The method comprises decoding a 3D mesh of an avatar from a template mesh obtained from a frame of the video stream. The template mesh (also called static mesh herein) is encoded according to a video-based dynamic mesh compression format. Avatar specific parameters for animating and rendering the avatar are decoded from first metadata associated with frames of the video stream. Using the decoded 3D mesh and the avatar specific parameters, the avatar is rendered and animated. In an embodiment, a 3D scene is decoded from a dynamic mesh sequence obtained from frames of the video stream. Second metadata associated with frames of the video stream are also decoded for locating and transforming the avatar in the 3D scene. The avatar is located and transformed in the 3D scene when rendering. The avatar specific parameters may be encoded as a type of parameters and at least a matrix of values.

[0014] The present principles also relate to a device comprising a processor and a memory associated with the processor that is configured to implement the method above.

[0015] 4. Brief Description of Drawings

[0016] The present disclosure will be better understood, and other specific features and advantages will emerge upon reading the following description, the description making reference to the annexed drawings wherein:

[0017] - Figure 1 illustrates an example of video-based encoding of a dynamic mesh sequence;

[0018] - diagrammatically depicts modules of a dynamic mesh sequence decoder according to the present principles;

[0019] - Figure 3 shows an example architecture of a device 30 which may be configured to implement encoding and / or decoding methods according to an embodiment of the present principles;

[0020] - Figure 4 shows an example of an embodiment of the syntax of a stream when the data are transmitted over a packet-based transmission protocol;

[0021] - Figure 5 illustrates an example scheme according to the present principles.

[0022] 5. Detailed description of embodiments

[0023] The present principles will be described more fully hereinafter with reference to the accompanying figures, in which examples of the present principles are shown. The present principles may, however, be embodied in many alternate forms and should not be construed as limited to the examples set forth herein. Accordingly, while the present principles are susceptible to various modifications and alternative forms, specific examples thereof are shown by way of examples in the drawings and will herein be described in detail. It should be understood, however, that there is no intent to limit the present principles to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present principles as defined by the claims.

[0024] The terminology used herein is for the purpose of describing particular examples only and is not intended to be limiting of the present principles. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises", "comprising," "includes" and / or "including" when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Moreover, when an element is referred to as being "responsive" or "connected" to another element, it can be directly responsive or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to other element, there are no intervening elements present. As used herein the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as" / ".

[0025] It will be understood that, although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element without departing from the teachings of the present principles.

[0026] Although some of the diagrams include arrows on communication paths to show a primary direction of communication, it is to be understood that communication may occur in the opposite direction to the depicted arrows.

[0027] Some examples are described with regard to block diagrams and operational flowcharts in which each block represents a circuit element, module, or portion of code which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in other implementations, the function(s) noted in the blocks may occur out of the order noted. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. Reference herein to “in accordance with an example” or “in an example” means that a particular feature, structure, or characteristic described in connection with the example can be included in at least one implementation of the present principles. The appearances of the phrase in accordance with an example” or “in an example” in various places in the specification are not necessarily all referring to the same example, nor are separate or alternative examples necessarily mutually exclusive of other examples.

[0028] Reference numerals appearing in the claims are by way of illustration only and shall have no limiting effect on the scope of the claims. While not explicitly described, the present examples and variants may be employed in any combination or sub-combination.

[0029] Figure 1 illustrates an example of video-based encoding of a dynamic mesh sequence. The encoder pre-processes the dynamic mesh content to convert it into components. For example, a base mesh 11, displacement vectors 12, 2D maps for different mesh attributes 13 and a 2D atlas map 14 are generated and encoded by dedicated encoders. Then, the components are multiplexed to generate a bitstream 15. To improve transmission bitrate, input mesh 10 in each frame of the dynamic mesh sequence is simplified to a lower resolution approximation called a base mesh 11, through a surface simplification procedure. A default midpoint subdivision scheme is then applied to iteratively up-sample the base mesh again to reach approximately the same resolution (number of vertices and faces) as the input mesh. The differences between the up-sampled mesh vertex (x, y, z) positions at the highest subdivision (resolution) level and the input mesh vertex positions (where the point correspondences between the input and up-sampled meshes are found by nearest-point searching) are then represented as a set of displacement vectors 12. These displacement vectors indicate how to deform the up-sampled mesh surface to approximate more closely the input mesh surface. The displacement vectors can be encoded by a profile or using SEI messages. The Attribute component 13 provide additional properties (e.g. texture or material data). All these components may be encoded using, for example, standard 2D video coders (e.g. V3C). Atlas component 14 provides information to a decoding and / or rendering system on how to perform inverse reconstruction. For example, how to perform the subdivision of the base mesh and how to apply the displacement vectors to the up-sampled mesh vertices and how to apply attributes to the reconstructed mesh. The decoder comprises similar modules to perform the inverse transformations to reconstruct a dynamic mesh content as close as possible to the input mesh for each frame.

[0030] Such a video-based encoding does not embed semantic features like rig blendshapes weights or color parameters. According to the present principles, some parameters representative of avatars are added to a format for video-based dynamic mesh sequence encoding. These parameters belong to a group comprising: affine translation matrix, covariance matrix, mouth matrix representing mouth motion, eye matrix representing the open close status and level of eyes, head rotation parameters representing the head rotation, head translation matrix representing head translation, head location matrix with size of representing the head location, compact feature matrix, rig parameters, blendshapes, rig blendshapes weights, color parameters, and sub-mesh semantic meaning (currently the sub-meshes in V-DMC are created without taking into account any semantic information about the different parts of the input mesh models).

[0031] Figure 2 diagrammatically depicts modules of a dynamic mesh sequence decoder 20 according to the present principles. In a first step, decoder checks in metadata whether a mesh has to be decoded for this frame and, if so, according to what prediction mode. If a skip flag is present in metadata, the first base mesh in the mesh buffer is retrieve by skip decoder module 21 and is used as the reference frame. No information of the reference frame index is parsed from the bitstream. Motion Decoder 22 extract from the bit stream motion parameters for a mesh and returns it to the reconstruction model with a base mesh to recover the sequence of encoded mesh or sub-mesh frames. It is directly used if the prediction mode is inter-prediction. Reconstruction of Base Mesh module 23 involves “INTRA” and “INTER” sub-meshes reconstruction process. It handles the general reconstruction process of the dynamic meshes given several possible settings. Mesh Buffer 24 is a storage container for a sequence of meshes. Static Mesh Decoder 25 generates a template mesh that can be reused in a sequence of frames and that is stored (cached) in mesh buffer 24 for later processing, using motion decoder 22 and reconstruction modules.

[0032] According to the present principles, Avatar Decoder 26 decodes the feature attributes related to avatar-specific information. Static Mesh Decoder 25 decodes geometrical information (template avatar mesh attributes) and provides it to Avatar Motion Decoder 27. For each sub- mesh or mesh decoded by the static mesh decoder 25, avatar decoder 26 retrieves avatar attributes that provide the ability to change the reconstruction of the base mesh at the stream level. Metadata locating and transforming (i.e. scaling, translating and rotating) the avatar within the 3D scene of the dynamic mesh sequence are also retrieved. Avatar motion decoder 27 receives a decoded static mesh and decoded avatar parameters as input. The avatar specific parameters provide additional information to deform and / or to animate the static mesh, bypassing motion decoder 22, to achieve the same results, for example a new reconstruction of the base-mesh. This pipeline reduces the payload required to transmit dynamic meshes by reducing the number of transmitted parameters and focusing on encoding and decoding of static meshes with dynamic animation parameters (blendshapes, parametric textures etc..). In an embodiment, there is no 3D scene encoded as a dynamic mesh sequence and the avatars are the only objects of the scene. In all embodiments, the mesh of an avatar is encoded as a static mesh in association with a single frame of the video.

[0033] To achieve this, metadata parameters and syntax are proposed below. According to the present principles, mesh attribute types are defined like in base mesh in video-based solutions. bmesh_attribute_type [ j ][ i ] indicates the attribute type of the Attribute Video Data unit with index i for the atlas with atlas index j.

[0034] ATTR TEXTURE indicates an attribute that contains texture information of a volumetric frame. For example, this may indicate an attribute that contains RGB (Red, Green, Blue) color information. ATTR MATERIAL ID indicates an attribute that contains supplemental information that identifies the material type of a point in a volumetric frame. For example, the material type may be used as an indicator for identifying an object or the characteristic of a point within a volumetric frame.

[0035] ATTR TRANSPARENCY indicates an attribute that contains transparency information that is associated with each point in a volumetric frame.

[0036] ATTR REFLECTANCE indicates an attribute that contains reflectance information that is associated with each point in a volumetric frame.

[0037] ATTR NORMAL indicates an attribute that contains a unit vector information associated with each point in a volumetric frame. The unit vector specifies the perpendicular direction to a surface at a point (i.e. direction a point is facing). An attribute frame with this attribute type shall have an ai attribute dimension minusl equal to 2. Each channel of an attribute frame with this attribute type shall contain one component of the unit vector (x, y, z), where the first component contains the x coordinate, the second component contains the y coordinate, and the third component contains the z coordinate.

[0038] ATTR_ TEXTCOORD indicates an attribute that contains texture coordinates of the mesh.

[0039] ATTR_ FACEGROUP ID indicates an attribute that contains face group indices.

[0040] ATTR AVATAR indicates an attribute that contains avatar specific information that is associated with each volumetric frame.

[0041] Values indicated as ATTR RESERVED or ATTR UNSPECIFIED are reserved for future use.

[0042] Mesh attribute types (not the base meshes) may be the following ones:

[0043]

[0044] MESH ATTR AVATAR indicates an attribute that contains avatar specific information that is associated with each mesh or sub-mesh frame sequence. The avatar attribute types may be defined as following: The following syntax is provided as an example for encoding, transmiting, and decoding an avatar in a video-based dynamic mesh sequence. This example is based on V- DMC syntax. A different format may be created or adapted according to the same principles.

[0045]

[0046] When the mesh attribute is of type MESH ATTR AVATAR, avatar specific parameters are provided in a avatar_parameter_set ( ) section, for example by using the following syntax.

[0047]

[0048] In this example syntax, types of avatar attributes are indicated and the attribute vales are encoded in matrices (when avatar_attribute_present_flag is 1). The decoding of the matrices requires the number of matrices (numMatrices[i]), their height (matrixHeight[i]), and their

[0049] 5 width (matrixWidth|i |) These values depend on the matrix type (avatar_attribute_type[i]). avatar_matrix_element_int [ i ] [ j ] [ k ] [ I ] indicates the integer part of the value of the matrix element at position (k, I ) of the j-th matrix of the i-th type. avatar_matrix_element_dec [ i ] [ j ] [ k ] [ I ] indicates the decimal part of the value of the matrix element at position (k, I ) of the j-th matrix of the i-th type. 0 avatar_matrix_element_sing_flag [ i ] [ j ] [ k ] [ I ] indicates the sign of the matrix element at position (k, I ) of the j-th matrix of the i-th type. When the sign is not present, it is inferred to be equal to zero.

[0050] The decoder is based on an avatar based on a mesh rig (including body and face rigs). Such a model has a three-dimensional morphable model, with all the material needed by the model to render a rig, such as vertices, blendshapes, UV maps, etc. So, the avatar can be remorphed and animated from the avatar specific parameters to be rendered.

[0051] Specifications from V-DMC format for base and static meshes may be used for he attribute types with an index between 0 and 6 in the bmesh_attribute_type table.

[0052] In the example syntax above, avatar_parameter_set_id indicates the index of the avatar mesh sub-bitstream associated with the i-th avatar attributes. avatar_attribute_present_flag indicates that the mesh sub-stream contains avatar mesh attributes, avatar attribute count indicates the number of attributes associated with the meshes, avatar attribute count shall be in the range of 0 to 127, inclusive. avatar_attribute_type indicates the attribute type of the attribute with index i for the mesh.

[0053] The encoding provides material_type[i], which may have the following values: 0 for diffuse, 1 for specular, and 2 for translucent. It also provides texture_size[i], which defines the size of the texture. The texture has equal width and height. For this case, numMatrices[i] is 1, matrixHeight[i] is 1, and matrixWidth[i] equals to the material model parameters. The encoding provides one value transparency _type[i]. If it equals to one, the transparency is present, otherwise is opaque. It also provides texture_size[i], which defines the size of the texture. The texture has equal width and height. For this case numMatrices[i] is 1, matrixHeight[i] equals to the texture size, and matrixWidth[i] equals to the texture size. The encoding provides one value reflectance _type[i] . If it equals to one, the reflectance model is present, otherwise there is no model. It also provides texture_size[i], which defines the size of the texture. The texture has equal width and height. For this case numMatrices[i] is 1, matrixHeight[i] is 1, and matrixWidth[i] equals to the model parameters chosen for the reflectance. The encoding provides colors_type[ i ], which may have the following values: 0 for vertex or 1 for faces. For this case, numMatrices[i] is 1, matrixHeight[i] is 1, and matrixWidth[i] equals to the number of vertices or faces. The encoding provides colors_basis_type[ i ], which may have the following values: 0 for vertex or 1 for faces. For this case, numMatrices[i] is 1, matrixHeight[i] is 1, and matrixWidth[i] equals to the number of vertices or faces. The encoding provides one value textcoord_type[i]. If it equals to one, the text coordinates are present, otherwise there is no mapping between vertices and the UV map. For this case, numMatrices[i] is 1, matrixHeight[i] is 1, and matrixWidth[i] equals to the number of vertices. The encoding provides one value facegroup_type[i]. If it equals to one, the faces are present, otherwise there is no face information. For this case, numMatrices[i] is 1, matrixHeight[i] is 1, and matrixWidthfi] equals to the number of polygon faces. The encoding provides two values blendshapes basis [ i ]. If it equals to one, the blendshapes correspond to expression, otherwise it is identity, blendshapes count equals to the number of blendshapes associated with avatar_parameter_set_id. For this case, numMatrices[i] equals to the number of blendshapes, matrixHeight[i] is 1, and matrixWidth[i] equals to the number of vertices. The encoding provides two values blendshapes_type[ i ] (if it equals to one, the blendshapes correspond to expression; otherwise it is identity) and blendshapes weights count minusl [ i ] that is the number of blendshapes weights minus one being provided for pose, expression or identify, depending on blendshapes_type[ i ]. If this number is lower than the total number of corresponding blendshapes, it defines the first blendshapes weights, the remaining being zero. This number cannot be higher than the total number of corresponding blendshapes. For this case, numMatrices[i] is 1, matrixHeight[i] is 1, and matrixWidthfi] equals to blendshapes_weights_count_minusl[ i ] + 1. The encoding provides colors_type[ i ], which may have the following values: 0 for provided colors are for reflectance, 1 for provided colors are for specularity, 2 for provided colors are for roughness, 3 for provided colors are for glossy. For this case, numMatrices[i] is 1, matrixHeight[i] is 1, and matrixWidthfi] equals to the number of vertices. The encoding provides texture_type[i], which may have the following values: 0 for provided texture is for reflectance, 1 for provided texture is for specularity, 2 for provided texture is for roughness, and 3 for provided texture is for glossy. It also provides texture_size[i], which defines the size of the texture. The texture has equal width and height. For this case, numMatrices[i] is 1, matrixHeight[i] equals the texture size, and matrixWidthfi] equals the texture size. The encoding provides colors_weights_type[ i ], which can have the following values: 0 for provided weights are for the parametric vertex colors reflectance, 1 for provided weights are for the parametric vertex colors specularity, 2 for provided weights are for the parametric vertex colors roughness, and 3 for provided weights are for the parametric vertex colors glossy. It also provides colors_weights_count_minusl[i], which defines the number of provided weights minus one. If it is lower than the maximum number, then the remaining weights are zero. It cannot be higher than the maximum number. For this case, numMatrices[i] is 1, matrixHeight[i] is 1, and matrixWidthfi] equals to the number of weights. The encoding provides texture_weights_type[i], which can have the following values: 0 for provided weights are for the parametric reflectance texture, 1 for provided weights are for the parametric specularity texture, 2 for provided weights are for the parametric roughness texture, and 3 for provided weights are for the parametric glossy texture. It also provides texture_weights_count_minusl[i], which defines the number of provided weights minus one. If it is lower than the maximum number, then the remaining weights are zero. It cannot be higher than the maximum number. For this case, numMatrices[i] is 1, matrixHeight[i] is 1, and matrixWidth[i] equals the number of weights. The encoding provides gaze_type[i], which may be 0 for 2D gaze or 1 for 3D gaze. 2D gaze has two values: one for the horizontal rotation of the eyes, and one for the vertical rotation of the eyes, before any transformation of the rig. 3D gaze has three values (x, y, z) that define the 3D points the eyes point to, before any rig transformation. For this case, numMatrices[i] is 1, matrixHeight[i] is 1, and matrixWidth[i] is 2 for 2D gaze and 3 for 3D gaze. The encoding provides semantics_type[i], which can have the following values: 0 for provided semantics are for each vertex, 1 for provided semantics are for each triangular face, 2 for provided semantics are for UV coordinate, 3 for provided semantics are for a single mesh, and 4 for provided semantics are for a sub mesh. For this case, numMatrices[i] is 1, matrixHeight[i] is 1, and matrixWidth[i] equals to the number of parameters given by semantics_type [ i ] (e.g. if equal to zero the size is equal to the number of vertices; if equal to one the size is equal to the number of faces etc.).

[0054] Figure 3 shows an example architecture of a device 30 which may be configured to implement encoding and / or decoding methods according to an embodiment of the present principles. The device is linked with other devices via their bus 31 and / or via I / O interface 36.

[0055] Device 30 comprises following elements that are linked together by a data and address bus 31 :D

[0056] - a processor 32 (or CPU), which is, for example, a DSP (or Digital Signal Processor);

[0057] - a ROM (or Read Only Memory) 33;

[0058] - a RAM (or Random Access Memory) 34;

[0059] - a storage interface 35;

[0060] - an I / O interface 36 for reception of data to transmit, from an application; and

[0061] - a power supply (not represented in Figure 2), e.g. a battery.

[0062] In accordance with an example, the power supply is external to the device. In each of mentioned memory, the word « register » used in the specification may correspond to area of small capacity (some bits) or to very large area (e.g. a whole program or large amount of received or decoded data). The ROM 33 comprises at least a program and parameters. The ROM 33 may store algorithms and instructions to perform techniques in accordance with present principles. When switched on, the CPU 32 uploads the program in the RAM and executes the corresponding instructions. The RAM 34 comprises, in a register, the program executed by the CPU 32 and uploaded after switch-on of the device 30, input data in a register, intermediate data in different states of the method in a register, and other variables used for the execution of the method in a register.

[0063] The implementations described herein may be implemented in, for example, a method or a process, an apparatus, a computer program product, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method or a device), the implementation of features discussed may also be implemented in other forms (for example a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus such as, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.

[0064] Device 30 is linked, for example via bus 31 to a set of sensors 37 and to a set of rendering devices 38. Sensors 37 may be, for example, cameras, microphones, temperature sensors, Inertial Measurement Units, GPS, hygrometry sensors, IR or UV light sensors or wind sensors. Rendering devices 38 may be, for example, displays, speakers, vibrators, heat, fan, etc.

[0065] In accordance with examples, the device 30 is configured to implement a method according to the present principles of encoding, decoding and rendering a 3D point clouds with attributes, and belongs to a set comprising:

[0066] - a mobile device;

[0067] - a communication device;

[0068] - a game device;

[0069] - a tablet (or tablet computer);

[0070] - a laptop;

[0071] - a still picture camera;

[0072] - a video camera. Figure 4 shows an example of an embodiment of the syntax of a stream when the data are transmitted over a packet-based transmission protocol. Figure 4 shows an example structure 4 of a stream encoding dynamic mesh sequences according to the present principle. The structure consists in a container which organizes the stream in independent elements of syntax. The structure may comprise a header part 41 which is a set of data common to every syntax element of the stream. For example, the header part comprises some of metadata about syntax elements, describing the nature and the role of each of them. The structure comprises a pay load comprising an element of syntax 42 and at least one element of syntax 43 (there may be an element of syntax 43 for each type of attribute data, for instance one for the color, one for the reflectance, one for the normal vectors, etc.). Syntax element 42 may comprise metadata describing each frame of the sequence.

[0073] Figure 5 illustrates an example scheme according to the present principles. Realistic 3D mesh sequences can be synthesized using semantic features, using template models like MORGAN (MPEG Reference avatar), for example. In these approaches, there is a parametric face model (the face rig), and a parametric body model (the body rig). The face and body rigs can be built in numerous ways. The solution described herein is an example of technical usecase. Others are possible, as long as the rig is a function of semantic parameters that returns a 3D mesh that can be rendered / encoded / transmitted. A face rig may be limited to the head but can also include upper parts of the body like shoulders. A face and body rig can be made of (but not limited to) a neutral 3D mesh 54 (e.g. a set of vertices and indices (e.g., a topology definition)), a set of identity blendshapes 51 (or morph targets), a set of expression blendshapes 51 (or morph targets), a spherical harmonics model 52, or a mesh for each eye 55. Once combined with weights 52 these rig components 51 lead to a 3D mesh 57 with almost all possible identities and expressions. Combined with a pose (translation, rotation) and the light generated by the spherical harmonics, a rendering of a face 53 is obtained. The rendered face is provided to a decoder 56 that will generate a realistic face.

[0074] Additional inputs can be provided to the decoder, like the driving mesh or the avatar rig parameters. The driving mesh can be one or several previously decoded frames or that will change the style of generated meshes. These additional inputs help decoder 56 to generate meshes 57 of better quality, or to orient the generation in a more appropriate way, depending on the final intent. So, avatar rig parameters may be categorized in three kinds. Highly stable parameters identity blendshapes weights or vertex colors, for example. Moderately stable parameters identify spherical harmonic weights, for instance. Unstable parameters are used for expression blendshapes weights, gaze parameters, pose, for example. Highly stable parameters rarely change. They do not need to be transmitted at every frame. In many cases, they are transmitted only once, for example in header part 41 of Figure 4 or with the first frame data. Moderately stable parameters may change sometimes, so they are transmitted when they change, for example in frame data of some frames. Unstable parameters need to change for every frame. So, they are transmitted with every frame, for example in frame data of syntax element 42 of Figure 4. If they are missing, they may be predicted from previous data since they are semantic and continuous.

[0075] The proposed encoding can do more than just compress general dynamic meshes and is not limited to a reproduction of the original sequence. Focusing the scene onto a prior template of a humanoid representation (avatar) improves compression efficiency and improves the bitstream transmission rate. In an embodiment, a generative neural network and / or traditional computer graphics techniques are used to drive the mesh data at the decoder side once decoded. For example, the present principles can be used in mesh stream in video conferencing use case to always output high-quality frames with more beautiful / pleasant people (meshes). The input video feed can be encoded and translated into an avatar parameter rig that can be paired with a V-DMC avatar mesh stream that provides a geometry to be rendered and rig parameters that can be manipulated to change dynamically the content of the V-DMC stream. This approach avoids encoding large sequences of dynamic mesh animations.

[0076] In the state of the art, an avatar representation is a file (often standalone) or a set of files that have to be integrally loaded before the animation can be rendered. By using the present video-based scheme, an avatar representation can be streamed, that is progressively decoded and the animation can start before the end of the animation is buffered by the decoder / renderer. Since the avatar rig parameters are semantic-based, they can be changed or updated on-the-fly. These changes can be made during the encoding (features are changed before the encoding), during the decoding (the decoded features can be adapted at the decoder side before rendering, that is what is rendered is not exactly what has been encoded), or defined in the stream, for instance, a filter can be set on blendshapes weights to ensure that only positive expressions are generated. Similar approaches allow to adapt the pose (e.g. an avatar always faces the camera), the eye gaze (e.g. the avatar is always looking at the user or at another avatar), or the spherical harmonics weights (e.g. manage lights, to avoid darkness), for example.

[0077] The implementations described herein may be implemented in, for example, a method or a process, an apparatus, a computer program product, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method or a device), the implementation of features discussed may also be implemented in other forms (for example a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus such as, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, Smartphones, tablets, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.

[0078] Implementations of the various processes and features described herein may be embodied in a variety of different equipment or applications, particularly, for example, equipment or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and related texture information and / or depth information. Examples of such equipment include an encoder, a decoder, a post-processor processing output from a decoder, a pre-processor providing input to an encoder, a video coder, a video decoder, a video codec, a web server, a set-top box, a laptop, a personal computer, a cell phone, a PDA, and other communication devices. As should be clear, the equipment may be mobile and even installed in a mobile vehicle.

[0079] Additionally, the methods may be implemented by instructions being performed by a processor, and such instructions (and / or data values produced by an implementation) may be stored on a processor-readable medium such as, for example, an integrated circuit, a software carrier or other storage device such as, for example, a hard disk, a compact diskette (“CD”), an optical disc (such as, for example, a DVD, often referred to as a digital versatile disc or a digital video disc), a random access memory (“RAM”), or a read-only memory (“ROM”). The instructions may form an application program tangibly embodied on a processor-readable medium. Instructions may be, for example, in hardware, firmware, software, or a combination. Instructions may be found in, for example, an operating system, a separate application, or a combination of the two. A processor may be characterized, therefore, as, for example, both a device configured to carry out a process and a device that includes a processor-readable medium (such as a storage device) having instructions for carrying out a process. Further, a processor-readable medium may store, in addition to or in lieu of instructions, data values produced by an implementation.

[0080] As will be evident to one of skill in the art, implementations may produce a variety of signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal may be formatted to carry as data the rules for writing or reading the syntax of a described embodiment, or to carry as data the actual syntax-values written by a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.

[0081] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made. For example, elements of different implementations may be combined, supplemented, modified, or removed to produce other implementations. Additionally, one of ordinary skill will understand that other structures and processes may be substituted for those disclosed and the resulting implementations will perform at least substantially the same function(s), in at least substantially the same way(s), to achieve at least substantially the same result(s) as the implementations disclosed. Accordingly, these and other implementations are contemplated by this application.

Claims

CLAIMS1. A method comprising:- obtaining a representation an avatar, wherein the representation comprises a three- dimensional (3D) mesh and avatar specific parameters for rendering the avatar;- encoding the 3D mesh of the avatar as a template mesh in a frame of a video stream; and- encoding the avatar specific parameters as first metadata associated with frames of the video stream to reflect changes in the rendering of the avatar.

2. The method of claim 1, comprising:- encoding a 3D scene as a dynamic mesh sequence in frames of the video stream; and- encoding second metadata associated with frames of the video stream, second metadata locating and transforming the avatar in the 3D scene.

3. The method of claim 1 or 2, wherein the avatar specific parameters are encoded as a type of parameters and at least a matrix of values.

4. The method of one of claims 1 to 3, wherein the avatar specific parameters belong to a group of types of parameters comprising: affine translation matrix, covariance matrix, mouth matrix representing mouth motion, eye matrix representing an open close status and level of eyes, head rotation parameters, head translation matrix representing head translation, head location matrix with size of representing head location, compact feature matrix, rig parameters, blendshapes, rig blendshapes weights, color parameters, and submesh semantic meaning.

5. A method for decoding a dynamic mesh sequence from a video stream, the method comprising:- decoding a 3D mesh of an avatar from a template mesh obtained from a frame of the video stream;- decoding avatar specific parameters for rendering the avatar from first metadata associated with frames of the video stream; and- rendering the avatar according to the 3D mesh and to the avatar specific parameters.

6. The method of claim 5, comprising:- decoding a 3D scene from a dynamic mesh sequence obtained from frames of the video stream;- decoding second metadata associated with frames of the video stream, second metadata locating and transforming the avatar in the 3D scene; and- locating and transforming the avatar in the 3D scene when rendering.

7. The method of claim 5 or 6, wherein the avatar specific parameters are encoded as a type of parameters and at least a matrix of values.

8. The method of one of claims 5 to 7, wherein the avatar specific parameters are updated or filtered after decoding.

9. The method of one of claim 5 to 8, wherein the avatar specific parameters belong to a group of types of parameters comprising: affine translation matrix, covariance matrix, mouth matrix representing mouth motion, eye matrix representing an open close status and level of eyes, head rotation parameters, head translation matrix representing head translation, head location matrix with size of representing head location, compact feature matrix, rig parameters, blendshapes, rig blendshapes weights, color parameters, and sub-mesh semantic meaning.

10. A device comprising a memory associated with at least one processor configured for:- obtaining a representation an avatar, comprising a three-dimensional (3D) mesh and avatar specific parameters for rendering the avatar;- encoding the 3D mesh of the avatar as a template mesh in a frame of a video stream; and- encoding the avatar specific parameters as first metadata associated with frames of the video stream to reflect changes in the rendering of the avatar.

11. The device of claim 10, wherein the at least one processor is configured for:- encoding a 3D scene as a dynamic mesh sequence in frames of the video stream; and- encoding second metadata associated with frames of the video stream, second metadata locating and transforming the avatar in the 3D scene.

12. The device of claim 10 or 11, wherein the avatar specific parameters are encoded as a type of parameters and at least a matrix of values.

13. The device of one of claims 10 to 12, wherein the avatar specific parameters belong to a group of types of parameters comprising: affine translation matrix, covariance matrix, mouth matrix representing mouth motion, eye matrix representing an open close status and level of eyes, head rotation parameters, head translation matrix representing head translation, head location matrix with size of representing head location, compact feature matrix, rig parameters, blendshapes, rig blendshapes weights, color parameters, and submesh semantic meaning.

14. A device for decoding a dynamic mesh sequence from a video stream, the device comprising a memory associated with at least one processor configured for:- decoding a 3D mesh of an avatar from a template mesh obtained from a frame of the video stream;- decoding avatar specific parameters for rendering the avatar from first metadata associated with frames of the video stream; and- rendering the avatar according to the 3D mesh and to the avatar specific parameters.

15. The device of claim 14, wherein the at least one processor is configured for:- decoding a 3D scene from a dynamic mesh sequence obtained from frames of the video stream;- decoding second metadata associated with frames of the video stream, second metadata locating and transforming the avatar in the 3D scene; and- locating and transforming the avatar in the 3D scene when rendering.

16. The device of claim 14 or 15, wherein the avatar specific parameters are encoded as a type of parameters and at least a matrix of values.

17. The device of one of claims 14 to 16, wherein the avatar specific parameters are updated or filtered after decoding.

18. The device of one of claims 14 to 17, wherein the avatar specific parameters belong to a group of types of parameters comprising: affine translation matrix, covariance matrix, mouth matrix representing mouth motion, eye matrix representing an open close status and level of eyes, head rotation parameters, head translation matrix representing head translation, head location matrix with size of representing head location, compact featurematrix, rig parameters, blendshapes, rig blendshapes weights, color parameters, and submesh semantic meaning.

Citation Information

Patent Citations

  • Base Mesh Data and Motion Information Sub-Stream Format for Video-Based Dynamic Mesh Compression

    US20240022765A1