Accessory representation for scene description
The method and device automatically align and scale accessories in 3D scenes by encoding key points in a node tree, addressing manual scaling issues and ensuring realistic rendering by avoiding clipping and dynamic deformations.
Patent Information
- Application Number
- PCT/EP2025/056033
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-05
- Filing Date
- 2025-03-05
- Publication Date
- 2025-10-09
AI Technical Summary
Existing methods for adapting accessories to animated objects in 3D scenes require manual scaling and alignment, which is time-consuming and lacks automation, leading to issues like scale and interpenetration (clipping) due to dynamic changes in the character's face or body.
A method and device that encode a 3D scene description with a node tree for objects and accessories, using key points to automatically align and scale accessories with characters by aligning corresponding key points, avoiding clipping through collision detection and non-rigid deformation algorithms.
Enables seamless and automated adaptation of accessories to animated objects, improving rendering quality by eliminating clipping and ensuring realistic positioning and scaling, even with dynamic deformations.
Smart Images

Figure EP2025056033_09102025_PF_FP_ABST
Abstract
Description
[0001] ACCESSORY REPRESENTATION FOR SCENE DESCRIPTION
[0002] 1. Technical Field
[0003] The present principles generally relate to the domain of formatting, transmitting and rendering of scene description when the scene comprises animated object which can be completed by accessories. In particular, the present principles relate to formatting scene descriptions in a way allowing the reading devices to automatically adapt an accessory to an animated object.
[0004] 2. Background
[0005] The present section is intended to introduce the reader to various aspects of art, which may be related to various aspects of the present principles that are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present principles. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art.
[0006] Three-dimensional scenes are commonly represented by scene descriptions. A scene description comprises different data comprising arrays of assets (e.g. 3D meshes, texture images, animation description) and a node tree that describes the organization of the elements of the scene. The scene may comprise animated objects (characters, avatars, machines or vehicles for example) and accessories for these animated objects (like hats, mugs, glasses for instance). In the following, the word “character” is used for any of referent animated object. The accessories are generally generic objects designed by different artists, stored in asset libraries and must be adapted to the animated object of the given scene. For instance, if generic glasses (from an available item list in the application) have to be added on a character’s head, they are not automatically adapted on the character. With the existing solution, the glasses would be added as a child node to the character’s node but with no adaptation to its body or facial morphology. That causes scale issue and / or interpenetration (also called clipping) between the 3D elements. To obtain a good rendering quality, the 3D asset of the glasses is manually scaled and aligned to the character’s head in a preprocessing step which requires artistic skills and time. So, there is a lack for a method to dynamically adapt an accessory when the character’s face or body is dynamically changing (e.g. adapting glasses’ position when the character is laughing).
[0007] 3. Summary
[0008] The following presents a simplified summary of the present principles to provide a basic understanding of some aspects of the present principles. This summary is not an extensive overview of the present principles. It is not intended to identify key or critical elements of the present principles. The following summary merely presents some aspects of the present principles in a simplified form as a prelude to the more detailed description provided below.
[0009] Methods, apparatus and data stream are provided to encode a 3D scene comprising a scene description that allows a digital content creation device or a rendering device to seamlessly stitch accessories on characters of the 3D scene. The scene description comprises a node tree for objects like characters and for accessories. Accessory nodes comprise a mesh and key points. The scene description is encoded in a data stream. When an accessory is to be stitched to a character, corresponding key points of the object are retrieved and aligned with the key points of the accessory.
[0010] The present principles relate to a method comprising obtaining a three-dimensional scene comprising one or more objects and one or more accessories. A scene description comprising a node tree is generated for the one or more obj ects and the one or more accessories. The nodes corresponding to the one or more accessories comprise a mesh or a reference to a mesh and one or more key points. The one or more key points comprise a type and localization information. A key point may be a vertex of the mesh, so its localization is indicated by its index. In a variant, the key point is located on a face of the mesh and its localization is indicated by the index of the face and by weights to compute a barycenter. Then, the scene description is encoded in a data stream. In an embodiment, the one or more key points are grouped by type in nodes corresponding to the one or more accessories.
[0011] The present principles also relate to a device comprising a processor and a memory associated with the processor that is configured to implement the method above.
[0012] The present principles also relate to a method comprising obtaining, from a data stream, a scene description comprising a node tree for one or more objects and one or more accessories. The nodes corresponding to the one or more accessories comprise a mesh or a reference to a mesh and one or more first key points. The one or more first key points comprise a type and localization information. A key point may be a vertex of the mesh, so its localization is indicated by its index. In a variant, the key point is located on a face of the mesh and its localization is indicated by the index of the face and by weights to compute a barycenter. When an accessory is to be stitched to an object like a character, corresponding key points of the object are retrieved and aligned with the key points of the accessory.
[0013] The present principles also relate to a device comprising a processor and a memory associated with the processor that is configured to implement the method above.
[0014] The present principles also relate to a data stream comprising scene description comprising a node tree is generated for the one or more obj ects and the one or more accessories. The nodes corresponding to the one or more accessories comprise a mesh or a reference to a mesh and one or more key points. The one or more key points comprise a type and localization information. A key point may be a vertex of the mesh, so its localization is indicated by its index. In a variant, the key point is located on a face of the mesh and its localization is indicated by the index of the face and by weights to compute a barycenter. Then, the scene description is encoded in a data stream. In an embodiment, the one or more key points are grouped by type in nodes corresponding to the one or more accessories.
[0015] 4. Brief Description of Drawings
[0016] The present disclosure will be better understood, and other specific features and advantages will emerge upon reading the following description, the description making reference to the annexed drawings wherein:
[0017] - Figure 1 illustrates the adaptation of accessories on a character;
[0018] - Figure 2 depicts key points necessary for a stitching operation;
[0019] - Figure 3 shows an example architecture of a device which may be configured to implement encoding and / or rendering methods according to an embodiment of the present principles;
[0020] - Figure 4 shows an example of an embodiment of the syntax of a stream when the data are transmitted over a packet-based transmission protocol;
[0021] - Figure 5 illustrates an encoding method 50 according to the present principles;
[0022] - Figure 6 illustrates a method 60 for rendering a 3D scene in which accessories are stitched on characters according to the present principles. 5. Detailed description of embodiments
[0023] The present principles will be described more fully hereinafter with reference to the accompanying figures, in which examples of the present principles are shown. The present principles may, however, be embodied in many alternate forms and should not be construed as limited to the examples set forth herein. Accordingly, while the present principles are susceptible to various modifications and alternative forms, specific examples thereof are shown by way of examples in the drawings and will herein be described in detail. It should be understood, however, that there is no intent to limit the present principles to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present principles as defined by the claims.
[0024] The terminology used herein is for the purpose of describing particular examples only and is not intended to be limiting of the present principles. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises", "comprising," "includes" and / or "including" when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Moreover, when an element is referred to as being "responsive" or "connected" to another element, it can be directly responsive or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to other element, there are no intervening elements present. As used herein the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as" / ".
[0025] It will be understood that, although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element without departing from the teachings of the present principles.
[0026] Although some of the diagrams include arrows on communication paths to show a primary direction of communication, it is to be understood that communication may occur in the opposite direction to the depicted arrows. Some examples are described with regard to block diagrams and operational flowcharts in which each block represents a circuit element, module, or portion of code which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in other implementations, the function(s) noted in the blocks may occur out of the order noted. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending on the functionality involved.
[0027] Reference herein to “in accordance with an example” or “in an example” means that a particular feature, structure, or characteristic described in connection with the example can be included in at least one implementation of the present principles. The appearances of the phrase in accordance with an example” or “in an example” in various places in the specification are not necessarily all referring to the same example, nor are separate or alternative examples necessarily mutually exclusive of other examples.
[0028] Reference numerals appearing in the claims are by way of illustration only and shall have no limiting effect on the scope of the claims. While not explicitly described, the present examples and variants may be employed in any combination or sub-combination.
[0029] Figure 1 illustrates the adaptation of accessories on a character. Accessories are an essential feature of 3D applications handling animated objects like characters. They are separated from the geometry of a digital character and come with their own geometry, animation setup (skeleton and skinning) and appearance. In the example of Figure 1, a basic character model 11 may be equipped with different accessories (a hat, glasses and earrings in the example) to obtain a new appearance 12 of the character. The same accessories may be used for other characters with different morphology, leading to clipping effects as shown on appearance 13.
[0030] Clipping issues also appear when the character’s mesh is animated without updating its accessories. This creates unrealistic situations (e.g. glasses colliding inside the face, etc.). Images taken from these configurations are not ideal, because they lack realism and credibility. Each accessory must be attached to the mesh and resilient to translations, rotations, scaling, and identity and expression deformations. When the character’s position is changing, accessories may have to adapt their geometry to fit the new mesh and copy uniform transformations. According to the principles, to establish an automatically usable link between a character’s mesh and its accessories, processing steps are set up to consider every transformation and deformation applied to the character’s mesh. In order to perform these automatic steps, nodes of the node tree of the scene description are completed with parameters specific to the present principles.
[0031] Figure 2 depicts key points necessary for a stitching operation. Key points are the points on the accessory’s mesh (and on the character’s mesh) that are used to align and scale the accessory with the character. An accessory can be fit in one or several ways. For example, let’s take a coffee mug. One could want to fit the mug into a character’s hand grasping the handle of the mug, one could want that the character grasp it from the side and another could want to place it on a tray. In the first case with the character, it becomes an accessory of the character and can be fit using two different sets of key points (one set defining the handle and another set defining the sides of the mug). In the second case with the tray. It becomes an accessory of the tray and can be fit using a set of key point defining the bottom of the mug. This is to illustrate that the stitching point list of the mug accessory would contain three elements (key points for the handle, key points for the sides and key points for the bottom).
[0032] According to the present principles, nodes of the node tree of the scene description associated with accessories are defined as follows. This extension comprises a name: a user- defined name of this object (ex: “sunglasses”); a description: describes the accessory and its purpose (to be used in UI for instance); a purpose: identifies the main purpose of the accessory (for the application); and stitchings: a list of stitching elements (used to align / scale the accessory in one or several ways). The stitchings list comprises one or several stitching element(s) defined as follows: a description: a semantic description of the key points (ex: “handle of the mug”); a type: an integer (or enumerate value or a character string) that defines the type of key points of the accessory; and keypoints: an array of the key points of the accessory.
[0033] These key points can be defined using two different ways. The vertex index (one single integer value) can be used to reference a point of the mesh of the accessory. Barycentric coordinates (face ID (one integer value) + 3 coefficients (floating values)) can be defined to locate a point on a face of the mesh of the accessory. Using barycentric coordinates leads to a higher bitrate with more values to encode but allows a more accurate position of the key point on the surface of the mesh. While using vertex ID makes the bitrate lower but only allows to position the key point on vertices. In a variant, a key point may also be defined out of the surface of the mesh of the accessory, for example with vertex IDs and Bezier handle vectors. The type of key points is indicated in the format. Each type defines a specific configuration for aligning the key points on the character. Herein, the type of key points may be vertex, barycenter or other (the “other” variant is not developed herein).
[0034] This key point information can be used by any Digital Content Creation (DCC) software or rendering engine (like a game engine) to add, for example, an accessory to a character (or to another element in the scene) and, when needed, to correctly align and / or scale the accessory on the character. The application starts with parsing the input 3D scene, that is parsing information about accessories and parsing information about key points of each accessory. At this step, an accessory is selected and attributed to another element of the scene (typically a character). The object to which the accessory is attributed has corresponding key points. Therefore, the application aligns the accessory with an adapted method, for example by using a rigid alignment function on the two sets of key points in 3D Euclidean space, both sets being composed of the same number of points. Given two sets of points in correspondence, such a function computes a scaling, a rotation, and a translation that define the transform TR that minimizes the sum of squared errors between TR(X) (X being the accessories’ key points) and its corresponding points in Y (Y being the character’s key point). In a further step, clipping is avoided. To do so, the application can, for example, use any collision detection algorithm (to detect clipping) coupled with a non-rigid deformation algorithm using as dense constraints the two sets of key points.
[0035] Example format provided herein is based on the glTF format and is compatible with the MPEG-I-SD effort to extend glTF with MPEG-I SD extensions. Of course, it is understood by a person skilled in the art that the meaning and use of the present principles are generic and can be coded in other formats (XML, USD, ... ). A format for accessory node “MPEG node accessory” can be added to the glTF “node” object as following: When parsing a node of this kind (in the example, the MPEG node accessory extension of a glTF node), an application parses the name, the description, the purpose and the stitchings. The name property is the name of the object provided by the user. The description property describes the accessory and its usage. This is for an illustrative purpose, for example it can be used in a user interface to help the user to understand what the accessory is and what it is useful for. The purpose property contains a string that identifies the purpose of the accessory. It must follow predetermined values, defined by an application or a standard. In the proposed format, this is string type, but an enumeration type or an integer can also be relevant. The stitchings list defines the properties of each stitching element. A stitching element is defined in the table below and can be seen as one potential way to stitch (in the meaning of “attach” or “link”) the accessory to another object or to a character.
[0036] When parsing a stitching element, the application parses the description property which describes the stitching and its usage. This is for illustrative purpose, for example it can be used in a user interface to help the user to understand what the accessory’s stitching is and what it is useful for (ex: “handle” for a mug or “pommel” for a sword). The purpose property contains a string that identifies the purpose of the stitching. It must follow predetermined values, defined by an application or a standard. In the proposed format, this is string type, but an enumeration type or an integer can also be relevant. The keypoints list defines the properties of each key point element. A key point element is defined in the table below and can be seen as one key point (in the meaning of “landmark” or “point of interest”) placed on the accessory’s mesh. The list may be empty. The reason is that if the key points are not declared, the application must get them from another source, for example from a glTF extension called “MPEG mesh rigid linking”. Key points may be defined according to the following syntax:
[0037] When parsing a keypoint element, the application parses the description property which describes the key point and its usage. This is for an illustrative purpose, for example it can be used in a user interface to help the user to understand what key point is and what it is useful for (ex: “left branch of glasses” for a pair of glasses). The purpose property contains a string that identifies the purpose of the key point. It must follow predetermined values, defined by an application or a standard. In the proposed format, this is string type, but an enumeration type or an integer can also be relevant. The type property defines the type of the key point whether it is located on a vertex or placed on the surface of the mesh (or other). An integer type is proposed in this disclosure, but an enumerated value or a string type can also be relevant. If type is 0, then the attribute vertexindex represents the index of the vertex and the attributes faceindex and weights are not defined. If type is 1, then the attribute faceindex represents the index of the face where the key point is located, and weights represents the weights to apply to the vertex positions of the face to determine the location of the key point. The attribute vertexindex is not defined. The “other” type of key points is not developed herein.
[0038] Figure 5 illustrates an encoding method 50 according to the present principles. At a step 51, a 3D scene is obtained. The 3D scene comprises regular objects like characters and accessory objects that are meant to be stitched to the regular objects. At this step, the accessories are not stitched on the characters, they are only possible parts of the 3D scene. The stitching of an accessory on a character depends on the scenario of the 3D scene, at the rendering stage. At a step 52, the scene description is generated. The scene description comprises scene data (meshes, textures, ...) and a node tree that represents the links between the objects. The nodes of the accessory objects are encoded according to the format described upper herein. At a step 53, the generated scene description is encoded in a data stream to be stored and / or transmitted to a rendering device. Figure 6 illustrates a method 60 for rendering a 3D scene in which accessories are stitched on characters according to the present principles. At a step 61, a scene description encoded according to method 50 is obtained. The scene description comprises regular object (like characters) nodes and accessory nodes. At a step 62, according to the 3D scene scenario, when an accessory is to be stitched to a character, the character key points corresponding to the accessory key points are retrieved and paired. At a step 63, the accessory key points are aligned with the paired character key points in accordance with a method following the present principles.
[0039] In the following example, there are two accessories: a coffee mug and a pair of glasses. The MPEG-I SD description of the scene features a mesh (mesh-0) for glasses and a mesh (mesh-1) for the coffee mug. For the mesh of the glasses, a set of three key points (for example points 21a to 21c of Figure 2) is provided in the Stitchings list. Using these three key points and corresponding key points on the character (that may have been designed by a different operator), the application can adapt the glasses’ mesh to an avatar head for example. An example scene description for this scenario is pesented below: vertexindex 1
[0040] "type": 0,
[0041] " vertexindex 12
[0042] "type": 0,
[0043] " vertexindex ": 3 [... ], { ode_accessory”: { "coffee_mug", tion”: “Authentic Californian sun glasses”, e”: “avatar:accessory:food”, gs": [ scription": “Handle” ypoints":[
[0044] "type": 1
[0045] " faceindex ": 17
[0046] "weights": [0.399, 0.41, 0.191] scription": “Side” ypoints":[
[0047] "type": 1,
[0048] " faceindex " : 2451
[0049] "weights": [0.781, 0.11, 0.0.109]
[0050] 'type": 0,
[0051] In another example, the MPEG mesh rigid linking extension is set, enabling to make the correspondence for animation between the two meshes. When defining the MPEG_node_accessory, two different stitchings are proposed. The first one is only composed by a description (e.g. “Stetson hat”) property meaning that it relies on the already defined points in the MPEG mesh rigid linking extension with “correspondence type” 0 and "correspondence_indices": [17,39,87], The second one (“Regular stitching on the head”) is defined as described above using keypoints. vertexindex 1 type": 0, vertexindex 12 type": 0, vertexindex ": 3 1] l_avatar", [ ": ION": 0, OORD 0": 1, 2 Figure 3 shows an example architecture of a device 30 which may be configured to implement encoding and / or rendering methods according to an embodiment of the present principles. The device is linked with other devices via their bus 31 and / or via I / O interface 36.
[0052] Device 30 comprises following elements that are linked together by a data and address bus 31 :D
[0053] - a processor 32 (or CPU), which is, for example, a DSP (or Digital Signal Processor);
[0054] - a ROM (or Read Only Memory) 33;
[0055] - a RAM (or Random Access Memory) 34;
[0056] - a storage interface 35;
[0057] - an I / O interface 36 for reception of data to transmit, from an application; and
[0058] - a power supply (not represented in Figure 2), e.g. a battery.
[0059] In accordance with an example, the power supply is external to the device. In each of mentioned memory, the word « register » used in the specification may correspond to area of small capacity (some bits) or to very large area (e.g. a whole program or large amount of received or decoded data). The ROM 33 comprises at least a program and parameters. The ROM 33 may store algorithms and instructions to perform techniques in accordance with present principles. When switched on, the CPU 32 uploads the program in the RAM and executes the corresponding instructions.
[0060] The RAM 34 comprises, in a register, the program executed by the CPU 32 and uploaded after switch-on of the device 30, input data in a register, intermediate data in different states of the method in a register, and other variables used for the execution of the method in a register.
[0061] The implementations described herein may be implemented in, for example, a method or a process, an apparatus, a computer program product, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method or a device), the implementation of features discussed may also be implemented in other forms (for example a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus such as, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.
[0062] Device 30 is linked, for example via bus 31 to a set of sensors 37 and to a set of rendering devices 38. Sensors 37 may be, for example, cameras, microphones, temperature sensors, Inertial Measurement Units, GPS, hygrometry sensors, IR or UV light sensors or wind sensors. Rendering devices 38 may be, for example, displays, speakers, vibrators, heat, fan, etc.
[0063] In accordance with examples, the device 30 is configured to implement a method according to the present principles of encoding and decoding a scene description of extended reality scene comprising accessories, and belongs to a set comprising:
[0064] - a mobile device;
[0065] - a communication device;
[0066] - a game device;
[0067] - a tablet (or tablet computer);
[0068] - a laptop;
[0069] - a still picture camera;
[0070] - a video camera.
[0071] Figure 4 shows an example of an embodiment of the syntax of a stream when the data are transmitted over a packet-based transmission protocol. Figure 4 shows an example structure 4 of a stream encoding a sequence of point clouds according to the present principle. The structure consists in a container which organizes the stream in independent elements of syntax. The structure may comprise a header part 41 which is a set of data common to every syntax element of the stream. For example, the header part comprises some of metadata about syntax elements, describing the nature and the role of each of them. The structure comprises a pay load comprising an element of syntax 42 and at least one element of syntax 43 (there may be an element of syntax 43 for each type of data, for instance one for the meshes, one for the textures, one for the normal vectors, etc.). Syntax element 42 comprises data representative of scene description, in particular the node tree according to the present principles.
[0072] The implementations described herein may be implemented in, for example, a method or a process, an apparatus, a computer program product, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method or a device), the implementation of features discussed may also be implemented in other forms (for example a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus such as, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, Smartphones, tablets, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.
[0073] Implementations of the various processes and features described herein may be embodied in a variety of different equipment or applications, particularly, for example, equipment or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and related texture information and / or depth information. Examples of such equipment include an encoder, a decoder, a post-processor processing output from a decoder, a pre-processor providing input to an encoder, a video coder, a video decoder, a video codec, a web server, a set-top box, a laptop, a personal computer, a cell phone, a PDA, and other communication devices. As should be clear, the equipment may be mobile and even installed in a mobile vehicle.
[0074] Additionally, the methods may be implemented by instructions being performed by a processor, and such instructions (and / or data values produced by an implementation) may be stored on a processor-readable medium such as, for example, an integrated circuit, a software carrier or other storage device such as, for example, a hard disk, a compact diskette (“CD”), an optical disc (such as, for example, a DVD, often referred to as a digital versatile disc or a digital video disc), a random access memory (“RAM”), or a read-only memory (“ROM”). The instructions may form an application program tangibly embodied on a processor-readable medium. Instructions may be, for example, in hardware, firmware, software, or a combination. Instructions may be found in, for example, an operating system, a separate application, or a combination of the two. A processor may be characterized, therefore, as, for example, both a device configured to carry out a process and a device that includes a processor-readable medium (such as a storage device) having instructions for carrying out a process. Further, a processor-readable medium may store, in addition to or in lieu of instructions, data values produced by an implementation. As will be evident to one of skill in the art, implementations may produce a variety of signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal may be formatted to carry as data the rules for writing or reading the syntax of a described embodiment, or to carry as data the actual syntax-values written by a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.
[0075] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made. For example, elements of different implementations may be combined, supplemented, modified, or removed to produce other implementations. Additionally, one of ordinary skill will understand that other structures and processes may be substituted for those disclosed and the resulting implementations will perform at least substantially the same function(s), in at least substantially the same way(s), to achieve at least substantially the same result(s) as the implementations disclosed. Accordingly, these and other implementations are contemplated by this application.
Claims
CLAIMS1. A method comprising:- obtaining (51) a three-dimensional scene comprising one or more objects and one or more accessories;- generating a scene description (52) comprising a node tree for the one or more objects and the one or more accessories, wherein nodes corresponding to the one or more accessories comprising information relative to a mesh and one or more key points, the one or more key points comprising a type and localization information; and- encoding the scene description (53) in a data stream.
2. The method of claim 1, wherein the type of the one or more key points indicates that the one or more key points is a vertex of the mesh and the localization information comprises an index of the vertex or wherein the type of the one or more key points indicates that the one or more key points is a barycenter on a face of the mesh and the localization information comprises an index of the face and weights for the barycenter.
3. The method of claim 1 or 2, wherein the one or more key points are grouped by type in nodes corresponding to the one or more accessories.
4. The method of one of claims 1 to 3, wherein nodes corresponding to the one or more accessories comprise information relative to semantics of the one or more accessories indicating relationships between the one or more accessories and other objects in the three- dimensional scene.
5. The method of one of claims 1 to 4, wherein the one or more key points of an accessory comprise a semantic information indicating how the one or more key points can stich with key points of other objects of the three-dimensional scene.
6. A device comprising a memory associated with one or more processor configured for:- obtaining a three-dimensional scene comprising one or more objects and one or more accessories;- generating a scene description comprising a node tree for the one or more objects and the one or more accessories, wherein nodes corresponding to the one or moreaccessories comprising information relative to a mesh and one or more key points, the one or more key points comprising a type and localization information; and- encoding the scene description in a data stream.
7. The device of claim 6, wherein the type of the one or more key points indicates that the one or more key points is a vertex of the mesh and the localization information comprises an index of the vertex or wherein the type of the one or more key points indicates that the one or more key points is a barycenter on a face of the mesh and the localization information comprises an index of the face and weights for the barycenter.
8. The device of claim 6 or 7, wherein the one or more key points are grouped by type in nodes corresponding to the one or more accessories.
9. The device of one of claims 6 to 8, wherein nodes corresponding to the one or more accessories comprise information relative to semantics of the one or more accessories indicating relationships between the one or more accessories and other objects in the three- dimensional scene.
10. The device of one of claims 6 to 9, wherein the one or more key points of an accessory comprise a semantic information indicating how the one or more key points can stich with key points of other objects of the three-dimensional scene.
11. A method comprising:- obtaining (61), from a data stream, a scene description comprising a node tree for one or more objects and one or more accessories, wherein nodes corresponding to the one or more accessories comprising information relative to a mesh and one or more first key points, the one or more first key points comprising a type and localization information;- when an accessory of the one or more accessories is to be stitched to an object of the one or more objects, retrieving (62) second key points on the object, the second key points corresponding to the first key points; and- aligning (63) the first key points with the second key points.
12. The method of claim 11, wherein the type of the one or more first key points indicates that the one or more first key points is a vertex of the mesh and the localization information comprises an index of the vertex or wherein the type of the one or more first key pointsindicates that the one or more first key points is a barycenter on a face of the mesh and the localization information comprises an index of the face and weights for the barycenter.
13. The method of claim 11 or 12, wherein the one or more first key points are grouped by type in nodes corresponding to the one or more accessories.
14. The method of one of claims 11 to 13, wherein nodes corresponding to the one or more accessories comprise information relative to semantics of the one or more accessories indicating relationships between the one or more accessories and other objects in the three- dimensional scene.
15. The method of one of claims 11 to 14, wherein the one or more key points of an accessory comprise a semantic information indicating how the one or more key points can stich with key points of other objects of the three-dimensional scene.
16. A device comprising a memory associated with one or more processor configured for:- obtaining, from a data stream, a scene description comprising a node tree for one or more objects and one or more accessories, wherein nodes corresponding to the one or more accessories comprising information relative to a mesh and one or more first key points, the one or more first key points comprising a type and localization information;- when an accessory of the one or more accessories is to be stitched to an object of the one or more objects, retrieving second key points on the object, the second key points corresponding to the first key points; and- aligning the first key points with the second key points.
17. The device of claim 16, wherein the type of the one or more first key points indicates that the one or more first key points is a vertex of the mesh and the localization information comprises an index of the vertex or wherein the type of the one or more first key points indicates that the one or more first key points is a barycenter on a face of the mesh and the localization information comprises an index of the face and weights for the barycenter.
18. The device of claim 16 or 17, wherein the one or more first key points are grouped by type in nodes corresponding to the one or more accessories.
19. The device of one of claims 16 to 18, wherein nodes corresponding to the one or more accessories comprise information relative to semantics of the one or more accessoriesindicating relationships between the one or more accessories and other objects in the three- dimensional scene.
20. The device of one of claims 16 to 19, wherein the one or more key points of an accessory comprise a semantic information indicating how the one or more key points can stich with key points of other objects of the three-dimensional scene.
21. A data stream comprising a scene description comprising a node tree for one or more objects and one or more accessories, wherein nodes corresponding to the one or more accessories comprising information relative to a mesh and one or more first key points, the one or more first key points comprising a type and localization information.
22. The data stream of claim 21, wherein the type of the one or more first key points indicates that the one or more first key points is a vertex of the mesh and the localization information comprises an index of the vertex or wherein the type of the one or more first key points indicates that the one or more first key points is a barycenter on a face of the mesh and the localization information comprises an index of the face and weights for the barycenter.
23. The data stream of claim 21 or 22, wherein the one or more first key points are grouped by type in nodes corresponding to the one or more accessories.
Citation Information
Patent Citations
3D model clothing image processing method and device and electronic equipment
CN115661324A
Method for providing immersive fashion metaverse service, device therefor, and non-transitory computer-readable recording medium for storing program for executing same method
WO2023096208A1