Animation mixer in scene descriptions

By grouping animation channels by node and property kind and applying predefined mixing strategies, the method addresses the inefficiencies in existing scene description formats, enabling efficient and seamless mixing of animations in 3D scenes without increasing memory or computational complexity.

WO2025214721A1PCT designated stage Publication Date: 2025-10-16INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/057185
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-09
Filing Date
2025-03-17
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Current scene description formats do not provide a generic solution for representing the mixing of different animations in a 3D scene, leading to increased memory usage and computational complexity, especially when multiple animations target the same properties of a node.

Method used

A method for mixing animations by grouping channels by targeted node and property kind, creating mixer channels that merge updates using predefined mixing strategies, and applying these strategies to hardware nodes without modifying the node hierarchy.

Benefits of technology

This approach reduces memory usage and computational complexity while enabling seamless mixing of animations, allowing for efficient rendering of complex 3D scenes with multiple animations targeting the same properties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025057185_16102025_PF_FP_ABST
    Figure EP2025057185_16102025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, apparatus and data stream are provided to mix two or more animations of a 3D scene. A list of two or more animations is obtained. Each animation comprises a list of one or more animation channels, and each channel targets one property of one node of a scene description. The channels of the list of animations is mixed to obtain a list of one or more mixer channels. This is performed by grouping the animation channels by targeted node and by property kind. Then, the mixer channels are merged to obtain a mixer animation comprising a list of one or more merged animation channels. Each merged channel targets a kind of property of one node of the scene description.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] ANIMATION MIXER IN SCENE DESCRIPTIONS

[0002] 1. Technical Field

[0003] The present principles generally relate to the domain of encoding for 3D scene representations. In particular, the present principles relate to the encoding of how different animations of a 3D scene can be seamlessly mixed.

[0004] 2. Background

[0005] The present section is intended to introduce the reader to various aspects of art, which may be related to various aspects of the present principles that are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present principles. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art.

[0006] Current Scene Description (SD) formats (like MPEG-I Scene Description) allow the encoding of animations (for example in an animations asset section of the scene description). Each animation is made of one or more channels. Each channel refers to a sampler. Each channel targets a property of the scene to animate and references a sampler that computes the property data for each animation frame. Channels can target different properties of the scene, but two channels of an animation cannot target the same property of a given node. For instance, two channels of an animation cannot target both the “translation” property of a node. Thanks to the information contained in the channels, a scene can be animated using the animations.

[0007] However, there is a need to improve the state of the art to set up a generic solution that allow to represent the mixing of different animations in a scene description.

[0008] 3. Summary

[0009] The following presents a simplified summary of the present principles to provide a basic understanding of some aspects of the present principles. This summary is not an extensive overview of the present principles. It is not intended to identify key or critical elements of the present principles. The following summary merely presents some aspects of the present principles in a simplified form as a prelude to the more detailed description provided below. The present principles relate to a method for mixing two or more animations of a 3D scene. A list of two or more animations is obtained. Each animation comprises a list of one or more animation channels, and each channel targets one property of one node of a scene description. The channels of the list of animations is mixed to obtain a list of one or more mixer channels. This is performed by grouping the animation channels by targeted node and by property kind. Then, the mixer channels are merged to obtain a mixer animation comprising a list of one or more merged animation channels. Each merged channel targets a kind of property of one node of the scene description.

[0010] The present principles also relate to a device comprising a processor and a memory associated with the processor that is configured to implement the above method according to any one of the embodiments described herein.

[0011] The present principles also relate to a data stream comprising data representative of a scene description, data representative of two or more animations, each animation comprising a list of one or more animation channels, wherein each channel targets one property of one node of a scene description, and data representative of a mixing strategy for one or more property kinds.

[0012] 4. Brief Description of Drawings

[0013] The present disclosure will be better understood, and other specific features and advantages will emerge upon reading the following description, the description making reference to the annexed drawings wherein:

[0014] - Figure 1 illustrates a method for mixing animations of a 3D scene according to the present principles;

[0015] - Figure 2 diagrammatically depicts an example in which two animations to move the eyes of an avatar horizontally and vertically are mixed according to the present principles;

[0016] - Figure 3 shows an example architecture of a device which may be configured to implement animation mixing methods according to embodiments of the present principles;

[0017] - Figure 4 shows an example of an embodiment of the syntax of a stream when the data are transmitted over a packet-based transmission protocol. 5. Detailed description of embodiments

[0018] The present principles will be described more fully hereinafter with reference to the accompanying figures, in which examples of the present principles are shown. The present principles may, however, be embodied in many alternate forms and should not be construed as limited to the examples set forth herein. Accordingly, while the present principles are susceptible to various modifications and alternative forms, specific examples thereof are shown by way of examples in the drawings and will herein be described in detail. It should be understood, however, that there is no intent to limit the present principles to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present principles as defined by the claims.

[0019] The terminology used herein is for the purpose of describing particular examples only and is not intended to be limiting of the present principles. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises", "comprising," "includes" and / or "including" when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Moreover, when an element is referred to as being "responsive" or "connected" to another element, it can be directly responsive or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to other element, there are no intervening elements present. As used herein the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as" / ".

[0020] It will be understood that, although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element without departing from the teachings of the present principles.

[0021] Although some of the diagrams include arrows on communication paths to show a primary direction of communication, it is to be understood that communication may occur in the opposite direction to the depicted arrows. Some examples are described with regard to block diagrams and operational flowcharts in which each block represents a circuit element, module, or portion of code which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in other implementations, the function(s) noted in the blocks may occur out of the order noted. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending on the functionality involved.

[0022] Reference herein to “in accordance with an example” or “in an example” means that a particular feature, structure, or characteristic described in connection with the example can be included in at least one implementation of the present principles. The appearances of the phrase in accordance with an example” or “in an example” in various places in the specification are not necessarily all referring to the same example, nor are separate or alternative examples necessarily mutually exclusive of other examples.

[0023] Reference numerals appearing in the claims are by way of illustration only and shall have no limiting effect on the scope of the claims. While not explicitly described, the present examples and variants may be employed in any combination or sub-combination.

[0024] Scene descriptions (like in the glTF 2.0 specification) do not specify whether several animations can be played at the same time and if so, how the animations can be mixed. Hence, rendering applications apply proprietary mixing methods. A default mixing strategy for animations that do not target the same scene components can be to create a new animation made of the channels from all animations. The resulting animation is then a standard animation that an application can run following the default rules. However, for animations with channels that target the same properties, no default method is generic. A mixing strategy may be based on how the animation works in glTF, where each channel replaces the component that it targets in the order they are listed. For instance, if a channel targets the “translation” property of a node, applying the corresponding animation means replacing any existing translation property with the one computed by the channel. In the mixing case, if a first channel targeting the same component and has already changed it, the method erases the previous change and replaces it with the new data. For instance, if a first animation which changes the translation property of node 0, and a second animation which also changes the translation property of node 0 ignores the first translation. So, the final translation property of node 0 will be the one of the second animation, as if the channel of the first animation which updates the translation of node 0 never existed. To avoid this drawback, new nodes can be introduced in the node tree of the scene description. For instance, if two rotations apply on the same node, with one around axis X (in animation 1) and the other around axis Y (in animation 2). With a replacement strategy, like the one presented in the previous paragraph, only the rotation around axis Y is applied. However, if a new node is inserted on top of the node to rotate, and that the second animation channel targets this new node, then the expected behavior applies. This solution is convenient for very simple scenes. For the complex ones, many new nodes must be inserted leading to a node tree very difficult to manage. Furthermore, it leads to high memory usage and computational complexity.

[0025] Figure 1 illustrates a method 10 for mixing animations of a 3D scene according to the present principles. A scene description (in many formats like glTF), is based on a hierarchy of nodes (also called node tree). Each node comprises, for example, 3D scene data (like a 3D mesh) or a reference to 3D data stored at another place, transform properties (like translation, rotation, etc.) for initial positioning and references to child nodes.

[0026] A common implementation of this representation is to map each node to a hardware equivalent. As a result, nodes of the node tree can be seen as software nodes and their hardware equivalent as hardware nodes. For instance, when using OpenGL, a node with a mesh is usually represented by a structure with Vertex Buffer Arrays (VBAs) which contains all the mesh data in the GPU. Transform properties are represented as a matrix stored as a uniform shader variable. Similar mapping can be found in other implementations, like the ones for Unity or Unreal Engine, in which case engine nodes are used as hardware nodes. Since running an animation corresponds to the update of software nodes, then most implementations directly update the equivalent hardware nodes. For example, morph target weights can be directly updated in the corresponding OpenGL data structure. For animations that change a transform property (rotation, translation, etc.), the value of the property is first converted into a transform matrix, and then copied into the shader variable.

[0027] According to the present principles, two or more animations are obtained from a file or a data stream. Each animation comprises one or more animation channels, and each channel targets one property of one node of a scene description. Animation mixing approach 10 comprises a step 11 that mix the two or more animations 13a- 13b of the 3D scene into a new list 14 of independent mixer channels, where each mixer channel modifies a different node. So, the node hierarchy is not modified on the hardware side. Thus, memory usage or computational complexity are not increased. The examples presented herein are based on glTF, but the present principles are generic and can be coded in other formats (e.g. XML, USD, . . . ).

[0028] Mixer channels 15 are defined as components resulting from the mixing of at least two animations. Like animation channels, each mixer channel modifies a single node. However, a mixer channel can run several updates on the node, and in an embodiment, several times the same property. Updates of the same property kind are mixed before updating the corresponding hardware node. For instance, if a mixer channel updates the translation, then the rotation, the translation again and finally the scale, the resulting transform matrix is computed and transmitted to the hardware node. For instance, morph target weights constitute a different property kind.

[0029] According to the present principles, a mixer channel is a list of mixer updates. Each mixer update 16 targets a node property (like standard animation channels). They do not reference a node, since they are embedded in a mixer channel that already references a node. Furthermore, a mixer update is associated with a mixing strategy, which defines how the property is mixed with the previous value. The following non-exhaustive list of strategies is proposed. Replace: the previous value is replaced by the current one. Add: the previous value is added to the current one. Multiply: the previous value is multiplied with the current one. The multiplication is consistent with the property type. For instance, for scale, it uses element-wise multiplication, for rotations a quaternion multiplication and for matrices a matrix multiplication. Mixing strategies of a mixer update can be set in different manners.

[0030] Two main update kinds are defined according to the present principles. Node transform update kind: updates for node transform (“rotation”, “translation”, “scale” and “matrix” properties of glTF nodes); and weights update kind: updates for the morph target weights. More update kinds can be considered for other properties that an animation channel can update (like in future glTF versions or in extensions like KHR_animation_pointer).

[0031] At a step 12, the mixer updates are merged. All mixer updates 16 in a mixer channel 15 of the same property kind are merged into a single one. After merging step 12, resulting animation channels 18 are equivalent to standard animation channels. Consequently, final mixer animation 17 is a list of animation channels where each channel updates a different node property in the scene. This mixer animation can be run as any other standard animation, where each channel replaces the node property it targets.

[0032] In the example of Figure 1, a first node can be rotated, and a second node has morph targets. At step 11, they are mixed into a list of mixer channels. There are two mixer channels 15 since animations target two different nodes: there is one mixer channel per node. Each mixer channel targets a specific node and holds a list of mixer updates. In the first mixer channel, the two updates rotate node 0. The first rotation (around axis X) comes from the first animation 13a, and the second rotation (around axis Y) comes from the second animation 13b. Similarly, the mixer updates in the second mixer channel change the weights with a value A, and then with a value B. Then, at step 12, all mixer updates in a mixer channel of the same type are merged. In this example, the first mixer channel has only node transform updates: the merge is a standard animation channel that targets the “rotation” property of node 0. The values of this rotation are the quatemionic product of the rotation around axis X and the rotation around axis Y. Similarly, the second standard animation channel of the final mixer animation updates the “weights” property computed as the sum of weights A and B.

[0033] Mixer updates can be merged in different ways. Mixing strategies can be set by default. For example, default mixing strategy can be set as following: for “scale” property and “rotation” property: Multiply mixing strategy; for “translation” property and “weights” property: Add mixing strategy. For other properties, the “Replace mixing strategy” is chosen unless specified. In other cases, the chosen strategies must be encoded in the input scene description, so the device in charge of the animation mixing can apply the mixing strategy decided by the designers. For glTF format, a glTF extension “MPEG animation mixer” that defines what strategy must be used to merge mixer updates can be introduced according to the present principles. Mixer updates are not objects contained in the glTF, but objects created from standard glTF animation channels by the mixer device. Consequently, “MPEG animation mixer” property is to be placed in the “extensions” property of glTF animation channels. An example of such a format is proposed below in relation to Figure 2. In this example, the extension contains a unique property “mixingStrategies” that is a plain JSON object, where each key corresponds to a component that a channel can target and each value is a mixing strategy. In glTF 2.0, the channel targets are limited to “rotation”, “scale”, “translation”, “matrix” and “weights”. In future glTF versions, any other new allowed value can be used as a key in “mixingStrategies”. It is also possible to use any value other extensions might allow, like the ones allowed by KHR_animation_pointer. Possible values for “mixingStrategies” items are: “REPLACE” for a Replace mixing strategy; “ADD” for a Add mixing strategy; “MUL” for a Multiply mixing strategy. The “MPEG animation mixer” can also be used in the “extensions” property of a standard glTF animation. In this case, it defines a default mixing strategy for all channels in the animation. For instance, if there is an “extension / MPEG_animation_mixer / mixingStrategies / rotation” with value “ADD” in a glTF animation, and none of the animation channels define a mixing strategy for “rotation”, then the mixing strategy for “rotation” of all these animation channels is “ADD”. However, if one animation channel defines a mixing strategy for “rotation”, then the definition of the animation scope is ignored. If a channel target is not listed in “mixingStrategies”, then the mixing strategy for this target is inherited if an ascendant item in the glTF hierarchy defines one, otherwise, it is the default strategy associated with the target. The “MPEG animation mixer” extension can be used in (path are JSON path from the root of the glTF): ’7animations / | animation index] / extensions”: overrides default mixing strategies for animation [animation index]; and “ / animations / [animation index] / channels / [channel index] / extensions”: overrides default mixing strategies for channel [channel index] in animation [animation index].

[0034] Figure 2 diagrammatically depicts an example in which two animations to move the eyes of an avatar horizontally and vertically are mixed according to the present principles. Two animations are prepared.

[0035] In this example, the gaze of an avatar is defined using up to three angles. Each angle is associated with an animation, and when using at least two angles, animation mixing is required. The controller’s disclosure assumes that the application handles how to mix the animations in the controllers. A mixing procedure is available, and mixing strategies can be specified in the animations and animation channels. Furthermore, the “gaze” property of “MPEG_ node” extensions can have an “extensions” property with a “MPEG animation mixer” property that overrides default mixing strategies for all animations used for the avatar gaze. The “MPEG animation mixer” extension can also be used in properties “horizontal”, “vertical” and “roll” to override default mixing strategies for each gaze angle. The “MPEG animation mixer” extension can also be used in (path are JSON path from the root of the glTF): “ / nodes / [node index] / extensions / MPEG_node_avatar / gaze”: overrides default mixing strategies for all animations used for avatar gaze in node [node index]; “ / nodes / [node index] / extensions / MPEG_node_avatar / gaze / horizontal / extensions”: overrides default mixing strategies for the animation used for horizontal avatar gaze in node [node index]; “ / nodes / [node index] / extensions / MPEG_node_avatar / gaze / vertical / extensions”: overrides default mixing strategies for the animation used for vertical avatar gaze in node [node index]; and ‘7nodes / [node index] / extensions / MPEG_node_avatar / gaze / roll / extensions”: overrides default mixing strategies for the animation used for roll avatar gaze in node [node index].

[0036] The first animation animates the eyes from left to right: the center row of Figure 2 shows the character state when time tl=Os (left), when time tl=ls (center) and when time tl=2s (right). The animation updates the head with two morph targets that move the eyelids, the left eye with a rotation around the Y axis, and the right eye with a rotation around the Y axis. The second animation animates the eyes from bottom to top: the center column of Figure 2 shows the character state when time t2=0s (bottom), when time t2=ls (center) and when time t2=2s (top). The animation updates the head with two morph targets that move the eyelids, the left eye with a rotation around the X axis, and the right eye with a rotation around the X axis. Both rotations are separated as they target a different node. These two animations can be mixed and merged using the present principles. According to one of the proposed mixing strategies, a mixer animation leads to an animation as depicted on the diagonals of Figure 2. For example, the top-right character in Figure 2 shows a state where both animations are at the same time, e.g., tl = 2s and t2=2s. Mixing and merging is not limited to the same time for all animations, for instance, the top avatar in Figure 2 is for first animation time tl=ls and second animation time t2=2s. Such a 3D scene can be encoded according to the following example: odes": [0] xtures": [{ ource": 0 ource": 1 ource": 2 imations": [{ hannels": [{ sampler": 0, target": {

[0037] "node": 1, "path": "weights" sampler": 1, target": {

[0038] "node": 2,

[0039] "path": "rotation" sampler": 2, target": {

[0040] "node": 3,

[0041] "path": "rotation" , ame": "gaze_0", amplers": [{ input": 15, interpolation": "CUBIC SPLINE", output": 16 input": 15, interpolation": "LINEAR", output": 17 input": 15, interpolation": "LINEAR", output": 17 , xtensions": { MPEG animation mixer": {

[0042] "mixingStrategies": {

[0043] "rotation": "MUL",

[0044] "weights": "ADD"

[0045] According to the present principles, an animation mixing device (for example according glTF) loads, parses and decodes all the content except the one in the “animations” root property. In this example, the result of this decoding is a single scene with three nodes: the first one corresponds to the head mesh, the second one to the left eye mesh, and the third one to the right eye mesh. Each mesh has a specific material with a texture. The head mesh has four morph targets: AU61_Eyes_tum_L: this morph target defines an animation for eyelids with a movement to the left; AU61_Eyes_tum_R: this morph target defines an animation for eyelids with a movement to the right; AU63_eyeUp: this morph target defines an animation for eyelids with a movement to the top; and AU63_eyeDown: this morph target defines an animation for eyelids with a movement to the bottom. Then, the animation mixing device decodes the “animations” root property that comprises two animations. The first animation has three channels: The first channel changes the morph target weights of the head mesh (node 1).

[0046] The changes only concern the two first morph targets AU61_Eyes_tum_L and AU61_Eyes_tum_R. The second channel rotates the left eye mesh (node 2) around the Y axis. The third channel rotates the right eye mesh (node 3) around the Y axis. The second animation has three channels: The first channel changes the morph target weights of the head mesh (node 1). The changes only concern the two last morph targets AU63_eyeUp and AU63_eyeDown. The second channel rotates the left eye mesh (node 2) around the X axis. The third channel rotates the right eye mesh (node 3) around the X axis. In this glTF example, the second animation has an “extensions” property with a “MPEG animation mixer” item, in which the “mixingStrategies” property has: a “rotation” property with value “MUL”: it indicates that a quaternion multiplication must be used when merging the rotation of two mixer updates. a weights” property with value “ADD”: it indicates that addition must be used when merging the weights of two mixer updates. Once the parsing is done, the mixing of the two animations can happen in the following way:

[0047] 1. A time t0between 0 and 2 is chosen for animation 0.

[0048] 2. A time between 0 and 2 is chosen for animation 1.

[0049] 3. Time t0is used to compute an output value for each channel of animation 0. We denote these output values: a. for the first channel: w0, a float vector with four values (one value for each morph target). b. for the second channel: Zo, a quaternion (to rotate the left eye). c. for the third channel: r0, a quaternion (to rotate the right eye).

[0050] 4. Time t±is used to compute an output value for each channel of animation 1. We denote these output values: a. for the first channel: w , a float vector with four values (one value for each morph target). b. for the second channel: l^, a quaternion (to rotate the left eye). c. for the third channel: , a quaternion (to rotate the right eye).

[0051] 5. Three mixer channels are created: a. A first mixer channel targeting node 1 (the head) with the following mixer updates: i. Update the morph target weights with w0. ii. Update the morph target weights with w±with mixing strategy “ADD”. b. A second mixer channel targeting node 2 (the left eye) with the following mixer updates: i. Update the rotation with l0. ii. Update the rotation with with mixing strategy “MUL”. c. A third mixer channel targeting node 3 (the right eye) with the following mixer updates: i. Update the rotation with r0. ii. Update the rotation with rrwith mixing strategy “MUL”.

[0052] 6. A mixer animation is created with the following animation channels: a. A first animation channel targets node 1 (the head) and update morph target weights with w = w0+ w1. b. A second animation channel targets node 2 (the left eye) and update rotation with I = lQl . c. A second animation channel targets node 3 (the right eye) and update rotation with r = ror .

[0053] 7. The mixer animation is applied to the scene as any usual glTF animation: a. The “weights” property of node 1 is replaced with w. b. The “rotation” property of node 2 is replaced with I. c. The “rotation” property of node 3 is replaced with r.

[0054] In another example, a glTF file can be used to store an avatar with two animations: one for walking, and the other for raising a hand. In an application, when the user presses an arrow key, the walking animation runs continuously until the user releases the key. Then, at any time, the user can press the space bar to start the hand animation. According to the present principles, the encoding of animation mixing strategies with the animations, the two animations are successfully mixed. The example of an avatar is to ease the understanding, but many other animated objects also benefit from an accurate animation mixing definition. Objects encoded in glTF files can be aggregated to build a large scene, like characters in a virtual world or game environment. The source of these objects can be outside the scope of the scene creators, like asset stores or other people / companies. These objects can contain animations that can be combined with different mixing strategies. According to the present principles, the user does not need to know in each case how to mix animations, and updates the objects, its software, scene, etc. accordingly. According to the present principles, information is inside the glTF file, and when there is no information, default mixing strategies are defined: the user has nothing to do when importing objects into the scene.

[0055] When creating 3D assets, a common approach consists of splitting the procedure into several stages. Each stage focuses on a specific task, like working on the mesh modeling, the textures, the animation, the lights, etc. These stages are usually handled by different people or teams, or by the same person but at different times. To manage this, a good solution is to store the result of each stage in a new file. Using the current glTF standard, one can already follow this approach for several tasks. It is not the case for others, like complex animations. Considering this context, a common approach consists of using a collection of basic animations (or controllers) per model. For instance, character animations like walk jump or idle. Faces are also associated with a set of facial expressions; each one being encoded as an animation. During an early step of the workflow, artists must design these animations, in which case the current glTF standard is sufficient. During later stages, artists must combine these basic animations to create a final animated scene. In this last case, if this is the last stage, the animations combinations can be baked into final animations into current glTF files. However, if this is not the last stage, and if another artist needs to update a bit the final animation (for instance, during the lighting stage), it is no longer possible with a standard approach: they must ask artists in the earlier stages to update animations and bake again. Using the animation mixing feature introduced according to the present principles, it becomes possible to store the animation creation process with animation mixing strategies in a glTF file. Then, when an artist receives a glTF file with these features, (s)he doesn’t have to assume or redefine the animation mixing strategies: they are defined in the glTF file, or if not present, default ones are defined, as well as an accurate mixing procedure. Figure 3 shows an example architecture of a device 30 which may be configured to implement animation mixing methods according to embodiments of the present principles. The device is linked with other devices via their bus 31 and / or via I / O interface 36.

[0056] Device 30 comprises following elements that are linked together by a data and address bus 31 :D

[0057] - a processor 32 (or CPU), which is, for example, a DSP (or Digital Signal Processor);

[0058] - a ROM (or Read Only Memory) 33;

[0059] - a RAM (or Random Access Memory) 34;

[0060] - a storage interface 35;

[0061] - an I / O interface 36 for reception of data to transmit, from an application; and

[0062] - a power supply (not represented in Figure 2), e.g. a battery.

[0063] In accordance with an example, the power supply is external to the device. In each of mentioned memory, the word « register » used in the specification may correspond to area of small capacity (some bits) or to very large area (e.g. a whole program or large amount of received or decoded data). The ROM 33 comprises at least a program and parameters. The ROM 33 may store algorithms and instructions to perform techniques in accordance with present principles. When switched on, the CPU 32 uploads the program in the RAM and executes the corresponding instructions.

[0064] The RAM 34 comprises, in a register, the program executed by the CPU 32 and uploaded after switch-on of the device 30, input data in a register, intermediate data in different states of the method in a register, and other variables used for the execution of the method in a register.

[0065] The implementations described herein may be implemented in, for example, a method or a process, an apparatus, a computer program product, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method or a device), the implementation of features discussed may also be implemented in other forms (for example a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus such as, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.

[0066] Device 30 is linked, for example via bus 31 to a set of sensors 37 and to a set of rendering devices 38. Sensors 37 may be, for example, cameras, microphones, temperature sensors, Inertial Measurement Units, GPS, hygrometry sensors, IR or UV light sensors or wind sensors. Rendering devices 38 may be, for example, displays, speakers, vibrators, heat, fan, etc.

[0067] In accordance with examples, the device 30 is configured to implement a method according to the present principles of mixing and merging animations of a 3D scene, and belongs to a set comprising:

[0068] - a mobile device;

[0069] - a communication device;

[0070] - a game device;

[0071] - a tablet (or tablet computer);

[0072] - a laptop;

[0073] - a still picture camera;

[0074] - a video camera.

[0075] Figure 4 shows an example of an embodiment of the syntax of a stream when the data are transmitted over a packet-based transmission protocol. Figure 4 shows an example structure 4 of a stream encoding a 3D scene comprising animated objects according to the present principle. The structure consists in a container which organizes the stream in independent elements of syntax. The structure may comprise a header part 41 which is a set of data common to every syntax element of the stream. For example, the header part comprises some of metadata about syntax elements, describing the nature and the role of each of them. The structure comprises a payload comprising an element of syntax 42 and at least one element of syntax 43 (there may be an element of syntax 43 for each animation for example). Syntax element 42 comprises data representative of the 3D scene. It comprises a scene description and data necessary to render the objects of the 3D scene (e.g. meshes, textures, lighting, etc.). Element of syntax 43 is a part of the payload of the data stream and may comprise data encoding the animations of the objects of the scene (channels and samplers). According to the present principles, a list of mixing strategies may be added to the animation list. If this list of mixing strategies is absent, default mixing strategies are used. In a variant, element of syntax 43 comprises an information indicating whether a list of mixing strategies is present. So, the data stream comprises data representative of a scene description, and data representative of two or more animations. Each animation comprises a one or more animation channels, and each channel targets one property of one node of a scene description. The data stream may also comprise data representative of a mixing strategy for one or more property kinds.

[0076] The implementations described herein may be implemented in, for example, a method or a process, an apparatus, a computer program product, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method or a device), the implementation of features discussed may also be implemented in other forms (for example a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus such as, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, Smartphones, tablets, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.

[0077] Implementations of the various processes and features described herein may be embodied in a variety of different equipment or applications, particularly, for example, equipment or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and related texture information and / or depth information. Examples of such equipment include an encoder, a decoder, a post-processor processing output from a decoder, a pre-processor providing input to an encoder, a video coder, a video decoder, a video codec, a web server, a set-top box, a laptop, a personal computer, a cell phone, a PDA, and other communication devices. As should be clear, the equipment may be mobile and even installed in a mobile vehicle.

[0078] Additionally, the methods may be implemented by instructions being performed by a processor, and such instructions (and / or data values produced by an implementation) may be stored on a processor-readable medium such as, for example, an integrated circuit, a software carrier or other storage device such as, for example, a hard disk, a compact diskette (“CD”), an optical disc (such as, for example, a DVD, often referred to as a digital versatile disc or a digital video disc), a random access memory (“RAM”), or a read-only memory (“ROM”). The instructions may form an application program tangibly embodied on a processor-readable medium. Instructions may be, for example, in hardware, firmware, software, or a combination. Instructions may be found in, for example, an operating system, a separate application, or a combination of the two. A processor may be characterized, therefore, as, for example, both a device configured to carry out a process and a device that includes a processor-readable medium (such as a storage device) having instructions for carrying out a process. Further, a processor-readable medium may store, in addition to or in lieu of instructions, data values produced by an implementation.

[0079] As will be evident to one of skill in the art, implementations may produce a variety of signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal may be formatted to carry as data the rules for writing or reading the syntax of a described embodiment, or to carry as data the actual syntax-values written by a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.

[0080] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made. For example, elements of different implementations may be combined, supplemented, modified, or removed to produce other implementations. Additionally, one of ordinary skill will understand that other structures and processes may be substituted for those disclosed and the resulting implementations will perform at least substantially the same function(s), in at least substantially the same way(s), to achieve at least substantially the same result(s) as the implementations disclosed. Accordingly, these and other implementations are contemplated by this application.

Claims

CLAIMS1. A method comprising:- obtaining two or more animations, wherein each animation comprises one or more animation channels, wherein each channel targets one property of one node of a scene description;- mixing (11) the channels of the two or more animations to obtain one or more mixer channels by grouping the one or more animation channels by targeted node and by property kind; and- merging (12) the one or more mixer channels to obtain a mixer animation comprising one or more merged animation channels, wherein each merged channel targets a kind of property of one node of the scene description.

2. The method of claim 1, wherein the merging is performed according to a mixing strategy, wherein the mixing strategy is a member of a group of mixing strategies comprising: replacing, adding and multiplying.

3. The method of claim 2, wherein the mixing strategy for a property kind is defined by default.

4. The method of claim 2, wherein the mixing strategy for a property kind is obtained from the two or more animations.

5. The method of one of claim 1 to 4, wherein the two or more animations are obtained from a data stream comprising data representative of a mixing strategy for one or more property kinds.

6. A device comprising a memory associated with at least one processor configured for:- obtaining two or more animations, wherein each animation comprises of one or more animation channels, wherein each channel targets one property of one node of a scene description;- mixing (11) the channels of the two or more animations to obtain one or more mixer channels by grouping the one or more animation channels by targeted node and by property kind; and- merging (12) the one or more mixer channels to obtain a mixer animation comprising one or more merged animation channels, wherein each merged channel targets a kind of property of one node of the scene description.

7. The device of claim 6, wherein the processor is configured to perform the merging according to a mixing strategy, wherein the mixing strategy is a member of a group of mixing strategies comprising: replacing, adding and multiplying.

8. The device of claim 7, wherein the mixing strategy for a property kind is defined by default.

9. The device of claim 7, wherein the mixing strategy for a property kind is obtained from the two or more animations.

10. The device of one of claim 9 to 9, wherein the two or more animations are obtained from a data stream comprising data representative of a mixing strategy for one or more property kinds.

11. A data stream comprising data representative of a scene description, data representative of two or more animations, wherein each animation comprises a one or more animation channels, wherein each channel targets one property of one node of a scene description, and data representative of a mixing strategy for one or more property kinds.

12. A method comprising: obtaining, from a data stream, two or more animations, wherein each animation comprises one or more animation channels, wherein each channel targets one property of one node of a scene description; wherein the channels of the two or more animations are mixed to obtain one or more mixer channels by grouping the one or more animation channels by targeted node and by property kind; and wherein the one or more mixer channels are merged to obtain a mixer animation comprising one or more merged animation channels, wherein each merged channel targets a kind of property of one node of the scene description; and rendering the mixer animation.

13. The method of claim 12, wherein the merging is performed according to a mixing strategy, wherein the mixing strategy is a member of a group of mixing strategies comprising: replacing, adding and multiplying.

14. The method of claim 13, wherein the mixing strategy for a property kind is defined by default.

15. The method of claim 13, wherein the mixing strategy for a property kind is obtained from the two or more animations.

16. The method of one of claim 12 to 15, wherein the two or more animations are obtained from a data stream comprising data representative of a mixing strategy for one or more property kinds.

17. A device comprising a processor configured for: obtaining, from a data stream, two or more animations, wherein each animation comprises one or more animation channels, wherein each channel targets one property of one node of a scene description; wherein the channels of the two or more animations are mixed to obtain one or more mixer channels by grouping the one or more animation channels by targeted node and by property kind; and wherein the one or more mixer channels are merged to obtain a mixer animation comprising one or more merged animation channels, wherein each merged channel targets a kind of property of one node of the scene description; and rendering the mixer animation.

18. The device of claim 17, wherein the merging is performed according to a mixing strategy, wherein the mixing strategy is a member of a group of mixing strategies comprising: replacing, adding and multiplying.

19. The device of claim 18, wherein the mixing strategy for a property kind is defined by default.

20. The device of claim 18, wherein the mixing strategy for a property kind is obtained from the two or more animations.

21. The device of one of claim 17 to 20, wherein the two or more animations are obtained from a data stream comprising data representative of a mixing strategy for one or more property kinds.