Avatar animation method and apparatus

By acquiring motion terms to generate control parameters for digital human animation, the problem of needing to manipulate underlying animation parameters one by one in existing technologies is solved, thus achieving efficient generation of digital human animation.

WO2026011860A1PCT designated stage Publication Date: 2026-01-15HISENSE VISUAL TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/087284
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-09
Filing Date
2025-04-03
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing digital human animation control methods require manipulating the underlying animation parameters one by one, resulting in low efficiency and making it difficult to achieve efficient digital human animation.

Method used

By acquiring motion terms, digital human animation control parameters are generated, and dynamic 3D models are generated based on these parameters and the digital human 3D model, avoiding the need to manipulate the underlying animation parameters one by one.

Benefits of technology

It improves the efficiency of digital human animation, simplifies the animation control process, and enhances the implementation efficiency of digital human animation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025087284_15012026_PF_FP_ABST
    Figure CN2025087284_15012026_PF_FP_ABST
Patent Text Reader

Abstract

Some embodiments of the present application provide an avatar animation method and apparatus, relating to the technical field of video processing. The method comprises: acquiring a motion term; acquiring an avatar animation control parameter on the basis of the motion term; and generating an avatar dynamic three-dimensional model on the basis of the avatar animation control parameter and an avatar three-dimensional model. Some embodiments of the present application are used for achieving avatar animation without performing one-by-one operations on lower-level animation parameters, thereby improving the efficiency of avatar animation.
Need to check novelty before this filing date? Find Prior Art

Description

Animation methods and devices for digital humans

[0001] This application claims priority to Chinese patent application filed on July 9, 2024, with application number 202410917366.9 and entitled "Animation Method and Apparatus for Digital Humans", the entire contents of which are incorporated herein by reference. Technical Field

[0002] Some embodiments of this application relate to the field of video processing technology. More specifically, they relate to a method and apparatus for animate a digital human. Background Technology

[0003] With the rapid development of immersive media technologies and the increasing acceptance of online work methods, immersive media technologies are becoming more widespread in people's work and lives. Compared to relying on large display screens to enhance immersion, Virtual Reality (VR) / Augmented Reality (AR) devices have the advantage of conforming to human physiological structure and can create vivid and realistic scenes through binocular stereoscopic vision. Therefore, VR / AR devices are more likely to serve as the hardware foundation for immersive media in the future. A crucial part of VR / AR technology is user representation, or simply avatar. When users wear VR / AR devices, they expect their real or virtual image to appear in the virtual scene presented by the headset and to move synchronously with their movements in the virtual scene. This is an important function of avatars and a key element in enhancing immersion.

[0004] Currently, in the scene description framework of immersive media, the animation control of digital humans is achieved by using the joint transformation matrix controlled by the skin description module to control the animation of the digital human's limbs and / or by using the weights of the deformation targets in the mesh description module (mash) combined with the facial blend shapes method to control the animation of the digital human's face. However, this method of animation control requires manipulating the underlying animation parameters such as the joint transformation matrix, deformation targets, and deformation weights one by one when implementing or editing each movement of the digital human. Manipulating these underlying animation parameters one by one is a very tedious and complex task. Therefore, the current method of digital human animation is very inefficient. Summary of the Invention

[0005] The exemplary embodiments of this application provide a digital human animation method and apparatus, which can realize digital human animation without manipulating the underlying animation parameters one by one, thereby improving the efficiency of digital human animation.

[0006] In a first aspect, some embodiments of this application provide a method for animate a digital human, including:

[0007] Acquire sports terminology;

[0008] Obtain digital human animation control parameters based on the aforementioned motion terminology;

[0009] Based on the digital human animation control parameters and the digital human 3D model, a dynamic 3D model of the digital human is generated.

[0010] As can be seen from the above technical solutions, the digital human animation method provided in some embodiments of this application first obtains motion terms, then obtains digital human animation control parameters based on the motion terms, and finally generates a dynamic 3D model of the digital human based on the digital human animation control parameters and the digital human 3D model. Since the digital human animation method provided in some embodiments of this application can obtain digital human animation control parameters based on the motion terms, and generate a dynamic 3D model of the digital human based on the digital human animation control parameters and the digital human 3D model reconstructed from the scene description file of the 3D scene to be rendered, some embodiments of this application can achieve digital human animation without operating on the underlying animation parameters one by one, thereby improving the efficiency of digital human animation. Attached Figure Description

[0011] To more clearly illustrate the implementation methods in some embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0012] Figure 1 shows a schematic diagram of the structure of the immersive media scene description framework in some embodiments of this application;

[0013] Figure 2 shows a schematic diagram of how skeletal points drive the rotation of skin apex in some embodiments of this application;

[0014] Figure 3 shows a schematic diagram of an animation control method based on facial fusion shape in some embodiments of this application;

[0015] Figure 4 shows a schematic diagram of the structure of the scene description file in some other embodiments of this application;

[0016] Figure 5 shows a framework diagram of the animation method for digital humans in some embodiments of this application;

[0017] Figure 6 shows a schematic flowchart of a digital human animation method in some embodiments of this application;

[0018] Figure 7 shows a schematic flowchart of a digital human animation method in some other embodiments of this application;

[0019] Figure 8 shows a flowchart of the steps of a digital human animation method in some embodiments of this application;

[0020] Figure 9 shows a schematic diagram of the structure of the animation device for a digital human in some embodiments of this application;

[0021] Figure 10 shows a structural block diagram of a display device in some embodiments of this application. Detailed Implementation

[0022] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.

[0023] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0024] Some embodiments of this application relate to a scene description framework for immersive media. Referring to the immersive media scene description framework shown in FIG1, in order to enable the display engine 11 to focus on media rendering, the immersive media scene description framework decouples media file access and processing from media file rendering, and designs a Media Access Function (MAF) 12 to be responsible for media file access and processing functions. A Media Access Function Application Programming Interface (API) is also designed, and the display engine 11 and the Media Access Function 12 interact with each other through the Media Access Function API. The display engine 11 can issue instructions to the Media Access Function 12 through the Media Access Function API, and the Media Access Function 12 can also request instructions from the display engine 11 through the Media Access Function API.

[0025] The general workflow of an immersive media scene description framework may include: 1) Display engine 11 obtains the scene description file provided by the immersive media service provider. 2) Display engine 11 parses the scene description file, obtains the access address of the media file, the attribute information of the media file (media type and encoding / decoding parameters, etc.), and the format requirements of the processed media file, and calls the media access function API to pass all or part of the information obtained from parsing the scene description file to media access function 12. 3) Media access function 12, based on the information passed by display engine 11, requests to download the specified media file from the media resource server or obtains the specified media file from the local machine, and establishes a corresponding pipeline for the media file. Subsequently, the media file is processed in the pipeline through decapsulation, decryption, decoding, post-processing, etc., to convert the media file from the encapsulated format to the format specified by display engine 11. 4) Media access function 12 stores the output data obtained after all processing in the specified cache. 5) Display engine 11 reads the fully processed data from the specified cache and renders the media file based on the data read from the cache.

[0026] Referring to Figure 1, in the workflow of the scene description framework for immersive media, the display engine 11 can include acquiring a scene description file, parsing the acquired scene description file to obtain the composition structure and detailed information of the 3D scene to be rendered, and rendering and displaying the 3D scene to be rendered based on the information obtained from parsing the scene description file. In some embodiments of this application, the specific workflow and principles of the display engine 11 are not limited, but are defined as follows: the display engine 11 can parse the scene description file, issue instructions to the media access function 12 through the media access function API, issue instructions to the cache management module 13 through the cache API, and retrieve processed data from the cache to complete the rendering and display of the 3D scene and its objects.

[0027] With the rapid development of immersive media technologies and the increasing acceptance of online work methods, immersive media technologies are becoming more widespread in people's work and lives. Compared to relying on large display screens to enhance immersion, Virtual Reality (VR) / Augmented Reality (AR) devices have the advantage of conforming to human physiological structure and can create vivid and realistic scenes through binocular stereoscopic vision. Therefore, VR / AR devices are more likely to serve as the hardware foundation for immersive media in the future. A key aspect of VR / AR technology is user representation, or simply avatar. When users wear VR / AR devices, they expect their real or virtual image to appear in the virtual scene presented by the headset and to move synchronously with their movements in the virtual scene. This is a crucial function of avatars and a vital element in enhancing immersion.

[0028] In the field of computer graphics (CG-based), a digital human is represented by the static geometric topology of a 3D mesh model, consisting of geometric points and a skeleton. The dynamic attributes of a digital human can be divided into two categories: limb movements, such as walking, running, jumping, turning, and picking up objects; and facial expression movements, such as smiling, frowning, and widening eyes. Limb movements have lower requirements for precision, while facial expression movements require far greater accuracy and complexity.

[0029] In the glTF 2.0-based 3D media resource format, animation is defined and implemented through syntax elements in the animation description module.

[0030] The dynamic properties of limbs are achieved through skinning to animate the body. Skinning establishes the relationship between skeletal points (or joints) and skin vertices. Each skeletal point is associated with many skin vertices, and the influence weight of each skin vertex is different. Through this relationship, the overall movement of the digital human can be controlled by controlling the rotation vectors of the skeletal points. That is, in a 3D media resource format based on glTF 2.0, by setting the animation type of "path" in "target" to "rotation", and finding the rotation information data of the corresponding nodes (joints) to be animated through "samplers", the motion information of the skeleton is obtained. Then, through skinning technology, the skeleton drives the skin animation, thereby realizing the animation of the digital human.

[0031] For example, referring to Figure 2, bone point 41 is associated with skin vertices 42 to 49. When the rotation vector of bone point 41 is rotation1 = [0.0, 0.0, 0.0, 0.1], the rotation vectors rotation2 to rotation9 of skin vertices 42 to 49 can be calculated based on the association between bone point 41 and skin vertices 42 to 49. Then, skin vertices 42 to 49 are rotated according to the rotation vectors rotation2 to rotation9 of skin vertices 42 to 49 respectively.

[0032] Facial dynamic attributes refer to facial animation-related attributes such as deformation and expression changes. Because facial skin changes are subtle and detailed, animate them using the same skeletal skinning method as the body, which would be redundant and cumbersome. Therefore, CG tools and glTF 2.0 both support animation based on facial blend shapes. A blend shape refers to a deformation relative to a base shape, usually represented as vertex displacements. Multiple blend shapes can be defined for the face, each considered an "expression base." By combining different blend shapes and their corresponding weights, facial skin deformation can be achieved. Specifically, in a glTF 2.0-based 3D media resource format, by setting the "path" animation style in the "target" to "weights," and using "samplers" to find the blend shape weight keyframe data of the mesh description module under the node to be animated, facial animation is achieved by combining the weights of each blend shape. For example, for a smiley face (morph target1) and a normal face (base shape), you can generate an animation of the transition between the smiley face and the normal face by using blend shape.

[0033] For example, as shown in FIG3, a transition animation 53 between facial expressions 51 and 52 can be generated by blend shape based on the weights of facial expressions 51 and 52, and a transition animation 55 between facial expressions 51 and 54 can also be generated by blend shape based on the weights of facial expressions 51 and 54.

[0034] Currently, the method of reconstructing dynamic facial expressions of digital humans using facial markers is quite prominent in the relevant technical field. The overall process of the facial marker method may include the following steps (1) to (4):

[0035] (1) Predefine a neutral digital human face 3D model and predefine a set of facial markers on the model.

[0036] The neutral digital human facial 3D model can be a static 3D model, whose topology can be a 3D mesh or a 3D point cloud, etc. Facial markers typically cover spatial locations with rich variations and significant impact on facial expressions, such as the eyes, eyebrows, mouth, nose, and jaw. Currently, there is no uniform rule for the total number of facial markers, but 68 points are commonly used. Of course, 21, 29, 98, 106, and 186 points can also be used. The larger the total number of facial markers, the more accurate the facial markers will be, but the corresponding data volume and computational load will also be greater. Therefore, in practical applications, the number of facial markers can be set according to the accuracy requirements of the facial 3D model and the performance of the relevant equipment.

[0037] (2) Obtain the input data at the target time and obtain a set of motion vectors of facial markers based on the input data.

[0038] The types of input data and the methods of processing it are highly diverse. For example, when the input data is depth video captured by an RGBD camera, the motion vectors of facial markers can be obtained by using computer vision and neural network methods to locate the facial markers. For instance, using a facial marker extraction network, the spatial coordinates of the facial markers can be obtained, and then the difference between these coordinates and those from previous moments can be used to obtain the motion vectors of the facial markers at the target moment. When the input data is text representing the emotions of a digital human, the motion vectors of the facial markers can be obtained by using a facial expression blending model. This involves first using a weighted superposition of basic expressions to obtain a realistic complex expression, and then locating the facial markers and calculating the motion vectors on this complex face.

[0039] (3) Using the motion vectors of this set of facial markers, obtain the motion vectors of all vertices in the digital human face 3D model.

[0040] When considering the amount of data and the number of points, facial markers are equivalent to sparse data, while all vertices of the static 3D model of a digital human face are equivalent to dense data. The mapping of motion vectors from facial markers to all vertices of the facial model is a mapping from sparse data to dense data. This step can use computer vision and neural network methods, such as using motion diffusion networks, to map the motion vectors of sparse data to the motion vectors of dense data.

[0041] (4) Superimpose the motion vectors of all vertices onto the neutral 3D model of the digital human face to obtain the data of the digital human's dynamic facial expression at the target time.

[0042] The process of superimposing vertex motion vectors onto a neutral 3D digital facial model varies depending on the topology of the static 3D model. When the topology is a 3D point cloud, the motion vector of each point is directly summed with the spatial coordinates of the corresponding point in the 3D point cloud, and the resulting spatial coordinates are the superimposed result. Additionally, the attribute information attached to the 3D point cloud needs to be mapped; that is, the spatial coordinates of the points in the point cloud change, but the attribute information remains unchanged. When the topology is a 3D mesh, the motion vector of each point is directly summed with the spatial coordinates of the corresponding vertex in the 3D mesh, and the resulting spatial coordinates are the superimposed result. Furthermore, the attribute information attached to the 3D mesh needs to be mapped one-to-one; the attribute information remains unchanged, and the connection relationships between vertices remain unchanged. That is, vertices that were previously connected still have the same connection relationships after coordinate transformation.

[0043] Through the above steps (1) to (4), the data of the digital human dynamic facial expression at the target time can be obtained. By continuously repeating the above process, the data of the digital human dynamic facial expression at each time can be obtained. By using the data at different times to drive the dynamic neutral digital human facial 3D model, the reconstruction and mapping of the digital human dynamic facial expression can be realized.

[0044] In some embodiments, after step (4) above, the dynamic facial expressions of the neutral digital human can be mapped onto a stylized digital human face to make the visual experience more realistic. The neutral digital human can be a digital human without any particular expression tendency, while the stylized digital human can be a digital human processed using specific techniques, forms, or features.

[0045] In the scene description framework of some embodiments of this application, the 3D scene service provider needs to provide a scene description file to the display engine, which then parses the scene description file. This scene description file includes, but is not limited to: description information of the entire 3D scene, description information of the digital human in the 3D scene, and information such as the location and index of facial marker points of the digital human in the 3D scene.

[0046] Referring to Figure 4, which is a schematic diagram of the structure of a scene description file that supports digital human representation and reconstruction in some embodiments of this application.

[0047] As shown in Figure 4, the scene description file may include MPEG media (MPEG_media) 601. The MPEG media (MPEG_media) 601 in the scene description file shown in Figure 4 can be used to describe the type of media file and provide necessary explanations for MPEG-type media files so that these MPEG-type media files can be used subsequently. Media files including facial marker data can be described in the MPEG media.

[0048] As shown in Figure 4, a scene description file may include a scene description module (scene), abbreviated as scene 602 in Figure 4. The scene description module (scene) 602 in the scene description file shown in Figure 4 is used to describe the 3D scenes contained in the scene description file. A scene description file may contain any number of 3D scenes, each represented by a separate scene description module.

[0049] As shown in Figure 4, the scene description file may include a node description module (node), which is abbreviated as node 603 in Figure 4. The node description module (node) 603 may include: a digital human node array ("MPEG_node_avatar":{}) 6031.

[0050] In some embodiments of this application, a digital human in a three-dimensional scene will be represented by a node, described by a corresponding node description module, and multiple levels of child nodes will be attached under the node. The specific parts of the digital human's body will be described layer by layer by the node description module corresponding to the child nodes. Through the hierarchical structure, the deepest unit in the upper body can be each finger joint. For example: Using the buttocks as a node to represent a digital human (layer 1), the node representing the digital human includes three child nodes: spine, left thigh, and right thigh (layer 2); selecting the child node representing the spine and continuing to extend it deeper, the child node representing the spine includes a child node representing the chest (layer 3); selecting the child node representing the chest and continuing to extend it deeper, the child node representing the chest includes a child node representing the upper chest (layer 4); selecting the child node representing the upper chest and continuing to extend it deeper, the child node representing the upper chest includes three child nodes representing the left shoulder, right shoulder, and neck (layer 5); selecting the child node representing the neck and continuing to extend it deeper, it includes a child node representing the head (layer 6).

[0051] In some embodiments, the hierarchical standard names for digital human body parts may be as follows:

[0052] In the node and child node hierarchy describing the specific parts of the digital human's body, each node and child node may have one or more 3D meshes attached to it, serving as 3D models of these body parts. Furthermore, using the displacement, rotation, scaling, and other parameters of the nodes and child nodes, these 3D models of body parts can be realistically pieced together to form a complete 3D model of the digital human.

[0053] Whether a node description module includes the extension "MPEG_node_avatar":{}" indicates whether the node corresponding to that node is used to represent a digital human. Specifically, if a node description module in a scene description file includes the extension "MPEG_node_avatar":{}", then it is determined that the node corresponding to that node description module represents a digital human.

[0054] It should be noted that in the scene description file, a digital human is represented by a node, and multiple levels of child nodes can be attached to this node. The specific parts of the digital human's body are described layer by layer by the node description module corresponding to the child nodes. In order to avoid information redundancy, in some embodiments of this application, only the digital human node array ("MPEG_node_avatar":{}) is added to the node description module corresponding to the top-level node representing the digital human, and the digital human node array is not added to the node description module corresponding to the child nodes attached to it.

[0055] In some embodiments, when a node represents a digital human, an array of digital human nodes ("MPEG_node_avatar":{}) can be added to the extension list ("extensions":{}) of the node description module corresponding to that node.

[0056] As shown in Figure 4, the digital human node array ("MPEG_node_avatar":{}) 6031 in the node description module 603 may include: active identifier syntax element ("isAvatar") 60311.

[0057] The active identifier syntax element ("isAvatar") in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) can be used to indicate whether the digital human represented by the node corresponding to the node description module (node) is active (whether the digital human needs to be rendered during the rendering process, and whether the display engine, media access function, etc. in the scene description box need to process the relevant data and information from the digital human).

[0058] In some embodiments, the method of indicating whether the digital person represented by the node corresponding to the node description module is active by using the active identifier syntax element ("isAvatar") may include: determining that the digital person represented by the node corresponding to the node description module is active when the value of the active identifier syntax element ("isAvatar") is a first value, and determining that the digital person represented by the node corresponding to the node description module is inactive when the value of the active identifier syntax element ("isAvatar") is a second value.

[0059] In some embodiments, the first value and the second value can be true and false, respectively. In other embodiments, the first value and the second value can be 1 and 0, respectively.

[0060] In some embodiments, the data type of the value of the active identifier syntax element ("isAvatar") can be boolean.

[0061] The digital human type syntax element ("type") in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) can be used to indicate the representation scheme of the digital human represented by the node corresponding to the node description module (node).

[0062] As shown in Figure 4, the digital human node array ("MPEG_node_avatar":{}) 6031 in the node description module 603 may include: digital human type syntax element ("type") 60312.

[0063] In some embodiments, the digital human type syntax element ("type") uses a Uniform Resource Name (URN) to describe the representation scheme of the digital human. The representation scheme of a digital human can be understood as the technical solution / architecture to which the digital human belongs. This technical solution / architecture defines the media format / type and other information of all body components of the digital human. Using this information, media access functions can establish corresponding pipelines for the reconstruction / driving of the corresponding digital human body components, or for the reconstruction / driving of the entire digital human body. For example, when the digital human represented by the node corresponding to the node description module (node) is an MPEG reference digital human, the value of the digital human type syntax element ("type") can be set to the Uniform Resource Name of the MPEG reference digital human. The URN of the MPEG reference digital human is urn:mpeg:sd:2023:avatar. Therefore, when the digital human represented by the node corresponding to the node description module (node) is an MPEG reference digital human, the digital human type syntax element and its value are: "type":"urn:mpeg:sd:2023:avatar".

[0064] In some embodiments, the data type of the value of the digital human type syntax element ("type") can be string.

[0065] As shown in Figure 4, the mapping list ("mappings":[]) 60313 may include: a part name syntax element ("path") 603131; in some embodiments, the mapping list ("mappings":[]) 60313 may include a node index syntax element ("node") 603132. The mapping list ("mappings":[]) in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) describes the mapping between digital human child nodes and hierarchical standard names of digital human body parts. The data type of the mapping list ("mappings":[]) is an array containing two syntax elements, namely the part name syntax element ("path") and the node index syntax element ("node"). The part name syntax element ("path") can be used to describe the standard names of the digital human body parts corresponding to the node / child node of the node description module, such as " / full_body / upper_body / head / face / eye_right", etc.; the node index syntax element ("node") can be used to index the nodes of the body parts identified by the part name syntax element ("path"), so that the body part names correspond to the nodes.

[0066] In summary, the syntax elements in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) are shown in Table 1 below:

[0067] Table 1

[0068] The syntax elements in the mapping list ("mappings":[]) of the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) are shown in Table 2 below:

[0069] Table 2

[0070] In some embodiments, the syntax elements in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) may not include the active identifier syntax element ("isAvatar"). When the syntax elements in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) do not include the active identifier syntax element ("isAvatar"), the syntax elements in the digital human node array ("MPEG_node_avatar":{}) of the node description module (node) are shown in Table 3 below:

[0071] Table 3

[0072] The scene description file shown in Figure 4 may include the definition of a digital human node array ("MPEG_node_avatar":{}) and provides a related reference model. The node description module including the digital human node array and the reference model provides a standardized way to define the geometry and semantics of the digital human's geometric description. Furthermore, it may include definitions of interactions between the digital human and the scene.

[0073] As shown in Figure 4, the scene description file may include a mesh description module (mesh), abbreviated as mesh 604 in Figure 4. The description method of mesh description module (mesh) 604 in the scene description file may include: for nodes representing 3D objects, the scene description file will use the data format of a 3D mesh for description, that is, use mesh description module (mesh) 604. The mesh description module contains various specific information, such as 3D coordinates and color information, all of which belong to the syntax elements in the attribute (mesh.primitives.attribute) of the primitive of mesh description module 604. The most basic mesh description module includes the 3D coordinates of vertices, the connection relationships between vertices, the color information of vertices, and the 3D coordinates requiring texture values. At the mesh description module level, the required 3D coordinate set texture coordinates (Texcoord) needs to be enabled so that the haptic values ​​obtained from the texture can be attached to the corresponding 3D coordinates of the 3D object. Texture coordinates (Texcoord) and color data, etc., belong to the attribute (mesh.primitives.attribute) of the primitive of mesh description module.

[0074] As shown in Figure 4, the scene description file may include an accessor description module (accessor), a buffer slice description module (bufferView), and a buffer description module (buffer), which are abbreviated as accessor 605, buffer slice 606, and buffer 607 in Figure 4. In some embodiments, the description method of the accessor description module (accessor) 605, buffer slice description module (bufferView) 606, and buffer description module (buffer) 607 in the scene description file may include: a part used to describe the output data of the media access function, and a part used to index the input data of the media access function (e.g., a media file including data related to facial marker points), and there is an interleaving relationship between the content used to describe the output data of the media access function and the content used to index the input data of the media access function.

[0075] In some embodiments, the method of describing the output data of a media access function through an accessor description module 605, a buffer slice description module 606, and a buffer description module 607 may include: a mesh description module 604 representing a dynamic 3D model of a digital human face in a 3D mesh for description points to a specific accessor description module 605; the accessor description module 605 then points to a corresponding buffer slice description module 606; and the buffer slice description module 606 then points to a corresponding buffer description module 607. This description method enables the storage and indexing of digital human facial expression model data. Specifically, the buffer described by the buffer description module 607 can be directly read by the display engine, and the stored data is data that can be directly used for rendering. In some embodiments of this application, the data stored in the buffer is the index information and position coordinates of digital human facial marker points and the dynamic 3D model of the digital human. The buffer slice description module (bufferView) 606, which describes the buffer slice, is responsible for slicing data in the buffer. This data slicing function can be achieved using two parameters: the starting byte offset (byteOffset) and the byte length (byteLength). The accessor description module (accessor) is responsible for adding additional information to the data in the buffer slice, such as the data type, the quantity of a certain type of data, and the numerical range of a certain type of data. The mesh description module (mesh) 604, which describes the 3D mesh representing the digital human face, will point to the accessor description module (accessor) 605 to retrieve the dynamic 3D model of the digital human to be rendered. The facial marker array, which describes the descriptive information of the digital human's facial markers, will point to the accessor description module (accessor) 605 to retrieve the index information and position coordinates of the digital human's facial markers.

[0076] Since the data related to digital facial markers is a time-varying medium, the accessor used to access this data needs to be capable of accessing time-varying media. Therefore, the accessor description module (accessor) 605 describing the accessor used to access the data related to digital facial markers should include an MPEG ring buffer ("MPEG_accessor_timed":{}) extension, and use the MPEG ring buffer to transform the described accessor into a time-varying accessor. When the accessor description module (accessor) 605 includes an MPEG ring buffer ("MPEG_accessor_timed":{}), the accessor description module (accessor) 605 contains two buffer slice index syntax elements ("bufferView"), one of which is outside the MPEG ring buffer, and the other is inside the MPEG ring buffer. The buffer slice index syntax element ("bufferView") outside the MPEG circular buffer will point to the relevant data of digital human facial markers or the dynamic 3D model of the digital human. The buffer slice index syntax element ("bufferView") inside the MPEG circular buffer will point to the header of the time-varying accessor. The header of the time-varying accessor, also known as the parameters of the time-varying accessor, describes the indexing method, data type, and data quantity of the time-varying data stored in the time-varying accessor. These parameters change over time and are therefore appended to the time-varying accessor as a header. The header may include: the buffer slice index syntax element ("bufferView") outside the MPEG circular buffer, the data type syntax element ("componentType"), and the value of the accessor type syntax element ("type") at different times.

[0077] Since the data related to digital facial markers is a time-varying medium, the buffer used to cache this data needs to be capable of caching time-varying media. Therefore, the buffer description module (buffer) 607, which describes the buffer used to cache the data related to digital facial markers, should include an MPEG circular buffer extension ("MPEG_buffer_circular":{}) and transform the buffer into a circular buffer using the MPEG circular buffer extension. The MPEG circular buffer ("MPEG_buffer_circular":{}) can contain syntax elements such as a media index syntax element ("media"), a track index syntax element ("tracks"), and a count syntax element ("count"). The media index syntax element ("media") and its value will point to the media file described in the MPEG media. The track index syntax element ("tracks") and its value are used to describe the track information of the source data of the data cached in the buffer. The count syntax element ("count") and its value are used to specify the number of storage counts in the circular buffer.

[0078] Through the above two extensions, different stages in the ring buffer store data of the time-varying medium at different times, and the different storage stages are separated in a slice manner by cache slicing. The time-varying accessor will access data from different cache slices (storage stages) at different times.

[0079] In some embodiments, the method of describing the input data of a media access function through an accessor description module (accessor) 605, a buffer slice description module (bufferView) 606, and a buffer description module (buffer) 607 may include: the input data of the media access function may be described as a media file in MPEG media (MPEG_media), and the media file declared in the MPEG media (MPEG_media) may be indexed by an MPEG circular buffer ("MPEG_buffer_circular":{}) in the buffer description module (buffer) 607, thereby delivering the input data of the media access function to the media access function.

[0080] In some embodiments, the method of describing the input data of a media access function through an accessor description module (accessor) 605, a buffer slice description module (bufferView) 606, and a buffer description module (buffer) 607 may include: instead of using a media file to carry the input data of the media access function, a scene description file is used to carry the input data of the media access function. The input data of the media access function will be delivered by the display engine to the media access function through the media access function API after the display engine completes the parsing of the scene description file. One implementation of directly describing the input data of the media access function using a scene description file may include: describing the input data of the media access function in the buffer description module (buffer) 607 of the scene description file.

[0081] In some embodiments, the method of describing the input data of a media access function through an accessor description module (accessor) 605, a buffer slice description module (bufferView) 606, and a buffer description module (buffer) 607 may include: describing a portion of the input data of the media access function as a media file in MPEG media (MPEG_media), and indexing the media file declared in MPEG media (MPEG_media) through an MPEG circular buffer ("MPEG_buffer_circular":{}) in the buffer description module (buffer) 607; and directly describing another portion of the input data of the media access function using a scene description file.

[0082] As shown in Figure 4, the scene description file may include a camera description module (camera), which is abbreviated as camera 608 in Figure 4. In some embodiments, the description method of the camera description module (camera) 608 in the scene description file may include: defining viewing-related visual information such as viewpoint and viewing angle of the node description module (node) 603 through the camera description module (camera) 608.

[0083] As shown in Figure 4, the scene description file may include a material description module 609, a texture description module 610, a sampler description module 611, and a texture map description module 612, which are abbreviated as material 609, texture 610, sampler 611, and texture map 612, respectively, in Figure 4. In some embodiments, the description method of the material description module 609 and texture description module 610 in the scene description file that supports obtaining the tactile values ​​of the tactile material properties of the reference type includes: describing additional information of the surface of the three-dimensional object through the material description module 609, texture description module 610, sampler description module 611, and texture map description module 612. The collaboration among the material description module 609, texture description module 610, sampler description module 611, and image description module 612 includes the following: the material description module 609 and texture description module 610 jointly define the color and physical information of the object's surface. The texture description module 610 specifies the sampler description module 611 and the image description module 612. The sampler description module 611 defines how to map the texture map onto the object's surface, implementing specific adjustments and wrapping of the texture. The image description module 612 uses ULRs to identify and index the texture map.

[0084] As shown in Figure 4, the scene description file may include a skin description module (skin), which is abbreviated as skin 613 in Figure 4. In some embodiments, the description method of the skin description module (skin) 613 in the scene description file includes: defining the motion and deformation relationship between the 3D mesh mounted by the node description module (node) 613 and the corresponding skeleton through the skin description module (skin) 613.

[0085] As shown in Figure 4, the scene description file may include an animation description module (animation 614) in Figure 4. In some embodiments, the description method of the animation description module (animation 614) in the scene description file includes: defining the animation description module (animation 614) as an animation added by the node description module (node) 603.

[0086] In some embodiments, the animation description module 614 can describe the animation added to the node description module 603 by one or more of position movement, angle rotation, and size scaling.

[0087] In some embodiments, the animation description module 614 may also indicate at least one of the start time, end time, and implementation method of the animation added to the node description module 603.

[0088] That is, in the scene description file provided in some embodiments of this application, animations can also be added to nodes representing objects in a three-dimensional object. The animation description module (animation) 614 describes the animations added to the nodes in three ways: position movement, angle rotation, and size scaling. It can also specify the start and end times of the animation and the implementation method of the animation.

[0089] As shown above, the scene description file has made the following contributions to digital humans: 1) The scene description file has newly extended the node description module (node), adding MPEG digital and node (MPEG_node_avatar) attributes to the extension list of the node description module to support the representation and reconstruction of digital humans in the scene. 2) The scene description file defines a default digital human reference representation format Morgan, which specifies the semantic geometric shape and other information of digital humans. However, the most important aspects of digital humans, such as animation control methods, the representation of animation data, and the encoding of animation information flow, are still lacking. In the scene description framework of immersive media, the animation control method of digital humans is to use the joint transformation matrix controlled by the skin description module (skin) to realize the animation control of the digital human's limbs and / or to use the weights ("weights":[]) of the deformation targets ("targets":[]) in the mesh description module (mash) combined with the facial blend shapes method to realize the animation control of the digital human's face. However, this method of controlling the animation of digital humans requires manipulating the underlying animation parameters such as joint transformation matrix, deformation target, and deformation weight one by one when implementing or editing each movement of the digital human. By constantly adjusting the values ​​of the animation parameters, the desired animation effect can be achieved. Manipulating these underlying animation parameters one by one is a very tedious and complex task. Therefore, the current method of digital human animation is very inefficient.

[0090] In view of this, some embodiments of this application provide a high-level animation control based on motion terminology. This high-level animation control based on motion terminology may include: defining motion primitives (or motion terms) for common human movements and their corresponding semantic descriptions; converting the motion primitives and corresponding semantic descriptions into digital human animation control parameters; and controlling the digital human to perform animation control through the converted digital human animation control parameters. The embodiments of this application can directly obtain the values ​​of each animation parameter through motion terminology, which can greatly improve animation efficiency compared to the scheme of constantly adjusting the values ​​of animation parameters.

[0091] Referring to Figure 5, which is a framework diagram of a digital human animation method provided in some embodiments of this application, the framework of the digital human animation method may include: a motion term acquisition module 71; a semantic acquisition module 72; an animation parameter acquisition module 73; a digital human animation control module 74; and a display device 75. The system includes a motion term acquisition module 71 for acquiring motion terms such as "walking," "running," "jumping," and "standing"; a semantic acquisition module 72 for acquiring semantic descriptions corresponding to the motion terms acquired by the motion term acquisition module 71; an animation parameter acquisition module 73 for acquiring corresponding animation control parameters based on the motion terms acquired by the motion term acquisition module 71 and / or the semantic descriptions corresponding to the motion terms acquired by the semantic acquisition module 72; a digital human animation control module 74 for acquiring a digital human 3D model and digital human animation control parameters, and for controlling the animation of the digital human 3D model based on the digital human animation control parameters to obtain a dynamic 3D model of the digital human; and a display device 75 for rendering the dynamic 3D model of the digital human obtained by the digital human animation control module 74 and displaying the rendering result of the dynamic 3D model of the digital human. In some embodiments, the animation parameter acquisition module 73 may include a predefined rule table 731; in some embodiments, the animation parameter acquisition module 73 may include an animation parameter generation model 732. Based on the animation parameter acquisition module 73 acquiring digital human animation control parameters through the predefined rule table 731 or the animation parameter generation model 732, the following two digital human animation methods are provided in some embodiments of this application:

[0092] In some embodiments, referring to FIG6, the animation method for a digital human includes the following steps S81 to S84:

[0093] S81, Sports Terminology Acquisition Module 71 acquires sports terminology.

[0094] In some embodiments, the sports term acquisition module 71 acquires sports terms, which may include receiving sports terms input by a user. For example, a user may input sports terms through a sports term input operation, which may be one or more of a touch click operation, a voice input operation, a special gesture, etc.

[0095] S82, the animation parameter acquisition module 73 acquires the animation control parameters corresponding to the motion terms according to the predefined rule table 731. Here, animation control parameters can refer to the values ​​of the animation control parameters.

[0096] In some implementations, the predefined rule table may include animation control parameters corresponding to each motion term. The animation parameter acquisition module 73 can directly obtain the values ​​of each animation control parameter based on the motion term. In other implementations, the predefined rule table may include a set of parameters corresponding to each motion term. The parameter values ​​in each parameter set may include multiple values. The animation parameter acquisition module 73 can determine a parameter set based on the motion term and then determine the values ​​of the parameters in that parameter set based on the motion information.

[0097] For example, taking the sports term "walking" as an example, the animation parameter acquisition module 73 can find the parameter set corresponding to "walking" as parameter set 1 in the predefined rule table. This parameter set 1 may include parameters related to arm and leg movements. Depending on the intensity of the exercise, the parameters in parameter set 1 can have multiple values. The animation parameter acquisition module 73 can further determine the corresponding animation parameter values ​​for the body parts to be animated (such as arms and legs) based on the different exercise intensities. For example, the animation control parameters for brisk walking and slow walking have different values.

[0098] For example, the correspondence between motion terms and animation control parameters is shown in Table 4. The animation control parameters in Table 4 can be a set of parameters.

[0099] Table 4. Correspondence between motion terms and animation control parameters

[0100] When the acquired motion term is "walking," the animation control parameter obtained according to the predefined rule table is animation control parameter 1. When the acquired motion term is "jumping," the animation control parameter obtained according to the predefined rule table is animation control parameter 2. When the acquired motion term is "running," the animation control parameter obtained according to the predefined rule table is animation control parameter 3. When the acquired motion term is "standing," the animation control parameter obtained according to the predefined rule table is animation control parameter 4. When the acquired motion term is "crouching," the animation control parameter obtained according to the predefined rule table is animation control parameter 5. When the acquired motion term is "rotating," the animation control parameter obtained according to the predefined rule table is animation control parameter 6. Figure 6 illustrates an example where the acquired motion input is "walking," and the animation control parameter obtained according to the predefined rule table is animation control parameter 1.

[0101] S83, the digital human animation control module 74 controls the animation of the digital human 3D model according to the animation control parameters to obtain the dynamic 3D model of the digital human.

[0102] S84, Display device 75 renders the dynamic 3D model of the digital human and displays the rendering result of the dynamic 3D model of the digital human.

[0103] In some embodiments, referring to FIG7, the animation method for a digital human may include the following steps S91 to S95:

[0104] S91, Sports Terminology Acquisition Module 71 acquires sports terminology.

[0105] Similarly, the sports terminology acquisition module 71 may acquire sports terms by receiving sports terms input by the user.

[0106] S92, Semantic Acquisition Module 72 acquires the semantic descriptions corresponding to the motion terms.

[0107] In some embodiments, the semantic acquisition module 72 may acquire the semantic description corresponding to the motion term by acquiring the semantic description corresponding to the motion term based on a preset mapping relationship.

[0108] The preset mapping relationship includes the semantic descriptions corresponding to each motion term.

[0109] For example, the preset mapping relationship can be shown in Table 5 below:

[0110] Table 5. Sports Terms and Corresponding Semantic Descriptions

[0111] In some embodiments, the semantic acquisition module 72 may further include receiving the semantic description corresponding to the motion term input by the user.

[0112] In some embodiments, receiving the semantic description corresponding to the sports term input by the user may include: after the user inputs the sports term, displaying the semantic description template corresponding to the sports term and prompting the user to fill in the semantic description template to obtain the semantic description corresponding to the sports term.

[0113] For example, if the obtained sports term is "jump," the following semantic description template corresponding to "jump" can be displayed, prompting the user to fill in the semantic description template to obtain the semantic description corresponding to the sports term:

[0114] The starting posture of the jump is (), the jump height is (), the jump distance is (), the duration is (), and the ending posture is (). Where () represents content to be filled in by the user.

[0115] S93. Input the motion terms and their corresponding semantic descriptions into the animation parameter generation model 732, and obtain the animation control parameters output by the animation parameter generation model 732.

[0116] In some embodiments, the animation parameter generation model 732 may be a model obtained by training a neural network model based on a training data set. The training data set may include multiple sets of training data, and any training data may include: motion terms, semantic descriptions corresponding to motion terms, and animation control parameters corresponding to motion terms.

[0117] S94, the digital human animation control module 74 controls the animation of the digital human 3D model according to the animation control parameters to obtain the dynamic 3D model of the digital human.

[0118] S95, display device 75 renders the dynamic 3D model of the digital human and displays the rendering result of the dynamic 3D model of the digital human.

[0119] The animation control parameters in some of the above embodiments are further described in detail below.

[0120] In some embodiments, animation control parameters may include the weights (mesh.weights) of the deformation targets of at least one mesh description module in the scene description file. The at least one mesh description module may be a mesh description module corresponding to a 3D mesh on which nodes representing parts of the digital human body are attached.

[0121] In some embodiments, animation control parameters may include: a local transformation matrix (node.matrix) of at least one node description module in the scene description file. The at least one node description module may be a mesh description module corresponding to a node representing a part of the digital human's body.

[0122] In some embodiments, animation control parameters may include the values ​​of local transformation syntax elements of at least one node description module in the scene description file. The local transformation syntax elements may include at least one of translation, rotation, and scale syntax elements.

[0123] The node corresponding to the node description module in the scene description file may contain a local transformation. This local transformation can be given as a local transformation matrix, or it can be given using separate translation, rotation, and scaling attributes, where rotation is given as a quaternion. The formula for calculating the local transformation matrix and the translation, rotation, and scaling attributes is: M = T * R * S

[0124] Where T, R, and S are matrices created by translation, rotation, and scaling. The global transformation of a node is given by the product of all local transformations on the path from the root node to the corresponding node.

[0125] In some embodiments, animation control parameters may include: the weights of deformation targets of at least one mesh description module in the scene description file and the local transformation matrix of at least one node description module in the scene description file.

[0126] In some embodiments, animation control parameters may include: the weights of deformation targets of at least one mesh description module in the scene description file and the values ​​of local transformation syntax elements of at least one node description module in the scene description file.

[0127] This application provides an animation method for a digital human in some embodiments. Referring to FIG8, the animation method for the digital human includes the following steps. It should be noted that the scheme in FIG8 can be used in combination with the scheme described above, and the contents not described in detail can be referred to the description above.

[0128] S101, Obtain sports terminology.

[0129] In some embodiments, acquiring sports terms includes receiving sports terms input by a user. That is, the user can input the sports terms into the display device via peripherals such as a mouse, keyboard, microphone, or touchscreen.

[0130] In some embodiments, the sports terminology includes one or more of walking, jumping, running, standing, squatting, rotating, etc.

[0131] S102. Obtain the digital human animation control parameters according to the aforementioned motion terminology. Step S102 can be executed by the animation parameter acquisition module 73 described above.

[0132] The control parameters for digital human animation include one or more of the following: the weights of the deformation targets in the mesh description module, the local transformation matrix of the node description module, and the values ​​of the local transformation syntax elements in the node description module. Wherein, the mesh description module is at least one mesh description module in the scene description file, and the node description module is at least one node description module in the scene description file.

[0133] In some embodiments, obtaining digital human animation control parameters based on the motion terminology may include: obtaining the weights of deformation targets of at least one mesh description module in the scene description file based on the motion terminology.

[0134] For example, a mesh description module in the scene description file of the 3D scene to be rendered is shown below:

[0135] Therefore, the digital human animation control parameters obtained according to the motion terminology may include: "weights":[x, y], thereby controlling the weight of the morph targets corresponding to "POSITION":1 in row n+10 of "targets":[] to be x, and controlling the weight of the morph targets corresponding to "POSITION":2 in row n+13 of "targets":[] to be y.

[0136] In some embodiments, obtaining digital human animation control parameters based on the motion terminology may include: obtaining the local transformation matrix of at least one node description module in the scene description file based on the motion terminology.

[0137] For example, a node description module in the scene description file of the 3D scene to be rendered is shown below:

[0138] Therefore, the digital human animation control parameters obtained based on the aforementioned motion terminology may include:

[0139] By replacing the local transformation matrix in rows n+03 to n+08 in the animation control parameters, and calculating the motion information of the skeleton based on the local transformation matrix in the animation control parameters, and then using skinning technology, the skeleton drives the skin animation, thereby realizing digital human animation.

[0140] In some embodiments, obtaining digital human animation control parameters based on the motion terms includes: obtaining the value of a local transformation syntax element of at least one node description module in the scene description file based on the motion terms; wherein, the local transformation syntax element may include at least one of a translation syntax element ("translation"), a rotation syntax element ("rotation"), and a scale syntax element ("scale").

[0141] For example, a node description module in the scene description file of the 3D scene to be rendered is shown below:

[0142] The digital human animation control parameters obtained based on the motion terms may include at least one of "translation":[x1,x2,x3], "rotation":[y1,y2,y3,y4], and "scale":[z1,z2,z3]. The values ​​of the syntax elements in rows n+3 to n+5 are replaced by the values ​​of the local transformation syntax elements in the animation control parameters. The motion information of the skeleton is calculated based on the values ​​of the local transformation syntax elements in the animation control parameters. Then, through skinning technology, the skeleton drives the skin animation, thereby realizing the digital human animation.

[0143] In some embodiments, obtaining digital human animation control parameters based on the motion terms may include: obtaining the local transformation matrix of at least one node description module in the scene description file and the weights of the deformation targets of at least one mesh description module in the scene description file based on the motion terms.

[0144] In some embodiments, obtaining digital human animation control parameters based on the motion terms may include: obtaining the values ​​of local transformation syntax elements of at least one node description module in the scene description file and the weights of deformation targets of at least one mesh description module in the scene description file based on the motion terms. The local transformation syntax elements may include at least one of translation syntax elements, rotation syntax elements, and scaling syntax elements.

[0145] That is, in some embodiments, the digital human animation control parameters can be a combination of at least two of the following: the weights of the deformation targets of the 3D mesh, the local transformation matrix of the nodes, and the values ​​of the local transformation syntax elements of the nodes.

[0146] S103. Generate the dynamic 3D model of the digital human based on the digital human animation control parameters and the digital human 3D model.

[0147] In some embodiments, before generating the dynamic 3D model of the digital human based on the digital human animation control parameters and the digital human 3D model, the digital human animation method further includes: obtaining a scene description file of the 3D scene to be rendered; parsing the scene description file to obtain description information of the digital human; and constructing the 3D model of the digital human based on the description information of the digital human.

[0148] In some embodiments, the scene description file of the 3D scene to be rendered can be obtained through the display engine in the scene description framework of immersive media, and the scene description file can be parsed to obtain information containing information for reconstructing the 3D model of the digital human. The information for reconstructing the 3D model of the digital human can be sent to the media access function via the media access function API. The media access function then requests the server to download or obtain the digital human-related media files of a specified quality level from the local machine according to the information passed by the media access function API, and establishes a corresponding pipeline for reconstructing the digital human in the media access function. Subsequently, the reconstruction of the digital human is performed in the pipeline to obtain the reconstructed 3D model of the digital human.

[0149] In some embodiments, generating the dynamic 3D model of the digital human based on the digital human animation control parameters and the digital human 3D model reconstructed from the scene description file of the 3D scene to be rendered may include: a media access function using the corresponding pipeline of the digital human to generate the dynamic 3D model of the digital human based on the digital human animation control parameters and the digital human 3D model reconstructed from the scene description file of the 3D scene to be rendered.

[0150] In some embodiments of this application, the animation method for the digital human further includes: rendering the dynamic 3D model of the digital human and displaying the rendering result of the dynamic 3D model of the digital human.

[0151] In some embodiments, rendering the digital human dynamic 3D model and displaying the rendering result of the digital human dynamic 3D model may include: a media access function writing the digital human dynamic 3D model to a cache, a display engine reading the digital human dynamic 3D model from the cache, rendering the digital human dynamic 3D model, and displaying the rendering result of the digital human dynamic 3D model.

[0152] The digital human animation method provided in some embodiments of this application first obtains motion terms, then obtains digital human animation control parameters based on the motion terms, and finally generates a dynamic 3D model of the digital human based on the digital human animation control parameters and the digital human 3D model. Since the digital human animation method provided in some embodiments of this application can obtain digital human animation control parameters based on the motion terms, and generate a dynamic 3D model of the digital human based on the digital human animation control parameters and the digital human 3D model reconstructed from the scene description file of the 3D scene to be rendered, some embodiments of this application can achieve digital human animation without manipulating the underlying animation parameters one by one, thereby improving the efficiency of digital human animation.

[0153] As an extension and refinement of the above embodiments, some embodiments of this application provide another method for animate a digital human, which includes the following steps:

[0154] S111, Obtain sports terminology.

[0155] S112. Obtain the digital human animation control parameters according to the motion terminology and the first mapping relationship.

[0156] The first mapping relationship may include the mapping relationship between the motion terms and the digital human animation control parameters. The first mapping relationship may be a predefined rule table 731.

[0157] For example, when the first mapping relationship is as shown in Table 4 above, then:

[0158] When the acquired motion term is "walking," the digital human animation control parameter is obtained as animation control parameter 1 based on the motion term and the first mapping relationship; when the acquired motion term is "jumping," the digital human animation control parameter is obtained as animation control parameter 2 based on the motion term and the first mapping relationship; when the acquired motion term is "running," the digital human animation control parameter is obtained as animation control parameter 3 based on the motion term and the first mapping relationship; when the acquired motion term is "standing," the digital human animation control parameter is obtained as animation control parameter 4 based on the motion term and the first mapping relationship; when the acquired motion term is "crouching," the digital human animation control parameter is obtained as animation control parameter 5 based on the motion term and the first mapping relationship; and when the acquired motion term is "rotating," the digital human animation control parameter is obtained as animation control parameter 6 based on the motion term and the first mapping relationship.

[0159] S113. Generate a dynamic 3D model of the digital human based on the digital human animation control parameters and the digital human 3D model.

[0160] S114. Render the dynamic 3D model of the digital human and display the rendering result of the dynamic 3D model of the digital human.

[0161] That is, a mapping table is pre-established that includes various motion terms and their corresponding digital human animation control parameters, and the corresponding digital human animation control parameters can be quickly obtained by looking up the table, thereby improving the efficiency of digital human animation.

[0162] As an extension and refinement of the above embodiments, some embodiments of this application provide another method for animate a digital human, which includes the following steps:

[0163] S121. Obtain sports terminology.

[0164] S122. Obtain the semantic description corresponding to the sports term based on the sports term.

[0165] In some embodiments, step S122 (obtaining the semantic description corresponding to the motion term based on the motion term) includes: obtaining the semantic description corresponding to the motion term based on the motion term and the second mapping relationship; wherein, the second mapping relationship includes the mapping relationship between the motion term and the semantic description corresponding to the motion term.

[0166] For example, the second mapping relationship can be as shown in Table 5 above. When the second mapping relationship is as shown in Table 5 above, then when the obtained motion term is "jump," the semantic description corresponding to the obtained motion term is: a rapid and brief upward movement, which typically includes both legs leaving the ground simultaneously, propelling the body into the air. The height, distance, and duration of the jump depend on the individual's strength and skill. The start and end of a jump are usually accompanied by squatting and standing movements.

[0167] In some embodiments, step S122 (obtaining the semantic description corresponding to the sports term based on the sports term) may include: receiving the semantic description corresponding to the sports term input by the user.

[0168] In some embodiments, receiving the semantic description corresponding to the sports term input by the user may include: outputting a semantic description template corresponding to the sports term; and generating a semantic description corresponding to the sports term based on the user's editing operation on the semantic description template.

[0169] For example: when the acquired motion term is walking, the output semantic description template for walking can be: the legs move alternately towards the direction of the body at a frequency of _____, the upper body maintains a ____ posture, the arm swing amplitude is _____, and the stride is ____.

[0170] The semantic description of walking generated based on the user's editing of the walking semantic description template can be as follows: the legs move forward alternately at a frequency of 90 times / minute, the upper body maintains a natural and upright posture, the arm swing amplitude is 40°, and the stride is 65cm.

[0171] S123. Input the motion term and the semantic description corresponding to the motion term into the pre-trained animation parameter generation model, and obtain the digital human animation control parameters output by the animation parameter generation model.

[0172] In some embodiments, before inputting the motion terms and their corresponding semantic descriptions into a pre-trained animation parameter generation model and obtaining the digital human animation control parameters output by the animation parameter generation model, the digital human animation method further includes:

[0173] Obtain a training dataset, which may include multiple sets of training data. Each set of training data includes: sample motion terms, sample semantic descriptions corresponding to the sample motion terms, and digital human animation control parameters corresponding to the sample motion terms and the sample semantic descriptions. Train a preset machine learning model based on the training dataset to obtain the animation parameter generation model.

[0174] In some embodiments, the preset machine learning model may be a neural network model.

[0175] S124. Generate a dynamic 3D model of the digital human based on the digital human animation control parameters and the digital human 3D model.

[0176] S125. Render the dynamic 3D model of the digital human and display the rendering result of the dynamic 3D model of the digital human.

[0177] That is, an animation parameter generation model is pre-trained. After obtaining motion terms, the semantic descriptions corresponding to the motion terms are first obtained. The motion terms and their corresponding semantic descriptions are then input into the pre-trained animation parameter generation model to obtain the digital human animation control parameters output by the animation parameter generation model, thereby improving the efficiency of digital human animation.

[0178] In some embodiments, this application provides an animation device for a digital human. Referring to FIG9, the animation device 1300 for the digital human may include:

[0179] Acquisition unit 131 is used to acquire sports terms;

[0180] Processing unit 132 is used to obtain digital human animation control parameters according to the motion terms;

[0181] The generation unit 133 is used to generate the dynamic three-dimensional model of the digital human based on the digital human animation control parameters and the digital human three-dimensional model.

[0182] The digital human animation device provided in the above embodiments can execute the digital human animation method provided in any of the above embodiments, and the implementation principle and the achieved technical effect are the same. To avoid redundancy, it will not be described in detail again.

[0183] In some embodiments of this application, a display device is provided, which may include:

[0184] Memory, configured to store computer programs;

[0185] The processor is configured to cause the display device to implement the digital human animation method described in any of the above embodiments when a computer program is invoked.

[0186] Figure 10 illustrates an exemplary configuration block diagram of the display device 1400 in some of the above embodiments. As shown in Figure 10, the display device 1400 may include at least one of the following: a tuner 141, a communicator 142, a detector 143, an external device interface 144, a controller 145, a display 146, an audio output interface 147, a memory 148, a power supply 149, and a user interface.

[0187] In some embodiments, controller 145 includes at least one of: a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM (random access memory), ROM (read-only memory), a first to an nth interface for input / output, a communication bus, etc.

[0188] The display 146 includes a display screen assembly for presenting images, a driving assembly for driving image display, a component for receiving image signals from the controller output, and a user control UI for displaying video content, image content, menu control interface, and user control UI.

[0189] The display 146 may be a liquid crystal display, an organic light-emitting diode (OLED) display, or a projection display, etc.

[0190] The communicator 142 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of the following: a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The display device 1400 can send and receive control signals and data signals with the control device or server through the communicator 142.

[0191] The user interface can be used to receive control signals input by the user through a control device (such as an infrared remote control) or by touch or gesture.

[0192] Detector 143 is used to acquire signals from the external environment or to interact with the external environment. For example, detector 143 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 143 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 143 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0193] The external device interface 144 may include, but is not limited to, one or more of the following: High-Definition Multimedia Interface (HDMI), analog or high-definition component input interface (component), Composite Video Broadcast Signal (CVBS), Universal Serial Bus (USB), RGB (Red, Green, Blue) port, etc. It may also be a composite input / output interface formed by multiple of the above interfaces.

[0194] The tuner / demodulator 141 receives broadcast television signals via wired or wireless means, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals. In some embodiments, the controller 145 and the tuner / demodulator 141 may be located in different separate devices, that is, the tuner / demodulator 141 may also be located in an external device of the main device where the controller 145 is located, such as an external set-top box.

[0195] The controller 145 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 145 controls the overall operation of the display device 1400. For example, in response to receiving a user command to select a UI object to display on the monitor 146, the controller 145 can perform operations related to the object selected by the user command.

[0196] In some embodiments, controller 145 includes at least one of: a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM (random access memory), ROM (read-only memory), a first to an nth interface for input / output, a communication bus, etc.

[0197] Some embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a computing device, causes the computing device to implement the animation method for a digital human as described in any of the above embodiments.

[0198] Some embodiments of this application provide a computer program product that, when run on a computer, enables the computer to implement the digital human animation method described in any of the above embodiments.

[0199] Some embodiments of this application provide a chip including a memory and a processor. The processor may be a logic circuit, an integrated circuit, or a general-purpose processor. The memory stores computer instructions, and the processor can implement the animation method of the digital human described in any embodiment by reading the computer instructions stored in the memory.

[0200] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0201] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.

Claims

1. A method for animate a digital human, characterized in that, include: Acquire sports terminology; Obtain digital human animation control parameters based on the aforementioned motion terminology; Based on the digital human animation control parameters and the digital human 3D model, a dynamic 3D model of the digital human is generated.

2. The method according to claim 1, characterized in that, The step of obtaining digital human animation control parameters based on the motion terminology includes: The digital human animation control parameters are obtained based on the motion terms and the first mapping relationship; The first mapping relationship includes the mapping relationship between the motion terms and the digital human animation control parameters.

3. The method according to claim 1, characterized in that, The step of obtaining digital human animation control parameters based on the motion terminology includes: Obtain the semantic description corresponding to the sports term based on the sports term; The motion terms and their corresponding semantic descriptions are input into a pre-trained animation parameter generation model, and the digital human animation control parameters output by the animation parameter generation model are obtained.

4. The method according to claim 3, characterized in that, Before inputting the motion terms and their corresponding semantic descriptions into a pre-trained animation parameter generation model, and obtaining the digital human animation control parameters output by the animation parameter generation model, the method further includes: Obtain a training dataset, which includes multiple sets of training data. Each set of training data includes: sample motion terms, sample semantic descriptions corresponding to the sample motion terms, and digital human animation control parameters corresponding to the sample motion terms and the sample semantic descriptions. The preset machine learning model is trained based on the training dataset to obtain the animation parameter generation model.

5. The method according to claim 4, characterized in that, The preset machine learning model is a neural network model.

6. The method according to claim 3, characterized in that, The step of obtaining the semantic description corresponding to the sports term based on the sports term includes: Based on the motion terminology and the second mapping relationship, obtain the semantic description corresponding to the motion terminology; The second mapping relationship includes the mapping relationship between the motion term and the semantic description corresponding to the motion term.

7. The method according to claim 3, characterized in that, The step of obtaining the semantic description corresponding to the sports term based on the sports term includes: Receive the semantic description corresponding to the motion term input by the user.

8. The method according to claim 7, characterized in that, The semantic description corresponding to the motion term input by the user includes: Output the semantic description template corresponding to the motion term; Based on the user's editing operations on the semantic description template, the semantic descriptions corresponding to the motion terms are generated.

9. The method according to any one of claims 1-8, characterized in that, The step of obtaining digital human animation control parameters based on the motion terminology includes: Based on the motion terminology, obtain the weights of the deformation targets of at least one mesh description module in the scene description file.

10. The method according to any one of claims 1-8, characterized in that, The step of obtaining digital human animation control parameters based on the motion terminology includes: Based on the motion terminology, obtain the local transformation matrix of at least one node description module in the scene description file.

11. The method according to any one of claims 1-8, characterized in that, The step of obtaining digital human animation control parameters based on the motion terminology includes: Based on the motion terminology, obtain the value of the local transformation syntax element of at least one node description module in the scene description file; The local transformation syntax element includes at least one of the following: translation syntax element, rotation syntax element, and scaling syntax element.

12. The method according to any one of claims 1-8, characterized in that, Before generating a dynamic 3D model of the digital human based on the digital human animation control parameters and the digital human 3D model, the method further includes: Obtain the scene description file of the 3D scene to be rendered; Parse the scene description file to obtain the description information of the digital human; The digital human's 3D model is constructed based on the digital human's description information.

Citation Information

Patent Citations

  • Animation generation method and device of digital object, electronic equipment and storage medium

    CN114898018A

  • Generation method and device of digital human animation, electronic equipment and storage medium

    CN116258799A

  • Interaction method and system of digital human and virtual whiteboard, and storage medium

    CN116797695A

  • Server, display device and digital human interaction method

    CN117809678A

  • Three-dimensional digital human generation method and device, electronic equipment, storage medium and program product

    CN118115642A