Digital human driving method, device, equipment and storage medium

Through the digital human standard markup protocol and parsing engine configured based on markup language, the standardization and unification of digital human rendering drivers are achieved, which solves the reusability problem of rendering drivers in different scenarios and improves the flexibility and adaptability of rendering drivers.

CN114494541BActive Publication Date: 2025-10-21ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210056508.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-18
Publication Date
2025-10-21
Estimated Expiration
2042-01-18

AI Technical Summary

Technical Problem

In the existing technology, digital human rendering drivers lack a unified standard interface, which makes it impossible to reuse rendering drivers in different scenarios and makes them sensitive to upgrades or changes of different rendering engines.

Method used

It adopts the digital human standard markup protocol (such as VAML) based on markup language configuration, receives the driving data packet through the parsing engine and parses it to obtain the driving information, calls the rendering engine to drive the digital human, realizes the unified driving of different rendering engines, and supports multiple rendering engines such as Unity and Unreal Engine.

Benefits of technology

The standardization and unification of digital human rendering drivers have been achieved, which can be reused in different service scenarios and is not affected by rendering engine upgrades or changes, thereby improving development efficiency and the flexibility of rendering drivers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494541B_ABST
    Figure CN114494541B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a digital human driving method, device, equipment and storage medium, related to a digital human standard markup protocol based on a markup language configuration. The method comprises: receiving, by a parsing engine, a driving data packet for a digital human, and parsing the driving data packet to obtain driving information; wherein the received driving data packet is a data packet based on the configured digital human standard markup protocol, used to control the digital human to perform a preset event at a preset time; calling, by the parsing engine, a rendering engine, and driving the pre-rendered digital human according to the driving information in the called rendering engine. The production of the digital human is standardized based on the digital human standard markup protocol, and the rendering and driving of the digital human are unified, so that the rendering and driving of the digital human do not exist in the logic of related service scenarios, the rendering and driving of the digital human are reused, and are not affected by the upgrade or change of the rendering and driving engine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a digital human driving method, a corresponding digital human driving device, a corresponding electronic device, and a corresponding computer storage medium. Background Art

[0002] In a narrow sense, digital humans are the product of the fusion of information science and life science. They utilize information science methods to simulate the human body's form and function at different levels. After completion, digital humans need to be rendered and driven. Rendering involves displaying the constructed digital human on a display. At this point, the rendered digital human is static, while driving the digital human involves animating the rendered static digital human like a real person.

[0003] The development of virtual digital humans has entered a period of rapid growth, and they can be applied to a variety of scenarios requiring digital humans, such as the gaming industry and live streaming. Currently, a set of rendering drivers is typically produced for each project to enable rendering and driving of digital humans. However, when rendering and driving digital humans in new projects, this process must be started from scratch. Even though digital human platforms exist for rendering and driving digital humans, the lack of standardized driver and rendering interfaces for different digital human drivers hinders the ability to render and drive digital humans in different scenarios. Summary of the Invention

[0004] In view of the above problems, the embodiments of the present application are proposed to provide a digital human driving method, a corresponding digital human driving device, a corresponding electronic device and a corresponding computer storage medium that overcome the above problems or at least partially solve the above problems.

[0005] The present application discloses a digital human driving method, which involves a digital human standard markup protocol configured based on a markup language, and is applied to a parsing engine adapted to the configured digital human standard markup protocol, wherein the parsing engine supports different rendering engines. The method includes:

[0006] The parsing engine receives a driving data packet for the digital human and parses the driving data packet to obtain driving information; wherein the received driving data packet is a data packet based on the configured digital human standard tagging protocol and is used to control the digital human to execute a preset event at a preset time;

[0007] The rendering engine is called by the parsing engine, and the pre-rendered digital human is driven in the called rendering engine according to the driving information.

[0008] Optionally, before receiving the driving data packet for the digital human, the method further includes:

[0009] Receive construction information for a digital human sent by a user system; the construction information includes digital human information applicable to a preset scene, digital human costume information, digital human voice information, and digital human movement information;

[0010] The parsing engine generates a digital human adapted to a preset scene based on the construction information, so as to pre-render the constructed digital human.

[0011] Optionally, after generating a digital human adapted to a preset scene based on the construction information, the method further includes:

[0012] Determining a rendering engine corresponding to the preset scene from rendering engines integrated by the parsing engine based on the digital human information applicable to the preset scene;

[0013] The rendering engine is called by the parsing engine, and the constructed digital human and the preset scene are rendered in advance by the called rendering engine, so as to drive the digital human in the pre-rendered preset scene.

[0014] Optionally, receiving a driving data packet for the digital human through the parsing engine includes:

[0015] The parsing engine receives a driving data packet for the digital human sent by the user system; the driving data packet for the digital human is a data packet converted by the user system into a digital human standard markup protocol by converting content information that the digital human needs to display based on the construction information; wherein, in the converted driving data packet, elements for representing preset events and the starting position attributes of the elements of the preset events for representing preset moments are used to configure the content information that the digital human needs to display, and the content information that the digital human needs to display includes speech events, action events and expression events that the digital human needs to perform at the preset moment, as well as card insertion events at the preset moment in the preset scene.

[0016] Optionally, parsing the driver data packet to obtain driver information includes:

[0017] The parsing engine performs real-time parsing on the elements in the driving data packet that represent preset events, as well as the starting position attributes of the elements of the preset events that represent preset moments, to obtain the driving information of the digital human; the driving information includes speech text information, action text information, card text information, and expression text information at the preset moment.

[0018] Optionally, driving the pre-rendered digital human in the called rendering engine according to the driving information includes:

[0019] The parsing engine converts the spoken text information into streaming voice data in real time, and generates corresponding mouth shape data, expression data and movement data at a preset time based on the spoken text information, movement information and expression information during the streaming voice conversion process;

[0020] The speech data, mouth shape data, expression data and action data generated at a preset time are sent in real time to the called rendering engine through the parsing engine;

[0021] The rendering engine obtains a pre-rendered digital human, and drives the pre-rendered digital human to play based on the voice data, mouth shape data, expression data and action data at a preset time, so as to push the played digital human to the user system.

[0022] Optionally, the driving information further includes card text information at a preset time, and the calling rendering engine driving the pre-rendered digital human according to the driving information further includes:

[0023] In the process of driving the pre-rendered digital human to play, the card text information at the preset moment is inserted into the preset scene according to the preset moment.

[0024] The present application also discloses a digital human driving device, which involves a digital human standard markup protocol configured based on a markup language, and is applied to a parsing engine adapted to the configured digital human standard markup protocol. The parsing engine supports different rendering engines. The device includes:

[0025] A driving data packet parsing module is used to receive a driving data packet for the digital human and parse the driving data packet to obtain driving information; wherein the received driving data packet is a data packet based on the configured digital human standard tagging protocol, and is used to control the digital human to execute a preset event at a preset time;

[0026] The digital human driving module is used to call the rendering engine and drive the pre-rendered digital human in the called rendering engine according to the driving information.

[0027] Optionally, before receiving the driving data packet for the digital human, the device further includes:

[0028] A construction information receiving module is used to receive construction information for a digital human sent by a user system; the construction information includes digital human information applicable to a preset scene, digital human costume information, digital human voice information, and digital human movement information;

[0029] The digital human generation module is used to generate a digital human adapted to a preset scene based on the construction information through the parsing engine, so as to pre-render the constructed digital human.

[0030] Optionally, after generating a digital human adapted to a preset scene based on the construction information, the apparatus further includes:

[0031] a rendering engine determination module, configured to determine, based on the digital human information applicable to the preset scene, a rendering engine corresponding to the preset scene from the rendering engines integrated by the parsing engine;

[0032] The digital human rendering module is used to call the rendering engine through the parsing engine, and pre-render the constructed digital human and the preset scene through the called rendering engine, so as to drive the digital human in the pre-rendered preset scene.

[0033] Optionally, the driver data packet parsing module includes:

[0034] The driving data packet receiving submodule is used to receive a driving data packet for the digital human sent by the user system through the parsing engine; the driving data packet for the digital human is a data packet converted by the user system into a digital human standard markup protocol by the content information that the digital human needs to display based on the construction information; wherein, the converted driving data packet uses elements used to represent preset events, and the elements of the preset events have a starting position attribute used to represent a preset moment to configure the content information that the digital human needs to display. The content information that the digital human needs to display includes speech events, action events and expression events that the digital human needs to perform at a preset moment, as well as a card insertion event at a preset moment in a preset scene.

[0035] Optionally, the driver data packet parsing module includes:

[0036] The driving data packet parsing submodule is used for the digital human to perform real-time parsing of the elements in the driving data packet used to represent preset events, as well as the starting position attributes of the elements of the preset events used to represent preset moments, through the parsing engine, to obtain the driving information of the digital human; the driving information includes the speech text information, action text information, card text information and expression text information at the preset moment.

[0037] Optionally, the driving information includes speech text information, action text information and expression text information at a preset time, and the digital human driving module includes:

[0038] a data conversion submodule for converting the spoken text information into streaming voice data in real time, and generating corresponding mouth shape data, expression data and motion data at a preset time based on the spoken text information, motion information and expression information during the streaming voice conversion process;

[0039] A data sending submodule is used to send the voice data, mouth shape data, expression data and action data generated at a preset time to the called rendering engine in real time through the parsing engine;

[0040] The digital human driving submodule is used to obtain a pre-rendered digital human through the rendering engine, and drive the pre-rendered digital human to play based on the voice data, mouth shape data, expression data and action data at a preset time, so as to push the played digital human to the user system.

[0041] Optionally, the driving information further includes card text information at a preset time, and the digital human driving module includes:

[0042] The card information insertion submodule is used to insert the card text information of the preset moment into the preset scene according to the preset moment during the process of driving the pre-rendered digital human to play.

[0043] An embodiment of the present application discloses an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements any step of the digital human driving method when executed by the processor.

[0044] An embodiment of the present application discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the digital human driving methods are implemented.

[0045] The embodiments of the present application include the following advantages:

[0046] In an embodiment of the present application, a digital human standard markup protocol configured based on a markup language is involved, which is applied to a parsing engine adapted to the configured digital human standard markup protocol. The parsing engine adapted to the digital human standard markup protocol can support different rendering engines. At this time, the received digital human driving data packet can be parsed by the parsing engine to obtain driving information, and then the parsing engine calls the rendering engine to drive the pre-rendered digital human according to the driving information. The driving data packet received by the parsing engine can be a data packet based on the configured digital human standard markup protocol, which can be used to control the digital human to execute preset events at preset times. The relevant driving standards of digital humans can be unified based on the proposed digital human standard tagging protocol and the parsing engine adapted to this tagging protocol. When driving digital humans, different rendering engines can be independent of the required rendering scenarios and related implementations. The parsing engine only needs to call the driving rendering engine based on the proposed digital human standard tagging protocol to drive the pre-rendered digital humans at a specified event at a certain moment. The production of digital humans can be standardized based on the digital human standard tagging protocol and then the rendering drive of digital humans can be unified, so that the rendering and driving of digital humans do not have the logic of related service scenarios, so as to cope with different upstream service scenarios and reuse the rendering drive aspects of digital humans. Moreover, the digital human standard tagging protocol based on this standardized unified and constrained rendering drive is not limited to any data parsing solution and any rendering drive system. It not only unifies the docking of upstream service scenarios, but also is not affected by the upgrade or change of downstream rendering drive engines. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is a flowchart of the steps of an embodiment of a digital human driving method of the present application;

[0048] Figure 2 This is a schematic diagram of the system framework driven by a digital human provided in an embodiment of the present application;

[0049] Figure 3 This is a flowchart of another embodiment of the digital human driving method of the present application;

[0050] Figure 4 This is a diagram of an application scenario driven by a digital human provided in an embodiment of the present application;

[0051] Figure 5 This is a structural block diagram of an embodiment of a digital human driving device of the present application. DETAILED DESCRIPTION

[0052] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0053] The development of virtual digital humans has entered a rapid growth stage. They can be applied to a variety of different scenarios with digital human needs, such as the gaming industry, live broadcast scenarios, etc., and manifested as different types of digital humans, such as digital virtual humans in live broadcasts, digital virtual humans in customer service, digital virtual humans in assistants, etc. However, in different scenarios, the rendering drivers for different types of virtual humans in different scenarios are universal. The non-reusable rendering driver solutions in the existing technology do not take advantage of this universality. That is, the existing digital human rendering driver solutions lack specifications for the rendering drivers of digital humans.

[0054] Reference Figure 1 , shows a flowchart of the steps of an embodiment of a digital human driving method of the present application, involving a digital human standard markup protocol configured based on a markup language, applied to a parsing engine adapted to the configured digital human standard markup protocol, the parsing engine supporting different rendering engines, and specifically including the following steps:

[0055] Step 101: receiving a driving data packet for a digital human through a parsing engine, and parsing the driving data packet to obtain driving information;

[0056] In an embodiment of the present application, in order to achieve the versatility of the driving rendering engine and the convenience of use in upstream service scenarios, the embodiment of the present application proposes to unify the relevant driving standards of digital humans based on the proposed digital human standard markup protocol and a parsing engine adapted to this markup protocol. The proposed language can be a standard VAML (Virtual Avatar Markup Language) language to control the rendering driving engine based on the proposed VAML language. It can be similar to the HTML language (Hyper Text Markup Language, which can be parsed and identified as a web page by a special browser) to describe the preset events performed by the digital human at the preset time, such as describing the digital human that the parsing engine needs to render within a period of time, the text spoken by the digital human, what actions and / or expressions are performed at what time, and the cards displayed, etc.

[0057] Just as HTML has a browser that parses it, the VAML language proposed in the embodiments of the present application also has a parsing engine that parses it. This parsing engine can be a VAML engine adapted to the configured digital human standard markup protocol. For clients related to upstream service scenarios, the VAML engine adapted to the VAML language is equivalent to a rendering driver engine, which can be used to quickly and easily render the digital human required by the service scenario, and the relevant clients of the upstream service scenario can drive the digital human according to demand. Internally, the VAML engine can refer to a set of digital human rendering driver frameworks that can organically unify digital human-related TTS (Text To Speech) technology, lip shape drive, expression drive, action drive, card drive, scene special effects and prop drive, etc., shielding internal implementation details from upstream service scenarios and supporting different rendering engines. The different rendering engines supported can refer to different rendering engines suitable for different service scenarios, such as the Unity game engine and the Unreal Engine. In this case, different rendering engines are called based on the VAML language and the called rendering engines drive the digital human according to the drive information contained in the VAML language, without being affected by upgrades or changes to the downstream rendering driver engine.

[0058] In one embodiment of the present application, the relevant driving standards of digital humans can be unified based on the proposed digital human standard markup protocol and the parsing engine adapted to this markup protocol. This can be manifested as the VAML parsing engine driving any different rendering engines based on the VAML protocol, thereby achieving the unification of digital human rendering drivers for different upstream service scenarios.

[0059] Specifically, the VAML parsing engine can receive a driving data packet for the digital human, and parse the received driving data packet to obtain driving information for rendering and driving the digital human. The received driving data packet can be a data packet based on the configured digital human standard markup protocol (i.e., VAML language), which can be used to control the digital human to perform preset events at preset times, that is, the driving information obtained by parsing the driving data packet based on the VAML language can realize the subsequent corresponding driving of the digital human.

[0060] In practical applications, refer to Figure 2, showing a schematic diagram of the system framework for driving a digital human provided by an embodiment of the present application. In the process of driving a digital human based on driving information, in addition to the parsing engine 12, the user system 11 may also be involved, and the parsing engine 12 may integrate different rendering engines 13. After receiving the construction information for the digital human sent by the user system 11, the parsing engine 12 may generate a digital human adapted to the preset scene based on the received construction information. The received construction information may include digital human information, digital human dressing information, digital human voice information and digital human action information suitable for the preset scene, so as to be used for subsequent digital human rendering and driving, and transmitted to the client of the upstream service scene for playback.

[0061] Among them, in addition to generating the digital human based on the received construction information, the driving data packet for the digital human sent by the user system received by the parsing engine can be a data packet of the digital human standard markup protocol converted by the user system into the content information that the digital human needs to display based on the construction information. That is, the driving data packet used for parsing can be based on the construction information for the digital human sent by the user system, encapsulated according to the VAML language protocol. Specifically, the main step is to convert the content information that the digital human needs to display based on the construction information into a data packet of the digital human standard markup protocol. In the converted driving data packet, elements used to represent preset events and the starting position attributes of the elements of preset events used to represent the preset time are used to configure the content information that the digital human needs to display. The content information that the digital human needs to display may include speech events, action events and expression events that the digital human needs to perform at the preset time, as well as card insertion events at the preset time in the preset scene, etc. At this time, the speech events, action events and expression events that the digital human needs to perform at the preset time, as well as the card insertion events at the preset time in the preset scene, etc., can be encapsulated as relevant elements of the corresponding events according to the protocol of the VAML language, and the starting position attribute representing the preset time and the attribute value are configured under this element to achieve the encapsulation of the driving data packet.

[0062] In a preferred embodiment, not only can the digital human's voice information and digital human's action information be encapsulated in the form of controlling preset events at preset moments in the VAML language, but after generating a digital human adapted to the preset scene based on the construction information, the rendering engine corresponding to the preset scene can also be determined from the rendering engines integrated by the parsing engine based on the digital human information applicable to the preset scene. The rendering engine is called by the parsing engine, and the constructed digital human and the preset scene are pre-rendered by the called rendering engine, thereby achieving pre-rendering of the rendered digital human. The pre-rendering here can refer to static rendering of the digital human based on the digital human's construction information, so that when the digital human is subsequently driven to render dynamically, the pre-statically rendered digital human can be directly driven, thereby reducing the playback delay of the digital human in the upstream preset scene client.

[0063] When parsing the received digital human driving data packet, the parsing engine can perform real-time analysis on the elements in the driving data packet used to represent preset events, as well as the starting position attributes of the elements of the preset events used to represent the preset time, to obtain the driving information of the digital human. The driving information may include speech text information, action text information, card text information, and expression text information at the preset time. At this time, the speech text information can also be converted into streaming voice data in real time. In the streaming voice conversion process, the corresponding mouth shape data, expression data, and action data at the preset time are generated based on the speech text information, action information, and expression information, so that the pre-rendered digital human can be driven according to the generated data.

[0064] Step 102: The rendering engine is called by the parsing engine, and the pre-rendered digital human is driven in the called rendering engine according to the driving information.

[0065] The driving data packet received by the parsing engine can be used to control the digital human to perform preset events at preset times. The driving information obtained based on the parsing of the driving data packet can also be used to describe the digital human that the parsing engine needs to render within a period of time, the text of the digital human's speech, the actions and / or expressions to be performed at what time, and the cards displayed, etc. At this time, the rendering engine can be called through the parsing engine, and the pre-rendered digital human can be driven in the called rendering engine according to the driving information, thereby realizing dynamic driving of the pre-statically rendered digital human.

[0066] The called rendering engine can be determined from multiple different rendering engines integrated by the parsing engine based on the digital human information applicable to the preset scenario. When calling different rendering engines to drive the digital human, it can be independent of the service scenario and related specific implementation required for rendering, such as the service scenario in which the digital human is located, how the digital human is rendered, how the action is issued, and other details. The parsing engine can simply call the driving rendering engine based on the proposed digital human standard tagging protocol to drive the pre-rendered digital human at a specified event at a certain moment, such as what to say in what scenario, what action to take when saying which word, what card to change after which word, etc., to complete the driving.

[0067] It should be noted that, for the internal part of the VAML engine, it can refer to a set of digital human rendering driver frameworks, in which the rendering engine, game engine, expression driver, etc. are integrated. The rendering engine is used to render digital humans. The rendering engine can be non-uniform, and the rendering engine for rendering each material can be different. It’s just that a standard set of protocols is set between the VAML engine and the rendering engine, namely the VAML language. At this time, you don’t need to care about what rendering engine is used. You only need to make self-decisions at different times to determine at a certain moment to trigger a certain action to render a certain special effect or prop. The rendering engine integrated into the VAML engine can render digital humans according to the VAML protocol, that is, the VAML engine can drive any rendering engine through the VAML protocol. When there is a need to replace the rendering engine in the future, you only need to replace the implementation part of the engine to realize the rendering and driving of the digital human part, which greatly improves the development efficiency of the digital human rendering driver in different service scenarios.

[0068] In one embodiment of the present application, when the pre-rendered digital human is driven in the called rendering engine according to the driving information, the voice data, mouth shape data, expression data and motion data generated at a preset time can be sent to the called rendering engine in real time through the parsing engine, and then the pre-rendered digital human is obtained through the rendering engine, and the pre-rendered digital human is driven to be played based on the voice data, mouth shape data, expression data and motion data at the preset time, so as to push the played digital human to the user system.

[0069] The driving information also includes card text information at a preset moment. In the process of driving the pre-rendered digital human to play, the card text information at the preset moment can also be inserted into the preset scene at the preset time. The inserted card information can include interactive cards and non-interactive cards (i.e., pure display). Interactive cards can be cards that users can respond to with touch. For example, in a live broadcast scene, product information can be displayed through cards in the area below the digital human. At this time, when the user clicks on the card with this product information, they can jump to the purchase interface, etc.

[0070] In an embodiment of the present application, a digital human standard markup protocol configured based on a markup language is involved, which is applied to a parsing engine adapted to the configured digital human standard markup protocol. The parsing engine adapted to the digital human standard markup protocol can support different rendering engines. At this time, the received digital human driving data packet can be parsed by the parsing engine to obtain driving information, and then the parsing engine calls the rendering engine to drive the pre-rendered digital human according to the driving information. The driving data packet received by the parsing engine can be a data packet based on the configured digital human standard markup protocol, which can be used to control the digital human to execute preset events at preset times. The relevant driving standards of digital humans can be unified based on the proposed digital human standard tagging protocol and the parsing engine adapted to this tagging protocol. When driving digital humans, different rendering engines can be independent of the required rendering scenarios and related implementations. The parsing engine only needs to call the driving rendering engine based on the proposed digital human standard tagging protocol to drive the pre-rendered digital humans at a specified event at a certain moment. The production of digital humans can be standardized based on the digital human standard tagging protocol and then the rendering drive of digital humans can be unified, so that the rendering and driving of digital humans do not have the logic of related service scenarios, so as to cope with different upstream service scenarios and reuse the rendering drive aspects of digital humans. Moreover, the digital human standard tagging protocol based on this standardized unified and constrained rendering drive is not limited to any data parsing solution and any rendering drive system. It not only unifies the docking of upstream service scenarios, but also is not affected by the upgrade or change of downstream rendering drive engines.

[0071] Reference Figure 3 , shows a flowchart of another embodiment of a digital human driving method of the present application, involving a digital human standard markup protocol configured based on a markup language, and applied to a parsing engine adapted to the configured digital human standard markup protocol, and specifically may include the following steps:

[0072] Step 301: Analyze the driving data packet in real time through the analysis engine to obtain the driving information of the digital human, and generate the mouth shape data, expression data and movement data at a preset time based on the driving information;

[0073] In an embodiment of the present application, the relevant driving standards of digital humans can be unified based on the proposed digital human standard markup protocol and a parsing engine adapted to the markup protocol. Specifically, the parsing engine can call the rendering engine for driving based on the proposed digital human standard markup protocol VAML language.

[0074] The proposed VAML language is similar to the HTML language. It can be used by the user system to encapsulate or convert data packets of the digital human standard markup protocol based on the content information that the digital human needs to display based on the construction information of the digital human. Among them, the construction information for the digital human can include digital human information applicable to the preset scene, digital human dressing information, digital human voice information and digital human action information.

[0075] Specifically, the main step is to convert the content information that the digital human needs to display based on the construction information into a data packet of the digital human standard markup protocol. In the converted driving data packet, elements used to represent preset events and the starting position attributes of the elements of preset events used to represent the preset time are used to configure the content information that the digital human needs to display. The content information that the digital human needs to display may include speech events, action events and expression events that the digital human needs to perform at the preset time, as well as card insertion events at the preset time in the preset scene, etc. At this time, the speech events, action events and expression events that the digital human needs to perform at the preset time, as well as the card insertion events at the preset time in the preset scene, etc., can be encapsulated as relevant elements of the corresponding events according to the protocol of the VAML language, and the starting position attribute representing the preset time and the attribute value are configured under this element to achieve the encapsulation of the driving data packet.

[0076] When the driving data packet is parsed in real time by the parsing engine to obtain the driving information of the digital human, the parsed driving information may include the spoken text information, action text information, card text information and expression text information at a preset moment. The lip shape data, expression data and action data at a preset moment may also be generated based on the driving information. Specifically, the spoken text information may be converted into streaming voice data in real time, and the corresponding lip shape data, expression data and action data at a preset moment may be generated based on the spoken text information, action information and expression information during the streaming voice conversion process, so that the digital human can be played and driven based on the generated data later.

[0077] Step 302: The rendering engine drives the pre-rendered digital human to play based on the voice data, mouth shape data, expression data and action data at the preset time, and inserts the card text information at the preset time into the preset scene according to the preset time.

[0078] The driving data packet received by the parsing engine can be used to control the digital human to perform preset events at preset times. The driving information obtained based on the parsing of the driving data packet can also be used to describe the digital human that the parsing engine needs to render within a period of time, the text of the digital human's speech, the actions and / or expressions to be performed at what time, and the cards displayed, etc. At this time, the rendering engine can be called through the parsing engine, and the pre-rendered digital human can be driven in the called rendering engine according to the driving information, thereby realizing dynamic driving of the pre-statically rendered digital human.

[0079] Specifically, it can be manifested as driving the pre-rendered digital human to play based on the voice data, mouth shape data, expression data and action data at the preset moment through the rendering engine, inserting the card text information at the preset moment into the preset scene at the preset moment, and pushing the played digital human to the user system.

[0080] For example, referring to Figure 4 , showing an application scenario diagram of the digital human drive provided by an embodiment of the present application, assuming that there is a preset scenario client upstream of the parsing engine 13 (i.e., the VAML engine), and its preset scenarios may include live broadcast scenarios, customer service scenarios, and assistant scenarios, etc. The digital human information applicable to these preset scenarios may include digital virtual human live broadcast, digital virtual human customer service, digital virtual human assistant, etc.

[0081] For the digital human-driven process, first, the user system can determine the digital human to be used. The user system has a built-in virtual human that provides the user with character construction according to the service selection, which can determine the digital human information suitable for the preset scene. At this time, the user can also upload a self-made digital human that follows the platform's 3D resource specifications in the user system. After determining the digital human corresponding to the preset service scene, the selected digital human can be dressed up. At this time, the built-in costumes and props can also be used to select the character's dress. The user can also upload self-made dress information that also needs to follow the platform's resource specifications to determine the digital human's dress information. At this time, the digital human can also be dressed up based on the built-in sounds and actions. By selecting the actions, sounds, etc., the digital human's sound information and digital human's action information can be determined. At this time, the user can also upload self-made sound information and action information that comply with the platform's 3D resource specifications to the user system; secondly, a rendering engine can be started, and the digital human selected in the above steps can be formulated through the rendering engine. The constructed digital human and the preset scene can be statically rendered in advance through the rendering engine, so that the digital human can be quickly determined during subsequent rendering, reducing playback delays. At this time, the user system can also calculate the content that needs to be displayed by the virtual human based on the construction information of the digital human, such as the digital human in the live scene, object display and other information Information, etc., and convert the displayed content into VAML protocol, so that the subsequent VAML engine can parse the received VAML protocol drive data packet to obtain the digital person's speech text information, action text information, expression text information and card text information at the preset time, and in the process of converting the speech text information into streaming voice data in real time, relevant algorithms can be added to the VAML engine, such as text-to-speech, text-to-action, text-to-face drive, etc., to generate the corresponding mouth shape data, expression data and action data at the preset time from the speech text information, action information and expression information. At this time, the VAML engine can generate The voice data, mouth shape data, expression data, and action data of each frame are pushed to the rendering engine in real time for playback. This rendering can be dynamic rendering. After receiving the VAML language, it can quickly obtain the corresponding digital human for relevant rendering control, reducing the playback delay of the digital human. During the playback process, the pre-rendered digital human can be driven to play based on the voice data, mouth shape data, expression data and action data at the preset moment, and the card text information at the preset moment can be inserted into the preset scene according to the preset time. For example, cards, special effects, props, etc. can be inserted at the time point specified by the service. Finally, the rendered digital human and sound can be pushed to the upstream service scene for playback.

[0082] Taking the live broadcast scenario as an example, the displayed content can be converted into the VAML protocol. For example, the converted VAML language can be as follows:

[0083]

[0084]

[0085] Based on the speech events, action events and expression events that the digital human needs to perform at the preset time, as well as the card insertion events at the preset time in the preset scene, the content that needs to be displayed can be mainly encapsulated based on the elements used to represent the preset events, such as the elements <scene> 、 <section> 、 <avatar> 、 <room> 、、 <interact> 、 <speech>as well as <action>Among them, the VAML language <scene>Can be used to represent a scene in a script. <section>Can be used to represent a scene fragment in the scene, <avatar>Can be used to represent the virtual anchor in the scene clip, <room>It can be used to represent the room information in the scene fragment and the h5 layer in the scene fragment. <interact>It can be used to represent the interactive information on the scene fragment end. <speech>Can be used to represent spoken content in a scene clip.

[0086] (1) Specifically, for a scene in a script <scene>Elements, such as live broadcast scenes:

[0087]

[0088]

[0089] Among them, located <scene>The child elements under the element may include a scene fragment <section>element.

[0090] (2) For a scene segment representing a scene <section>Elements, such as:

[0091]

[0092] in, <section>The unique_code attribute of the element can be a uuid, which can be used to represent the unique value bound to the section. <section>Sub-elements under the element can include <avatar> 、 <room> 、、 <interact>as well as <speech>.

[0093] (3) For the virtual anchor in the scene segment <avatar>, for example:

[0094] avatar character_code="CH_h9HtQeNIAU">

[0095] <action begin_index="2" end_index="4" intent="TAG_VAML_Emo_Pos"

[0096] interrupt="false" / >

[0097] <action begin_index="12" end_index="15" intent="TAG_VAML_Emo_Pos"

[0098] interrupt="false" / >

[0099] <action begin_index="17" end_index="17" intent="TAG_VAML_indicate_default"

[0100] interrupt="false" / >

[0101] < / avatar>

[0102] in, <avatar>The attribute of the element is character_code, and its value can be character code, which is used as the unique code of the character in the virtual human construction platform. <avatar>Sub-elements under the element can include <action>.

[0103] (4) For elements used to represent the executed action event <action>,For example:

[0104] <action begin_index="2"end_index="4"intent="TAG_VAML_Emo_Pos"

[0105] interrupt="false" / >

[0106] in, <action>Element attributes can include attributes for presetting the starting position of an event at a preset time, such as begin_index = the starting trigger position in the text, which indicates the starting position of the text in the speech; end_index = the ending position in the text, which indicates the starting and ending positions of the text in the speech; and intent = the intent, which indicates the action tag value. Attribute values ​​can also include interrupt = true / false, where true indicates that the action can be interrupted and false indicates that the action cannot be interrupted.

[0107] (5) For room information in scene segments <room>, which is coupled with the virtual human and can only be implemented in Unity, for example:

[0108]

[0109] in, <room>The attributes of the element can include layout=code, which mainly refers to the layout code; background=url, which is used to describe the 2D background image; bgm=url, which is used to describe the background music. <room>Sub-elements under the element can include <layout> 、 <background>and <bgm>.

[0110] (6) For elements representing layout <layout>,For example:

[0111] <layout code="bizspace::default" / >

[0112] in, <layout>The attributes of the element may include code=bizspace::default, which is used to describe the layout code.

[0113] (7) For elements representing background <background>,For example:

[0114] <background url="" / >

[0115] in, <background>The attributes of the element may include url=background image, which is used to describe a 2D background image.

[0116] (8) For elements representing scenes <scene>,For example:

[0117] <scene url="" / >

[0118] in, <scene>The attributes of the element include url=packaged scene resource file URL, which is used to describe the scene resource file.

[0119] (9) For the personalized cards of service scenes in the scene fragments, they do not rely on the capabilities of Unity and are stored as display cards. If you need to interact with the user, move to the tag. If you need Unity capabilities to achieve this, move to <room>Labels, such as:

[0120]

[0121] <card begin_index="1" end_index="8" code="bizspace::subtitle" data="Household high-power hair dryer," / >

[0122] <card begin_index="9" end_index="15" code="bizspace::subtitle" data="Does not damage hair and dries hair quickly." / >

[0123] <card begin_index="15" end_index="26" code="bizspace::subtitle" data="Then its additional function is quick drying." / >

[0124]

[0125] Among them, the subelements under the element include <card>.

[0126] (10) Elements describing card information <card>, which is a custom card for different service scenarios. Here, the code is the unique identity on the card platform, and the data is the service data required for card rendering. The data format is defined by the card itself. For example:

[0127] <card begin_index="15" end_index="26" code="bizspace::subtitle" data="Then its additional function is quick-drying." / >

[0128] Among them, <card>The attributes of the element may include begin_index = the starting trigger position in the text, used to describe the starting position of the text in the speech; end_index = the ending position in the text, used to describe the ending position of the text in the speech; code = the unique ID of the card, used to describe the unique ID in the card platform; data = the data required for the card exhibition, used to describe the data required for the card exhibition.

[0129] (11) For interactive information on the middle end of the scene segment <interact>, the interactive cards are handled by each end, and VAML is only responsible for triggering and destroying the cards at the corresponding time, for example:

[0130] <interact>

[0131] <signal begin_index="2"end_index="4"code=""data="" / >

[0132] < / interact>

[0133] in, <interact>Sub-elements under the element can include <signal>element.

[0134] (12) For elements <signal>,For example:

[0135] <signal begin_index="2"end_index="4"code=""data="" / >

[0136] in, <signal>The attributes of the element may include begin_index = the starting trigger position in the text, which is used to describe the starting position of the text in the speech; end_index = the ending position in the text, which is used to describe the ending position of the text in the speech; code = the unique value of the card, which is used to describe that the unique value of the card is managed by the end scene itself and is used to control its card display and destruction; data = the data required for the card exhibition, which is used to describe the data required for the card exhibition.

[0137] (13) For the speech content in the scene segment <speech>,For example:

[0138] <speech>

[0139] A high-power hair dryer for home use that dries hair quickly without damaging it.

[0140] < / speech>

[0141] in, <speech>Sub-elements under the element can include text.

[0142] In a preferred embodiment, after real-time parsing of the driving data packet based on VAML protocol conversion, the driving information of the digital human can be obtained, for example, the attribute values ​​of begin_index and end_index are parsed to obtain the preset time, and the element <speech>The digital human's spoken text message, "A high-power hair dryer for home use, no hair damage, and quick hair drying," is parsed and converted into streaming voice data in real time. During the streaming voice conversion process, corresponding lip shape data, facial expression data, and motion data at a preset moment are generated based on the spoken text message, motion information, and facial expression information. At the preset moment when the text begins, the digital human is driven to produce the speech message "A high-power hair dryer for home use, no hair damage, and quick hair drying" according to the lip shape data, facial expression data, and motion data. It should be noted that the parsing and activation processes for other elements representing other preset events are similar to those described above and are not limited in this application.

[0143] In an embodiment of the present application, the relevant driving standards of digital humans can be unified based on the proposed digital human standard tagging protocol and the parsing engine adapted to this tagging protocol. When driving digital humans, different rendering engines can be driven regardless of the required rendering scenarios and related implementations. The parsing engine only needs to call the driving rendering engine based on the proposed digital human standard tagging protocol to drive the pre-rendered digital humans at a specified event at a certain moment. The production of digital humans can be standardized based on the digital human standard tagging protocol and then the rendering drive of digital humans can be unified, so that the rendering and driving of digital humans do not have the logic of related service scenarios, so as to cope with different upstream service scenarios and reuse the rendering drive aspects of digital humans. The digital human standard tagging protocol based on this standardized, unified and constrained rendering drive is not limited to any data parsing solution and any rendering drive system. It not only unifies the docking of upstream service scenarios, but also is not affected by the upgrade or change of the downstream rendering drive engine.

[0144] It should be noted that for the method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the order of the actions described, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.

[0145] Reference Figure 5 , shows a structural block diagram of an embodiment of a digital human driving device of the present application, involving a digital human standard markup protocol configured based on a markup language, and a parsing engine adapted to the configured digital human standard markup protocol. The parsing engine supports different rendering engines and may specifically include the following modules:

[0146] The driving data packet parsing module 501 is used to receive a driving data packet for the digital human and parse the driving data packet to obtain driving information; wherein the received driving data packet is a data packet based on the configured digital human standard tagging protocol, and is used to control the digital human to execute a preset event at a preset time;

[0147] The digital human driving module 502 is used to call the rendering engine and drive the pre-rendered digital human in the called rendering engine according to the driving information.

[0148] In one embodiment of the present application, before receiving the driving data packet for the digital human, the device may further include the following modules:

[0149] A construction information receiving module is used to receive construction information for a digital human sent by a user system; the construction information includes digital human information applicable to a preset scene, digital human costume information, digital human voice information, and digital human movement information;

[0150] The digital human generation module is used to generate a digital human adapted to a preset scene based on the construction information through the parsing engine, so as to pre-render the constructed digital human.

[0151] In one embodiment of the present application, after generating a digital human adapted to a preset scene based on the construction information, the apparatus may further include the following modules:

[0152] a rendering engine determination module, configured to determine, based on the digital human information applicable to the preset scene, a rendering engine corresponding to the preset scene from the rendering engines integrated by the parsing engine;

[0153] The digital human rendering module is used to call the rendering engine through the parsing engine, and pre-render the constructed digital human and the preset scene through the called rendering engine, so as to drive the digital human in the pre-rendered preset scene.

[0154] In one embodiment of the present application, the driver data packet parsing module 501 may include the following submodules:

[0155] The driving data packet receiving submodule is used to receive a driving data packet for the digital human sent by the user system through the parsing engine; the driving data packet for the digital human is a data packet converted by the user system into a digital human standard markup protocol by the content information that the digital human needs to display based on the construction information; wherein, the converted driving data packet uses elements used to represent preset events, and the elements of the preset events have a starting position attribute used to represent a preset moment to configure the content information that the digital human needs to display. The content information that the digital human needs to display includes speech events, action events and expression events that the digital human needs to perform at a preset moment, as well as a card insertion event at a preset moment in a preset scene.

[0156] In one embodiment of the present application, the driver data packet parsing module 501 may include the following submodules:

[0157] The driving data packet parsing submodule is used for the digital human to perform real-time parsing of the elements in the driving data packet used to represent preset events, as well as the starting position attributes of the elements of the preset events used to represent preset moments, through the parsing engine, to obtain the driving information of the digital human; the driving information includes the speech text information, action text information, card text information and expression text information at the preset moment.

[0158] In one embodiment of the present application, the digital human driving module 502 may include the following submodules:

[0159] a data conversion submodule for converting the spoken text information into streaming voice data in real time, and generating corresponding mouth shape data, expression data and motion data at a preset time based on the spoken text information, motion information and expression information during the streaming voice conversion process;

[0160] A data sending submodule is used to send the voice data, mouth shape data, expression data and action data generated at a preset time to the called rendering engine in real time through the parsing engine;

[0161] The digital human driving submodule is used to obtain a pre-rendered digital human through the rendering engine, and drive the pre-rendered digital human to play based on the voice data, mouth shape data, expression data and action data at a preset time, so as to push the played digital human to the user system.

[0162] In one embodiment of the present application, the driving information further includes card text information at a preset time. The digital human driving module 502 may include the following submodules:

[0163] The card information insertion submodule is used to insert the card text information of the preset moment into the preset scene according to the preset moment during the process of driving the pre-rendered digital human to play.

[0164] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0165] An embodiment of the present application further provides an electronic device, including:

[0166] It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned digital human driving method embodiment are implemented and the same technical effects can be achieved. To avoid repetition, it will not be described here.

[0167] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned digital human driving method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0168] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0169] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0170] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0171] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0172] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0173] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0174] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0175] The above is a detailed introduction to the digital human driving method, device, equipment and storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the contents of this specification should not be understood as limiting the present application.< / speech> < / speech> < / speech> < / signal> < / signal> < / signal> < / interact> < / interact> < / card> < / card> < / card> < / room> < / scene> < / scene> < / background> < / background> < / layout> < / layout> < / bgm> < / background> < / layout> < / room> < / room> < / room> < / action> < / action> < / action> < / avatar> < / avatar> < / speech> < / interact> < / room> < / avatar> < / section> < / section> < / section> < / section> < / scene> < / scene> < / speech> < / interact> < / room> < / avatar> < / section> < / scene> < / action> < / speech> < / interact> < / room> < / avatar> < / section> < / scene>

Claims

1. A digital human driving method, characterized in that: This method involves a digital human standard markup protocol configured based on a markup language, and is applied to a parsing engine adapted to the configured digital human standard markup protocol. The parsing engine supports rendering engines for different service scenarios. The markup language is VAML. The method includes: The parsing engine receives a driving data packet for the digital human and parses the driving data packet to obtain driving information; wherein the received driving data packet is a data packet based on the configured digital human standard tagging protocol and is used to control the digital human to execute a preset event at a preset time; Calling a rendering engine through the parsing engine, and driving the pre-rendered digital human in the called rendering engine according to the driving information; The step of parsing the driver data packet to obtain driver information includes: The parsing engine performs real-time parsing on the elements in the driving data packet that represent preset events, as well as the starting position attributes of the elements of the preset events that represent preset moments, to obtain the driving information of the digital human; the driving information includes speech text information, action text information, card text information, and expression text information at the preset moment.

2. The method according to claim 1, characterized in that Before receiving the driving data package for the digital human, it also includes: Receive construction information for a digital human sent by a user system; the construction information includes digital human information applicable to a preset scene, digital human costume information, digital human voice information, and digital human movement information; The parsing engine generates a digital human adapted to a preset scene based on the construction information, so as to pre-render the constructed digital human.

3. The method according to claim 2, characterized in that After generating a digital human adapted to a preset scene based on the construction information, the method further includes: Determining a rendering engine corresponding to the preset scene from the integrated rendering engines based on the digital human information applicable to the preset scene; The rendering engine is called by the parsing engine, and the constructed digital human and the preset scene are rendered in advance by the called rendering engine, so as to drive the digital human in the pre-rendered preset scene.

4. The method according to claim 2, characterized in that The receiving of a driving data packet for the digital human by the parsing engine includes: The parsing engine receives a driving data packet for the digital human sent by the user system; the driving data packet for the digital human is a data packet converted by the user system into a digital human standard markup protocol by converting content information that the digital human needs to display based on the construction information; wherein, in the converted driving data packet, elements for representing preset events and the starting position attributes of the elements of the preset events for representing preset moments are used to configure the content information that the digital human needs to display, and the content information that the digital human needs to display includes speech events, action events and expression events that the digital human needs to perform at the preset moment, as well as card insertion events at the preset moment in the preset scene.

5. The method according to claim 1, wherein The driving information includes speech text information, action text information, and expression text information at a preset time. The driving of the pre-rendered digital human in the called rendering engine according to the driving information includes: The parsing engine converts the spoken text information into streaming voice data in real time, and generates the mouth shape data, expression data and movement data of the digital human at a preset moment based on the spoken text information, movement information and expression information during the streaming voice conversion process; The speech data, mouth shape data, expression data and action data generated at a preset time are sent in real time to the called rendering engine through the parsing engine; The rendering engine obtains a pre-rendered digital human, and drives the pre-rendered digital human to play based on the voice data, mouth shape data, expression data and action data at a preset time, so as to push the played digital human to the user system.

6. The method according to claim 5, characterized in that The driving information also includes card text information at a preset time. The calling rendering engine drives the pre-rendered digital human according to the driving information, including: In the process of driving the pre-rendered digital human to play, the card text information at the preset moment is inserted into the preset scene according to the preset moment.

7. A digital human driving device, characterized in that: It involves a digital human standard markup protocol configured based on a markup language, applied to a parsing engine adapted to the configured digital human standard markup protocol, the parsing engine supporting rendering engines for different service scenarios, the markup language being VAML, and the device comprising: A driving data packet parsing module is used to receive a driving data packet for the digital human and parse the driving data packet to obtain driving information; wherein the received driving data packet is a data packet based on the configured digital human standard tagging protocol, and is used to control the digital human to execute a preset event at a preset time; A digital human driving module, configured to call a rendering engine and drive the pre-rendered digital human in the called rendering engine according to the driving information; The driver data packet parsing module includes the following submodules: The driving data packet parsing submodule is used for the digital human to perform real-time parsing of the elements in the driving data packet used to represent preset events, as well as the starting position attributes of the elements of the preset events used to represent preset moments, through the parsing engine, to obtain the driving information of the digital human; the driving information includes the speech text information, action text information, card text information and expression text information at the preset moment.

8. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the steps of the digital human driving method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the digital human driving method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Information processing method and device, equipment and medium

    CN112306324A

  • Digital human rendering method and device, storage medium and electronic equipment

    CN113886551A