Digital human image driving method and device, equipment, storage medium and product

By obtaining image feature data through semantic analysis of dialogue text, the digital human is controlled for animation driving, which solves the problem of single image change of the digital human and improves the user experience.

CN120707709APending Publication Date: 2025-09-26TIANJIN FAW TOYOTA MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510863538.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the existing technology, the response actions of digital humans are fixed, resulting in relatively simple image changes and poor matching with the broadcast content, resulting in a poor user experience.

Method used

By performing semantic analysis on the dialogue text input by the user, image feature data is obtained, and based on this data, the digital human is controlled to make image animations that match the semantic content of the text, including movements of the mouth, limbs and facial expressions.

Benefits of technology

It enriches the content of digital human image changes, improves the degree of anthropomorphism, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707709A_ABST
    Figure CN120707709A_ABST
Patent Text Reader

Abstract

The invention discloses a digital human image driving method and device, equipment, a storage medium and a product, and relates to the technical field of computers. The method comprises the following steps: obtaining a dialogue text, wherein the dialogue text is a text responded by an exponential word person based on user input; semantic analysis is conducted on the dialogue text, image feature data matched with semantic content of the dialogue text are obtained, and the image feature data are used for indicating image change features of the digital human. And based on the image feature data, controlling the digital human to make an image animation matched with the semantic content of the dialogue text. According to the method, the content expressed by the digital person is matched with the image change item of the digital person by performing semantic analysis on the responded text, so that the image change content of the digital person is enriched, several fixed actions are not used any more, the anthropomorphic degree of the digital person is improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to digital human image driving methods, devices, equipment, storage media and products. Background Art

[0002] With the rapid development of emerging smart cars, in-car voice assistants have evolved from virtual avatars to humanoid digital humans. When users interact with a digital human, it can perform several fixed actions in response to voice commands. When a user issues a specific voice command, the digital human identifies the scenario it corresponds to and responds accordingly, performing the corresponding action. For example, the command: "Xiao A, turn on the air conditioner." The car computer identifies this as a "command-type" scenario. Xiao A's response: the air conditioner is turned on, possibly accompanied by a simple gesture such as a nod or a smile. Another example: the command: "Xiao A, play music." The car computer identifies this as a "command-type" scenario. Xiao A's response: the music starts playing, and Xiao A may also say "OK" or "I've played it for you."

[0003] In related technologies, a preset action library is built into the vehicle computer, storing actions and corresponding triggering commands. The digital human's responses rely primarily on this preset, fixed action library to handle different situations. For example, if a user says "play music," the vehicle computer will interpret this as an "entertainment" scenario and might trigger actions like "smile" or "nod to the beat."

[0004] However, in the above-mentioned related technologies, the response actions of the digital human are fixed, that is, according to the scene to which the trigger instruction belongs, the action corresponding to the scene is executed, resulting in the image changes of the digital human being being relatively simple and programmed, and the matching between the image of the digital human and the content of the digital human's broadcast is poor. Summary of the Invention

[0005] This application provides a digital human image driving method, device, equipment, storage medium and product, which enriches the content of digital human image changes, makes the digital human image and the digital human broadcast content more compatible, and improves the user experience.

[0006] To achieve the above objectives, this application adopts the following technical solutions:

[0007] In a first aspect, a method for driving a digital human image is provided, which is applied to a digital human image driving engine. The method comprises: obtaining a dialogue text, which is a text in which a digital human responds to user input; performing semantic parsing on the dialogue text to obtain image feature data that matches the semantic content of the dialogue text, wherein the image feature data is used to indicate image change characteristics of the digital human; and controlling the digital human to produce an image animation that matches the semantic content of the dialogue text based on the image feature data.

[0008] The solution provided in this application performs semantic analysis on the responsive text, so that the content expressed by the digital human matches the image changes of the digital human, enriching the image changes of the digital human instead of using a few fixed actions, thereby improving the degree of anthropomorphism of the digital human and enhancing the user experience.

[0009] In a possible implementation, the image feature data includes one or more of mouth feature data, body movement feature data, or facial expression feature data.

[0010] In another possible implementation, the digital human image driving engine includes a text-to-speech engine. The aforementioned semantic parsing of the conversation text to obtain image feature data matching the semantic content of the conversation text can be specifically implemented by: using the text-to-speech engine to perform speech conversion on the conversation text to obtain speech information that conforms to the digital human character setting. The text-to-speech engine decomposes the text in the conversation text into phonemes to obtain the phonemes corresponding to each word, and records the start and end times of the phonemes corresponding to each word in the speech information. A phoneme is a pronunciation unit corresponding to a word. Based on the start and end times of the phonemes corresponding to each word in the speech information, the speech information and the phonemes corresponding to each word are time-aligned to obtain the mouth feature data.

[0011] In another possible implementation, the above-mentioned control of the digital human to produce an image animation that matches the semantic content of the dialogue text based on the image feature data can be specifically realized as follows: based on the mouth feature data, the digital human is controlled to emit the language information, and the digital human's mouth is controlled to produce lip movements corresponding to the phonemes corresponding to the respective characters.

[0012] In another possible implementation, the digital human image driving engine includes a behavior inference engine. The semantic parsing of the dialogue text to obtain image feature data matching the semantic content of the dialogue text can be specifically implemented by using the behavior inference engine to break down the text in the dialogue text into text segments. Based on these text segments, selection is made in a behavior tree to obtain the body movement feature data matching the semantic content of the dialogue text. The behavior tree includes different action information corresponding to different text segments.

[0013] In another possible implementation, the above-mentioned control of the digital human to produce an image animation that matches the semantic content of the dialogue text based on the image feature data can be specifically realized as follows: based on the body movement feature data, the digital human's body is controlled to produce body movements corresponding to the text segment.

[0014] In another possible implementation, the digital human image driving engine includes: an emotion analysis engine; the semantic parsing of the dialogue text to obtain image feature data that matches the semantic content of the dialogue text can be specifically implemented as follows: using the emotion analysis engine to filter out emotional text in the dialogue text, where the emotional text is text used to express emotions; and obtaining the facial expression feature data based on the type and emotional level of the emotional text.

[0015] In another possible implementation, obtaining the facial expression feature data based on the type and emotional level of the emotional text may be specifically implemented by: determining the facial expression corresponding to the emotional text based on the type of the emotional text; obtaining the magnitude of the facial expression based on the emotional level of the emotional text; and obtaining the facial expression feature data based on the facial expression and the magnitude of the facial expression.

[0016] In another possible implementation, the above-mentioned controlling the digital human to produce an image animation that matches the semantic content of the dialogue text based on the image feature data can be specifically realized as follows: based on the facial expression feature data, controlling the face of the digital human to produce a facial expression corresponding to the emotional text.

[0017] In another possible implementation, the digital human image driving engine is deployed on the vehicle side and / or in the cloud.

[0018] In another possible implementation, the digital human image driving method provided herein further includes: when the image feature data includes at least two of the mouth feature data, the body movement feature data, or the facial expression feature data, synchronizing the at least two feature data according to a timestamp. Based on the synchronized at least two feature data, controlling the digital human to produce an image animation that matches the semantic content of the dialogue text.

[0019] In a second aspect, a digital human image driving device is provided, which is applied to a digital human image driving engine. The device includes: an acquisition module, a parsing module, and a driving module.

[0020] The acquisition module is used to acquire the dialogue text, which refers to the text in which the digital human responds based on the user input.

[0021] The above-mentioned analysis module is used to perform semantic analysis on the dialogue text to obtain image feature data that matches the semantic content of the dialogue text. The image feature data is used to indicate the image change characteristics of the digital human.

[0022] The driving module is used to control the digital human to produce an image animation that matches the semantic content of the dialogue text based on the image feature data.

[0023] In a possible implementation, the image feature data includes one or more of mouth feature data, body movement feature data, or facial expression feature data.

[0024] In another possible implementation, the digital human image driving engine includes a text-to-speech engine. The parsing module is further configured to perform speech conversion on the dialogue text using the text-to-speech engine to obtain speech information that conforms to the digital human character settings. The text-to-speech engine decomposes the text in the dialogue text into phonemes to obtain the phonemes corresponding to each word, and records the start and end times of the phonemes corresponding to each word in the speech information. A phoneme is a pronunciation unit corresponding to a word. Based on the start and end times of the phonemes corresponding to each word in the speech information, the speech information and the phonemes corresponding to each word are time-aligned to obtain the mouth feature data.

[0025] In another possible implementation, the driving module is further configured to control the digital human to emit the language information based on the mouth feature data, and to control the digital human's mouth to make lip movements corresponding to the phonemes corresponding to the characters.

[0026] In another possible implementation, the digital human avatar driving engine includes a behavior inference engine. The parsing module is further configured to use the behavior inference engine to decompose the text in the dialogue text into text segments. Based on the text segments, a selection is made in a behavior tree to obtain the body movement feature data that matches the semantic content of the dialogue text. The behavior tree includes different action information corresponding to different text segments.

[0027] In another possible implementation, the driving module is further configured to control the limbs of the digital human to perform limb movements corresponding to the text segment based on the limb movement feature data.

[0028] In another possible implementation, the digital human image driving engine includes: an emotion analysis engine; the parsing module is further used to filter out emotional text in the dialogue text through the emotion analysis engine, where the emotional text is text used to express emotions; and obtain the facial expression feature data based on the type and emotional level of the emotional text.

[0029] In another possible implementation, the parsing module is further configured to determine a facial expression corresponding to the emotional text based on the type of the emotional text; obtain a magnitude of the facial expression based on the emotional level of the emotional text; and obtain facial expression feature data based on the facial expression and the magnitude of the facial expression.

[0030] In another possible implementation, the driving module is further configured to control the face of the digital human to make a facial expression corresponding to the emotional text based on the facial expression feature data.

[0031] In another possible implementation, the digital human image driving engine is deployed on the vehicle side and / or in the cloud.

[0032] In another possible implementation, the parsing module is further configured to, when the image feature data includes at least two of the mouth feature data, the body movement feature data, or the facial expression feature data, synchronize the at least two feature data according to the timestamp. Based on the synchronized at least two feature data, the digital human is controlled to produce an image animation that matches the semantic content of the dialogue text.

[0033] The digital human image driving device provided in the second aspect is used to execute the digital human image driving method provided in the first aspect or any possible implementation method of the first aspect. The technical effects corresponding to any implementation method in the second aspect can be referred to the technical effects corresponding to any implementation method in the first aspect, and will not be repeated here.

[0034] In a third aspect, a computer device is provided, comprising: a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the digital human image driving method as described in the first aspect or any possible implementation method of the first aspect.

[0035] In a fourth aspect, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the digital human image driving method as described in the first aspect or any one of the implementation methods of the first aspect.

[0036] In a fifth aspect, a computer program product is provided, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the digital human image driving method described in the first aspect or any one of the implementation methods of the first aspect is implemented.

[0037] In the sixth aspect, an embodiment of the present application provides a chip system comprising at least one processor and at least one interface circuit, wherein the at least one interface circuit is used to perform transceiver functions and send instructions to the at least one processor. When the at least one processor executes the instructions, the at least one processor executes to implement the digital human image driving method as described in the first aspect or any one of the implementation methods in the first aspect.

[0038] In a seventh aspect, an embodiment of the present application provides a vehicle, comprising a display screen on which a digital human is displayed, wherein the vehicle drives the digital human based on the digital human image driving method described in the first aspect or any possible implementation of the first aspect.

[0039] The solutions provided in aspects 3 to 7 above are used to implement the first aspect or the method provided in aspect 1 above, and their specific implementations are not described in detail here. The technical effects corresponding to any implementation of the solutions provided in aspects 3 to 7 above can be referred to the technical effects corresponding to the first aspect or any implementation of aspect 1 above, and are not described in detail here.

[0040] It should be noted that various possible implementations of any of the above aspects can be combined under the premise that the solutions are not contradictory. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a schematic diagram of triggering a digital human to perform an action provided by an exemplary embodiment;

[0042] Figure 2 is a schematic diagram of the architecture of a computing system provided by an exemplary embodiment of the present application;

[0043] Figure 3 is a flow chart of a digital human image driving method provided by an exemplary embodiment of the present application;

[0044] Figure 4 is a flow chart of a digital human image driving method provided by an exemplary embodiment of the present application;

[0045] Figure 5 This is a schematic diagram of driving a digital human image on a vehicle side provided by an exemplary embodiment of the present application;

[0046] Figure 6 This is a schematic diagram of driving a digital human image in the cloud provided by an exemplary embodiment of the present application;

[0047] Figure 7 This is a structural diagram of a digital human image driving device provided by an exemplary embodiment of the present application;

[0048] Figure 8 It is a structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0049] In the embodiments of the present application, in order to clearly describe the technical solutions of the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity or execution order, and words such as "first" and "second" do not necessarily mean different. There is no order of precedence or priority between the technical features described by "first" and "second".

[0050] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner to facilitate understanding.

[0051] In the embodiments of the present application, at least one can also be described as one or more, and multiple can be two, three, four or more, which is not limited in this application.

[0052] In addition, the network architecture and scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0053] It should be noted that the information (including but not limited to device information, personal information of the subject, etc.), data (including but not limited to data used for analysis, storage, and display, etc.), and signals involved in this application are all authorized by the subject or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the conversation text involved in this application was obtained with full authorization.

[0054] With the rapid development of emerging smart cars, in-car voice assistants have evolved from virtual avatars to humanoid digital humans. When users interact with these digital humans, they can respond to voice commands and perform actions accordingly. When users issue specific voice commands, the digital humans not only execute the commands but also perform actions simultaneously.

[0055] The schematic diagram of triggering the digital human to perform actions in the conventional method is as follows Figure 1 The process of triggering a digital human to perform an action includes: user voice input – local signal processing – local or cloud voice processing – dialogue management – ​​digital human-driven processing – returning dialogue text and converting local text to speech – and outputting the speech to the user.

[0056] Specifically, the user inputs voice input through the in-vehicle voice assistant, triggering the in-vehicle voice assistant to start working. After receiving the voice input, the in-vehicle voice assistant first performs signal processing on the voice input locally (on the vehicle side). Signal processing includes pre-processing operations such as noise reduction and framing. After signal processing, the voice input can be processed locally or in the cloud.

[0057] Among them, cloud-based speech processing or local speech processing includes: automatic speech recognition (ASR) and natural language understanding (NLU).

[0058] When performing local voice processing, the vehicle-side processing engine performs ASR and NLU on the voice input. After understanding the intent and semantics of the voice input, it responds to the voice input and generates a conversational text. Based on the response conversational text, the vehicle-side processing engine converts the conversational text into speech through text-to-speech (TTS) for playback. Furthermore, the vehicle-side processing engine calls the local large model and searches the preset action library for the digital human's image movements, executing the image movements while the digital human speaks.

[0059] For example, a user says into the car's microphone, "Xiao A, I want to listen to songs by celebrity C." The local ASR module receives the user's voice signal and converts it into text: "I want to listen to songs by celebrity C." The local NLU module analyzes this text, understanding the user's intent to "play music" and specifying the artist "celebrity C." It extracts key information: Intent = play music, Artist = celebrity C. Based on the NLU interpretation, the on-board processing engine generates an appropriate conversation text. For example, "Okay, I'm playing songs by celebrity C for you." The local TTS module receives the conversation text "Okay, I'm playing songs by celebrity C for you," synthesizes it into speech, and plays it through the car's speakers. Simultaneously, the on-board processing engine processes the actions of the digital human Xiao A. It performs simple semantic analysis on the conversation text (for example, identifying words like "Okay" and "Play") and then calls a pre-stored "action library" on the car. In the action library, it might find several preset actions based on the word "OK," indicating agreement or confirmation, and the action "Play." For example: Action a: Little A nods slightly, raises the corners of her mouth, and gestures "start" with both hands. Action b: Little A leans forward slightly, her eyes focused, and gestures in the air to select / confirm. At this point, the user hears Little A say in a synthesized voice, "Okay, I'm playing Star C's song for you." Simultaneously, the user sees the digital Little A on the screen nod slightly, smile, and gesture to start playing with both hands.

[0060] When voice processing is performed in the cloud, the cloud-side processing engine processes the voice input, performing ASR and NLU. After understanding the intent and semantics of the voice input, it responds to the voice input and generates a dialogue text. Simultaneously, it uses a large model to drive the digital human's image, inducing movements such as lip movement and body language. Based on the response dialogue text, the cloud-side processing engine converts the dialogue text into speech for playback using text-to-speech transmission (TTS). Furthermore, the cloud-side processing engine searches for the digital human's image movements in a pre-set action library and drives the digital human's lip movement and body language, executing the image movements while the digital human speaks.

[0061] For example, a user says into the car's microphone, "Xiao A, help me find the nearest gas station." The voice signal "Xiao A, help me find the nearest gas station" is transmitted to the cloud-side processing engine via the car's network connection (such as 4G / 5G / Wi-Fi). The ASR module in the cloud-side processing engine receives the voice signal and converts it into text: "Help me find the nearest gas station." The NLU module in the cloud-side processing engine analyzes this text and understands that the user's intent is to "find a location," specifically "a gas station," with the qualifier "nearest." Based on the NLU interpretation, the cloud-side processing engine generates an appropriate conversational text. For example, "Okay, we're looking for a nearby gas station for you. We'll be right there." The cloud-side processing engine sends the conversational text to the car's TTS module. Upon receiving the conversational text "Okay, we're looking for a nearby gas station for you. We'll be right there." The TTS module synthesizes it into speech and plays it through the car's speakers. Simultaneously, the cloud-side processing engine processes the movements of the digital human Xiao A. The cloud-side processing engine uses a large model to analyze the conversation text and generate precise lip-syncing commands, ensuring that the digital human Xiao A's mouth animations are perfectly synchronized with the TTS-synthesized speech. The large model further analyzes the text's semantics and sentiment. It recognizes that "OK" indicates agreement. Based on this analysis, the large model searches for or generates appropriate actions in real time from a larger and more sophisticated cloud-based action library. It may select a nod to indicate "confirmation / agreement." The cloud-side processing engine sends the conversation text and the generated animation data (lip-syncing and body-syncing commands) to the vehicle computer. The vehicle computer's audio player receives and plays the speech converted from the conversation text. The vehicle computer renders the digital human and executes the lip-syncing and body-syncing commands, ensuring that its mouth and body movements are synchronized with the cloud-side processing engine. At this point, the user hears Xiao A say in a synthesized voice, "Okay, I'm looking for a nearby gas station for you. I'll be right there." Simultaneously, the user sees the digital human Xiao A on the screen smiling and nodding, its mouth shape precisely changing with the speech (lip syncing).

[0062] However, in the above technologies, the response actions of the digital human are fixed, that is, according to the trigger instruction, the action corresponding to the trigger instruction is executed, resulting in the image changes of the digital human being relatively simple and programmed, and the matching between the image of the digital human and the content of the digital human's broadcast is poor.

[0063] Based on this, this application provides a method for driving a digital human image. This method obtains the dialogue text that the digital human responds to based on user input. The dialogue text is semantically parsed to obtain image feature data that matches the semantic content of the dialogue text. Based on this image feature data, the digital human is controlled to produce image animations that match the semantic content of the dialogue text. This method semantically parses the response text to ensure that the content expressed by the digital human matches the digital human's image changes, enriching the digital human's image changes and eliminating the need for a fixed number of movements. This enhances the digital human's anthropomorphism and improves the user experience.

[0064] The solutions provided by the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0065] The solution provided in this application can be applied to Figure 2 In the computing system shown in FIG. Figure 2 As shown, the computing system includes a digital human image driving engine. The digital human image driving engine can run in different locations, such as the vehicle side and the cloud.

[0066] The in-vehicle digital human image driving engine runs locally in the vehicle. It can operate even without a network connection, making it suitable for basic vehicle control and situations with poor network conditions. The cloud-based digital human image driving engine runs on a remote server. A network connection is required between the vehicle and the cloud; the cloud-based digital human image driving engine can only operate when connected to the network.

[0067] Optionally, the vehicle can be a new energy vehicle, a gasoline vehicle, a motorcycle, a tricycle, an autonomous driving logistics vehicle, an electric truck, etc., but is not limited thereto, and the embodiments of the present application do not make specific limitations on this.

[0068] The vehicle supports an in-vehicle voice assistant, meaning it can support voice interaction. The vehicle also includes a display screen that displays a digital human that can interact with the user.

[0069] Optionally, the cloud can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides cloud computing services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and other basic cloud computing services.

[0070] Figure 3The embodiment of the present application provides a digital human image driving method, which is applied to a digital human image driving engine. The digital human image driving engine can be Figure 2 The computing system shown in the figure is operated in the vehicle side, and can also be used in Figure 2 Schematic diagram of a computing system running in the cloud.

[0071] like Figure 3 As shown, the digital human image driving method may include:

[0072] Step 302: The digital human image driving engine obtains the dialogue text.

[0073] Among them, dialogue text refers to the text that the digital human responds to based on user input, which can also be called "response text".

[0074] For example, a user interacts with an in-vehicle voice assistant in a vehicle, and the conversation text is the text that responds to the user's input. For example, the user says to the in-vehicle voice assistant, "Xiao A, I want to listen to the songs of celebrity C." The in-vehicle voice assistant responds, "Okay, I'm playing the songs of celebrity C for you." The text "Okay, I'm playing the songs of celebrity C for you" is the conversation text.

[0075] Optionally, the conversation text may be a reply to the user's voice input or a response to the user's instruction.

[0076] The digital human image driving engine is a software system that runs on the vehicle side, cloud or other computing environments. It generates the digital human's movements, expressions and language feedback by parsing the user's instructions or dialogue text, thereby realizing natural interaction between the digital human and the user.

[0077] Digital humans are constructed virtual characters. They not only have anthropomorphic or real-life appearance and form, but can also interact with humans through voice, facial expressions, and body movements. Digital humans can be either 3D or 2D.

[0078] In addition, the style of the digital human is not limited to humans, that is, the digital human can be an animated dog, an animated cat, a plant, etc., but is not limited to this. The embodiments of this application do not make specific limitations on this.

[0079] Step 304: The digital human image driving engine performs semantic analysis on the dialogue text to obtain image feature data that matches the semantic content of the dialogue text.

[0080] Among them, the image feature data is used to indicate the image change characteristics of the digital human.

[0081] Exemplarily, the digital human image driving engine performs semantic parsing on the dialogue text to obtain the semantic content of the dialogue text. For example, whether the dialogue text is happy, surprised, questioning, or complaining. And based on this, a set of image feature data is generated to drive the image of the digital human.

[0082] Optionally, the image feature data includes one or more of the following: mouth feature data, limb movement feature data, or facial expression feature data, but is not limited thereto. The embodiments of the present application do not make specific limitations in this regard.

[0083] Among them, the mouth feature data is the data used to indicate the morphological features of the digital human's mouth. That is, through the mouth feature data, the morphological features of the digital human's mouth can be determined.

[0084] Optionally, the mouth morphology includes: mouth shape morphology, mouth opening degree, mouth corner state, etc., but is not limited thereto. The embodiments of the present application do not make specific limitations in this regard.

[0085] The mouth shape morphology is the core part of the mouth morphology and is mainly used to match the speech. Different phonemes require different mouth shapes. For example: when pronouncing / a / (ah), the mouth is usually opened wider,呈圆形或椭圆形。When pronouncing / i / (yi), the corners of the mouth are stretched to both sides, and the lips may become slightly thinner. When pronouncing / u / (wu), the mouth is closed and protruded forward,呈圆形。When pronouncing / s / (si) or / f / (fo), the lips may be close or separated to form a gap.闭嘴(用于元音之间的过渡或无声部分)。The digital human image driving engine will select or generate the corresponding mouth shape morphology according to the phoneme to be pronounced.

[0086] The mouth opening degree refers to the size of the mouth opening.

[0087] The mouth corner state includes:上扬: Usually indicates positive emotions such as smiling, happy, and proud.下垂: May indicate negative emotions such as sadness, dissatisfaction, and exhaustion.保持水平: Neutral state, or indicates seriousness and neutral statement.不对称: May indicate疑惑,嘲笑, or just the natural distortion when speaking.

[0088] The limb movement feature data is the data used to indicate the limb movement features of the digital human. That is, through the limb movement feature data, the limb movements of the digital human can be determined.

[0089] Optionally, the limbs include the hands, feet, upper body, lower body, etc. of the digital human. The limb morphology refers to the specific posture or movement state presented by the limbs of the digital human. It includes the following main aspects:

[0090] (1)关节角度 / 旋转。

[0091] Joint angles / rotations are the most core data for limb movements. Each movable joint (such as shoulders, elbows, wrists, hips, knees, ankles, etc.) has one or more angle values ​​to define its bending, extension, and rotation states. For example: Arms: the degree of bending of the elbow joint (bend vs. straighten), the lifting angle and forward and backward swing angle of the shoulder joint. Legs: the degree of bending of the knee joint (bend vs. straighten), the lifting, lowering, and forward and backward swing angles of the hip joint. Torso: the bending and twisting angles of the spine (leaning forward, leaning back, twisting left and right). Fingers and wrists: the degree of bending of the fingers (clenching fists vs. opening fists), the rotation angle of the wrists. Body posture / balance: refers to the posture of the entire body relative to the ground, including the position of the center of gravity, the inclination angle of the body, etc. For example: whether to stand straight or slightly leaning forward.

[0092] (2) Gestures / hand movements.

[0093] Gestures / hand movements include: pointing, waving, clenching fists, making gestures (such as indicating "OK", "stop", etc.), picking up or manipulating objects, etc.

[0094] Facial expression feature data is data used to indicate the facial expression features of the digital human. That is, the facial expression of the digital human can be determined through the mouth feature data.

[0095] Optionally, facial expressions typically include: happiness (expression: corners of the mouth raised, eyes squinted), sadness (eyebrows furrowed and downward, corners of the mouth drooping, possibly accompanied by tears), anger (expression: eyebrows down and gathered, eyes widened, lips closed or pursed), fear (eyebrows raised, eyes widened, mouth slightly open), surprise (eyebrows suddenly raised, eyes widened, mouth open), disgust (upper lip raised, nose wrinkled, eyebrows slightly wrinkled), contempt (expression: one side of the mouth corner raised, eyebrows slightly drooping), etc., but are not limited to these, and the embodiments of the present application do not make specific limitations on this.

[0096] For example, consider a conversation text: "Oops! I just spilled water on my keyboard. What should I do?" The digital human image-driven engine parses the conversation text, understanding that the user is likely surprised, a little flustered, and seeking help. Therefore, it generates the following image feature data that matches the semantic content of the conversation text: Facial expression feature data: eyebrows raised (indicating surprise), slightly drooping corners of the mouth (indicating worry or annoyance), and wide-open eyes. Mouth feature data: Generates corresponding lip shape changes based on the conversation text (for example, the lip shape for pronouncing "ah," "ban," "shui," etc.). Body movement feature data: The body leans slightly forward to express concern, one hand may be raised as if simulating wiping, and the head may be slightly tilted to one side to indicate thought or confusion.

[0097] Step 306: The digital human image driving engine controls the digital human to produce an image animation that matches the semantic content of the dialogue text based on the image feature data.

[0098] Among them, image animation is the image change animation after the digital human image driving engine adjusts the digital human image according to the image feature data.

[0099] For example, after the digital human image driving engine obtains the image feature data, the digital human image driving engine drives the digital human image according to the image feature data and renders it, so that the digital human image changes according to the image feature data.

[0100] For example, the acquired image feature data includes: Facial expression data: eyebrows raised (indicating surprise), slightly drooping corners of the mouth (indicating worry or annoyance), and wide-open eyes. Mouth feature data: corresponding lip shape changes are generated based on the conversation text (for example, the mouth shape for pronouncing "ah," "ban," and "shui"). Body movement feature data: the body slightly leans forward to express concern, one hand may be raised as if wiping, and the head may be slightly tilted to one side to indicate thought or confusion. The digital human image driving engine controls the digital human to animate the image based on the image feature data. After rendering, the user sees the following image animation: the digital human's eyebrows raised and eyes wide open. The mouth opens and closes, and changes shape in response to the pronunciation of the sentence "Oops, I just spilled water on the keyboard. What should I do?" The body slightly leans forward, one hand is raised, and the head is slightly tilted to one side.

[0101] To sum up, the solution provided in this application, by performing semantic analysis on the responding text, makes the content expressed by the digital human match the image change items of the digital human, enriches the image change content of the digital human, no longer uses a few fixed actions, improves the degree of anthropomorphism of the digital human, and enhances the user experience.

[0102] Furthermore, if Figure 4 In the flowchart of the digital human image driving method shown, since the image feature data includes mouth feature data, body movement feature data, or facial expression feature data, step 304 can be implemented as steps 3041, 3042, and 3043 for different image feature data. Step 306 can be implemented as steps 3061, 3062, and 3063.

[0103] When the image feature data is mouth feature data, step 304 and step 306 are implemented as: step 3041 and step 3061.

[0104] In the case where the image feature data is body movement feature data, step 304 and step 306 are implemented as: step 3042 and step 3062.

[0105] In the case where the image feature data is facial expression feature data, step 304 and step 306 are implemented as: step 3043 and step 3063.

[0106] It should be noted that steps 3041, 3042, and 3043 can be combined, and steps 3061, 3062, and 3063 can also be combined, that is, different image feature data can be centrally reflected in a single digital human. Specifically, when the image feature data includes at least two of mouth feature data, body movement feature data, or facial expression feature data, the at least two types of feature data are synchronized based on the timestamp. Based on the at least two types of synchronized feature data, the digital human is controlled to produce an image animation that matches the semantic content of the conversation text.

[0107] In a possible implementation, when the image feature data is mouth feature data, step 304 and step 306 are implemented as: step 3041 and step 3061.

[0108] Step 3041: Perform semantic analysis on the dialogue text to obtain mouth feature data that matches the semantic content of the dialogue text.

[0109] For example, the text-to-speech engine within the digital human avatar driver engine converts the conversation text into speech, generating voice information that matches the digital human character settings. The text-to-speech engine decomposes the text in the conversation text into phonemes, obtaining the phonemes corresponding to each word, and records the start and end times of each phoneme in the voice information. Based on the start and end times of each phoneme in the voice information, the voice information and the phoneme corresponding to each word are time-aligned to obtain mouth feature data.

[0110] The text-to-speech engine is a TTS engine, which is used to convert text into sound.

[0111] Voice messages that match the digital human's character settings are generated by adjusting the timbre, speaking speed, pitch, and tone of the voice based on the digital human's character settings. For example, if Xiao A is set as a young and lively male, the TTS engine will generate a voice message with a youthful, energetic voice and a slightly faster speaking speed based on the conversation text. If Xiao A is set as a mature and stable female, the TTS engine will generate a voice message with a slower, deeper, and gentler voice based on the conversation text.

[0112] Phonemes refer to the pronunciation units corresponding to words;

[0113] Specifically, while generating speech information, the TTS engine breaks down the conversation text into its most basic pronunciation units: phonemes. (For example, "a," "o," "e," "b," and "p" in Chinese.) It then precisely records the time at which each phoneme begins and ends in the generated speech information. This is like adding a time stamp to each "syllable fragment" of speech. Based on the speech information and phonemes, the TTS engine generates mouth feature data.

[0114] For example, if the conversation is: "Hello, what a nice day!", the TTS engine will break the conversation down into its most basic pronunciation units: the phonemes: ni, hao, jin, tian, tian, qi, zhen, hao, and accurately record the start and end times of each phoneme in the audio file. For example, "ni" starts at 0.1 seconds and ends at 0.3 seconds. "hao" starts at 0.3 seconds and ends at 0.6 seconds. "jin" starts at 0.6 seconds and ends at 0.9 seconds, and so on.

[0115] Step 3061: Based on the mouth feature data, control the digital human to emit language information, and control the digital human's mouth to make lip movements corresponding to the phonemes corresponding to each word.

[0116] For example, based on the mouth feature data, the digital human is controlled to emit language information, and the digital human's mouth is controlled to make lip movements corresponding to the phonemes corresponding to each word.

[0117] Specifically, after obtaining the voice information and phonemes from the mouth feature data and aligning the phonemes and voice information, the TTS engine can use this information to control the digital human's mouth movements. For example, when the TTS engine announces the sound "a," the digital human's mouth will assume the shape of an "a" sound (e.g., opening its mouth wide); when an "n" sound is announced, the digital human's mouth will assume the shape of an "n" sound (e.g., raising the corners of the mouth and pressing the tip of the tongue against the upper gum). This process is performed in real time, ensuring that the digital human's lip movements perfectly match the sounds it produces.

[0118] For example, consider the following conversation: "Hello, what a nice day!" The TTS engine uses phonemes to control the digital human's mouth, making the corresponding lip movements for each phoneme in the text. For example, at 0.1 seconds into the speech, the TTS engine determines based on alignment that the first phoneme "n" in "ni" should be pronounced. It then moves Little A's mouth into the shape of an "n" sound (the corners of the mouth may be slightly drawn inward, with the tip of the tongue close to the upper gum). At 0.3 seconds, the mouth shape switches to that of an "i" sound (flattened mouth, open corners). At 0.3 seconds into the sound "hao," the mouth changes to that of an "h" sound (slightly parted lips, air flowing out). This then switches to the shape of an "a" sound (mouth wide open). Then, it switches to the shape of an "o" sound (mouth rounded). This process continues until the "o" phoneme in "hao" is finished, with Little A's mouth shape changing smoothly and in real time to meet the pronunciation requirements of each phoneme.

[0119] In another possible implementation, when the image feature data is mouth feature data, step 304 and step 306 are implemented as: step 3042 and step 3062.

[0120] Step 3042: Perform semantic analysis on the dialogue text to obtain body movement feature data that matches the semantic content of the dialogue text.

[0121] For example, the behavior inference engine within the digital human avatar driver engine breaks down the text in the dialogue text into text segments. Based on the text segments, the behavior tree is selected to obtain body movement feature data that matches the semantic content of the dialogue text.

[0122] The behavior tree includes different action information corresponding to different text segments. The behavior inference engine is used to understand and infer the semantic content of the conversation text, that is, to infer the actual meaning of the text in the conversation text, including the emotions, intentions, and specific content.

[0123] For example, consider the conversation text: "What a nice day today! Let's go have a picnic in the park!" The behavior inference engine breaks down "What a nice day today! Let's go have a picnic in the park!" into the following text segments: "What a nice day (positive emotion)" and "Picnic in the park (activity proposal)." The text segment "What a nice day" indicates a positive emotion, so the behavior tree can be used to identify these: gestures of appreciation or approval, such as a gentle nod or an upward gesture, as if pointing toward the clear sky. The text segment "Picnic in the park" indicates a specific example, so the behavior tree can be used to identify more specific and imaginative gestures, such as holding something with both hands or pointing a finger off into the distance to indicate "go over there." The action data obtained from the behavior tree is aggregated to generate body movement feature data.

[0124] Step 3062: Based on the body movement feature data, control the digital human's body to make body movements corresponding to the text segment.

[0125] For example, after analyzing the text in a conversation, the behavioral inference engine will plan a series of appropriate "hand gestures" and "leg movements" based on the analysis results. These movements need to be associated with the analyzed text segment, for example, expressing happiness, confusion, emphasizing a point, or describing a specific action. The planned movements will instruct the digital human to perform the corresponding body movements.

[0126] For example, the dialogue text is: "The weather is so nice today, let's go to the park for a picnic!" The behavioral inference engine will control the digital human's limbs to make corresponding body movements based on the body movement feature data. The picture the user sees is: the digital human nods slightly, points his finger at the clear sky, and then makes the gesture of "holding" something with both hands.

[0127] In another possible implementation, when the image feature data is mouth feature data, step 304 and step 306 are implemented as: step 3043 and step 3063.

[0128] Step 3043: Perform semantic analysis on the dialogue text to obtain facial expression feature data that matches the semantic content of the dialogue text.

[0129] For example, the emotional text in the conversation text is filtered out by the emotion analysis engine, and facial expression feature data is obtained based on the type and emotional level of the emotional text.

[0130] Among them, emotional text is text used to express emotions, and emotional level is used to indicate the degree of emotional expression.

[0131] Optionally, the facial expression feature data includes: facial expression and the magnitude of the facial expression. The emotion analysis engine determines the facial expression corresponding to the emotion text based on the type of emotion text. Based on the emotional level of the emotion text, the magnitude of the facial expression is obtained. Based on the facial expression and the magnitude of the facial expression, the facial expression feature data is obtained.

[0132] For example, consider the following conversation: "The new restaurant we went to today was amazing! The food was delicious, the service was great, and I was so happy!" The sentiment analysis engine filters out the sentiment text in the conversation: "great," "delicious," "good," and "super happy." Based on these sentiment texts, it can be determined that the emotion is very happy, indicating a high level of emotion. Therefore, based on the type and level of sentiment text, the sentiment analysis engine can derive facial expressions such as "a big smile, eyes possibly narrowed (like a smile), eyebrows slightly raised," and the range of expression can include "the corners of the mouth raised high."

[0133] Step 3063: Based on the facial expression feature data, control the digital human's face to make a facial expression corresponding to the emotional text.

[0134] For example, after analyzing the emotional text in the conversation text, the emotion analysis engine will plan appropriate facial expressions based on the analysis results. These facial expressions need to be associated with the emotional text just analyzed.

[0135] For example, the conversation text is: "The new restaurant we went to today was amazing! The food was delicious, the service was great, I was so happy!" The sentiment analysis engine will control the digital human's body movements to make corresponding expressions based on the facial expressions and the range of expression. The image the user sees is: the digital human's mouth corners are raised high, showing a big smile, the eyes may be narrowed (like smiling eyes), and the eyebrows are slightly raised.

[0136] The above embodiment describes the digital human image driving method. The following will describe the digital human image driving method with specific examples.

[0137] For example, a user inputs a voice command through the in-vehicle voice assistant, triggering the in-vehicle voice assistant to start working. After receiving the voice command, the in-vehicle voice assistant first processes the voice command locally (on the vehicle side). This processing includes pre-processing operations such as noise reduction and framing. The processed voice command can then be processed locally or in the cloud.

[0138] When the digital human image driving engine is deployed on the vehicle side, Figure 5The diagram shows how the digital human image is driven on the vehicle's computer. The digital human image driving engine in the vehicle performs local speech processing (ASR and NLU) on the signal-processed speech input. After understanding the intent and semantics of the speech input, it responds to the speech input and generates a conversational text. The digital human image driving engine in the vehicle performs semantic analysis on the conversational text, obtaining image feature data that matches the semantic content of the conversational text. Based on this image feature data, the animation engine generates an animation corresponding to the digital human. The rendering engine then renders the animation to produce a video file, which is then presented to the user.

[0139] When the digital human image driving engine is deployed in the cloud, Figure 6 The diagram shows a digital human avatar driven in the cloud. The digital human avatar driving engine in the cloud performs cloud-based voice processing (i.e., voice recognition) on the signal-processed voice input. After understanding the intent and semantics of the voice input, it responds to the voice input and generates a conversational text. The digital human avatar driving engine in the cloud performs semantic understanding on the conversational text to determine the context of the conversational text, for example, determining that the user is chatting with the in-vehicle voice assistant. The digital human avatar driving engine in the cloud converts the conversational text into speech using a text-to-speech engine, generating a voice message that matches the digital human character setting. Simultaneously, the text-to-speech engine decomposes the text in the conversational text into phonemes, obtaining the phonemes corresponding to each word and recording the start and end times of each phoneme in the voice message. Based on the start and end times of each phoneme in the voice message, the voice message and the phoneme corresponding to each word are time-aligned to obtain mouth feature data. Furthermore, the digital human avatar driving engine in the cloud decomposes the text in the conversational text into text segments using a behavioral inference engine. Based on the text segments, selections are made within the behavior tree to obtain body movement feature data that matches the semantic content of the conversation text. Furthermore, the cloud-based digital human image driver engine uses the emotion analysis engine to filter out emotional text within the conversation text. Based on the type and emotional level of the emotional text, facial expression feature data is obtained. The cloud-based digital human image driver engine stores the mouth feature data, body movement feature data, and facial expression feature data in the asset library. The animation engine then generates the corresponding animation for the digital human based on the mouth feature data, body movement feature data, and facial expression feature data in the asset library. This animation is then rendered by the rendering engine to produce a video file of the rendered digital human, and the visual animation of the video file is presented to the user.

[0140] For example, a user chats with Xiao A through the car microphone. The user says: "What a nice day today. What are you going to do?" Xiao A: "What a nice day today. Let's go for a walk in the park!" The conversation text is: "What a nice day today. Let's go for a walk in the park!"

[0141] The digital human image driving engine calculates all the phoneme sequences and corresponding precise mouth shape change time points required to say "The weather is so nice today, let's go for a walk in the park!" through the text-to-speech engine, and obtains the mouth feature data.

[0142] The behavioral inference engine analyzes that the content is about "good weather" and "going for a walk in the park" and plans possible gestures, such as pointing out the window (if the interface allows) or making a gesture indicating "departure" or "enjoyment" (such as opening both hands to simulate embracing the sun), to obtain body movement feature data.

[0143] The sentiment analysis engine analyzes that the text is full of positive and happy emotions ("That's great", "Let's go for a walk in the park"), and plans out happy expressions (smiling, curved eyes) and the possible accompanying relaxed and happy body posture (the body is slightly leaning forward, appearing energetic), and obtains facial expression feature data.

[0144] After receiving one or more of the following data: mouth feature data, body movement feature data, or facial expression feature data, the animation engine begins to drive the digital human. For example, the mouth needs to accurately pronounce each sound according to the mouth feature data, making the corresponding mouth shape. When saying "That's great," the expression needs to switch to a happy smile, and the eyes should light up. When saying "Let's go for a walk in the park," you can accompany it with a relaxed gesture, such as raising one hand, slightly bending the fingers, and then extending them outward, simulating an invitation or walking towards the distance. At the same time, the body posture should remain relaxed and cheerful.

[0145] The final comprehensive performance of the digital human Xiao A is as follows: his mouth opens and closes naturally in accordance with the content of the speech, clearly uttering "The weather is so nice today, let's go for a walk in the park!" From the beginning of the speech, a happy smile gradually appears on his face, and his eyes appear bright. When he says the second half of the sentence, Xiao A may make a relaxed, inviting gesture. Throughout the process, Xiao A's posture is relaxed and positive, and his eyes are full of joy.

[0146] The above mainly introduces the solution provided by this application. Correspondingly, this application also provides a digital human image driving device, which is used to implement the above method embodiment.

[0147] In some embodiments, in order to realize the above functions, the digital human image driving device includes hardware structures and / or software modules corresponding to the execution of each function. It should be easily appreciated by those skilled in the art that, in combination with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0148] The embodiments of the present application can divide the digital human image driving device into functional modules based on the above-mentioned method embodiments. For example, each functional module can be divided into corresponding functional modules, or two or more functions can be integrated into a single processing module. The above-mentioned integrated modules can be implemented in the form of hardware or software functional modules. It should be noted that the module division in the embodiments of the present application is illustrative and only represents a logical functional division. In actual implementation, other division methods may be used.

[0149] In some embodiments, the present application provides a digital human image driving device, which is used to implement the functions of the digital human image driving device in the above-mentioned digital human image driving method embodiment. Figure 7 The digital human image driving device shown in the figure may include an acquisition module 701 , an analysis module 702 and a driving module 703 .

[0150] Among them, the acquisition module 701 is used to execute Figure 3 、 Figure 4 The operation of step 302 in the illustrated method. The parsing module 702 is used to perform Figure 3 、 Figure 4 The operation of steps 304, 3041, 3042, and 3043 in the illustrated method. The driving module 703 is used to execute Figure 3 、 Figure 4 The operations of steps 306, 3061, 3062, and 3063 in the illustrated method.

[0151] like Figure 8 As shown, the computer device provided in the embodiment of the present application may include a processor 801, a bus 802, a communication interface 803, and a memory 804. The processor 801, the memory 804, and the communication interface 803 communicate with each other via the bus 802. It should be understood that the present application does not limit the number of processors and memories in the computer device.

[0152] The bus 802 may be a PCI bus or an extended industry standard architecture (EISA) bus, or a USB bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 The bus 802 may include a path for transmitting information between various components of a computer device (eg, memory 804, processor 801, communication interface 803).

[0153] The processor 801 may include any one or more processors such as a CPU, a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0154] The memory 804 may include a volatile memory, such as a random access memory (RAM). The processor 801 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0155] The communication interface 803 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computer device and other devices or a communication network.

[0156] The memory 804 stores executable program code, and the processor 801 executes the executable program code to implement the functions of the digital human image driving device or the CPU core in the above-mentioned method embodiment. In other words, the memory 804 stores the program code for executing the above-mentioned digital human image driving method.

[0157] On the other hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the digital human image driving method provided in the above-mentioned method embodiments.

[0158] On the other hand, a computer program product is provided. The computer program product includes a computer program or instructions. When the computer program or instructions are executed by a processor, the digital human image driving method as described above is implemented.

[0159] On the other hand, a chip system is provided, comprising at least one processor and at least one interface circuit, wherein the at least one interface circuit is used to perform transceiver functions and send instructions to the at least one processor. When the at least one processor executes the instructions, the at least one processor executes to implement the digital human image driving method as described above.

[0160] The method steps in this embodiment can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC. In addition, the ASIC can be located in a computer device. Of course, the processor and storage medium can also exist in a computer device as discrete components.

[0161] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application is performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device or other programmable device. The computer program or instruction can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instruction can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD). The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A digital human image driving method, characterized in that: The method is applied to a digital human image driving engine; the method comprises: Acquiring a dialogue text, wherein the dialogue text refers to the text in which the digital human responds based on user input; Performing semantic analysis on the dialogue text to obtain image feature data that matches the semantic content of the dialogue text, wherein the image feature data is used to indicate image change characteristics of the digital human; Based on the image feature data, the digital human is controlled to produce an image animation that matches the semantic content of the dialogue text.

2. The method according to claim 1, characterized in that The image feature data includes one or more of mouth feature data, body movement feature data or facial expression feature data.

3. The method according to claim 2, characterized in that The digital human image driving engine includes: a text-to-speech engine; The semantic analysis of the dialogue text to obtain image feature data that matches the semantic content of the dialogue text includes: Converting the dialogue text into speech using the text-to-speech engine to obtain speech information that matches the character setting of the digital human; Decomposing the characters in the conversation text into phonemes by the text-to-speech engine to obtain phonemes corresponding to each character, and recording the start time and end time of the phonemes corresponding to each character in the voice information, wherein the phonemes refer to the pronunciation units corresponding to the characters; According to the start time and end time of the phonemes corresponding to the respective characters in the voice information, the voice information and the phonemes corresponding to the respective characters are time-aligned to obtain the mouth feature data.

4. The method according to claim 3, characterized in that The step of controlling the digital human to produce an image animation that matches the semantic content of the dialogue text based on the image feature data includes: Based on the mouth feature data, the digital human is controlled to emit the language information, and the mouth of the digital human is controlled to make lip movements corresponding to the phonemes corresponding to the various characters.

5. The method according to claim 2, characterized in that The digital human image driving engine includes: a behavior reasoning engine; The semantic analysis of the dialogue text to obtain image feature data that matches the semantic content of the dialogue text includes: Decomposing the text in the conversation text by the behavior inference engine to obtain text segments; According to the text segment, selection is made in a behavior tree to obtain the body movement feature data that matches the semantic content of the dialogue text, wherein the behavior tree includes different action information corresponding to different text segments.

6. The method according to claim 5, characterized in that The step of controlling the digital human to produce an image animation that matches the semantic content of the dialogue text based on the image feature data includes: Based on the limb movement feature data, the limbs of the digital human are controlled to perform limb movements corresponding to the text segment.

7. The method according to claim 2, characterized in that The digital human image driving engine includes: an emotion analysis engine; The semantic analysis of the dialogue text to obtain image feature data that matches the semantic content of the dialogue text includes: Filtering out emotional text from the conversation text by the emotion analysis engine, wherein the emotional text is text used to express emotions; The facial expression feature data is obtained according to the type and emotion level of the emotional text.

8. The method according to claim 7, characterized in that The facial expression feature data is obtained according to the type and emotion level of the emotional text, including: Determining, according to the type of the emotional text, a facial expression corresponding to the emotional text; Obtaining the magnitude of the facial expression according to the emotional level of the emotional text; The facial expression feature data is obtained based on the facial expression and the expression amplitude.

9. The method according to claim 8, characterized in that The step of controlling the digital human to produce an image animation that matches the semantic content of the dialogue text based on the image feature data includes: Based on the facial expression feature data, the face of the digital human is controlled to make a facial expression corresponding to the emotional text.

10. The method according to any one of claims 1 to 9, characterized in that The digital human image driving engine is deployed on the vehicle side and / or the cloud.

11. The method according to any one of claims 2 to 9, characterized in that The method further comprises: When the image feature data includes at least two of the mouth feature data, the body movement feature data, or the facial expression feature data, synchronizing the at least two feature data according to a timestamp; According to the at least two kinds of feature data after data synchronization, the digital human is controlled to produce an image animation that matches the semantic content of the dialogue text.

12. A digital human image driving device, characterized in that: The device comprises: An acquisition module, configured to acquire a dialogue text, wherein the dialogue text refers to the text in which the digital human responds based on the user input; An analysis module, configured to perform semantic analysis on the dialogue text to obtain image feature data that matches the semantic content of the dialogue text, wherein the image feature data is used to indicate image change characteristics of the digital human; The driving module is used to control the digital human to produce an image animation that matches the semantic content of the dialogue text based on the image feature data.

13. A vehicle, characterized in that: The vehicle comprises a display screen on which a digital human is displayed, and the vehicle drives the digital human based on the digital human image driving method according to any one of claims 1 to 11.

14. A computer device, characterized in that: The computer device includes: a processor and a memory, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the digital human image driving method according to any one of claims 1 to 11.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the digital human image driving method according to any one of claims 1 to 11.

16. A computer program product, characterized in that The computer program product includes a computer program or instructions, and when the computer program or instructions are executed by a processor, the digital human image driving method according to any one of claims 1 to 11 is implemented.