Digital human interaction control method and device, electronic equipment and storage medium

By acquiring users' voice commands and facial expressions, determining emotional characteristics, and using a pre-set library of animation clips and speech recognition results to generate response animations that match the user's emotions, the problem of latency and inconsistency in digital human interaction is solved, improving interaction efficiency and user experience.

CN116524923BActive Publication Date: 2026-04-28AVATAR WORKS INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AVATAR WORKS INC
Filing Date
2023-04-21
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as delayed interaction, inconsistency between actions and statements, and monotonous actions when digital humans interact with users, resulting in low interaction efficiency and affecting user experience.

Method used

By acquiring the user's voice commands and facial expressions, the system determines emotional characteristics, generates response animations using a pre-set animation clip library and speech recognition results, and adjusts the animation clips based on the emotional characteristics to generate response animations that match the user's emotions.

Benefits of technology

It enables efficient interaction with digital humans, enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116524923B_ABST
    Figure CN116524923B_ABST
Patent Text Reader

Abstract

The application discloses a digital person interactive control method and device, electronic equipment and storage medium, the method comprises the following steps: obtaining a target digital person for interaction; determining the emotional characteristics of the user according to the voice instruction issued by the user and the expression image of the user; obtaining a response text according to the voice recognition result of the voice instruction, and obtaining a response animation segment corresponding to the response text from a preset animation segment library, wherein the preset animation segment library comprises a plurality of preset animation segments corresponding to the target digital person; adjusting the response animation segment based on the emotional characteristics to obtain a target animation segment, and generating a response animation corresponding to the target digital person based on the target animation segment, thereby adjusting the response animation segment based on the emotional characteristics of the user, so that the digital person exhibits a response animation that meets the emotional characteristics of the user, and more efficient interaction with the digital person is realized, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to a method, apparatus, electronic device, and storage medium for interactive control of a digital human. Background Technology

[0002] With the continuous development of artificial intelligence, digital human interaction has begun to be applied in various fields to achieve intelligent human-computer interaction. In existing technologies, when interacting with digital humans, there are often problems such as delays in the connection between language interaction and body movements, inconsistencies between actions and expressions, and monotonous actions, resulting in low interaction efficiency and affecting user experience.

[0003] Therefore, how to interact with digital humans more efficiently and improve user experience is a technical problem that needs to be solved.

[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] This application provides an interactive control method, device, electronic device, and storage medium for digital humans, enabling more efficient interaction with digital humans and improving user experience.

[0006] In a first aspect, a method for interactive control of a digital human is provided. The method includes: acquiring a target digital human for interaction; determining the emotional characteristics of a user based on a voice command issued by the user and an image of the user's facial expression; acquiring a response text based on a speech recognition result of the voice command, and acquiring a response animation clip corresponding to the response text from a preset animation clip library, wherein the preset animation clip library includes multiple preset animation clips corresponding to the target digital human; adjusting the response animation clip based on the emotional characteristics to obtain a target animation clip, and generating a response animation corresponding to the target digital human based on the target animation clip.

[0007] Secondly, an interactive control device for a digital human is provided, the device comprising: a first acquisition module for acquiring a target digital human for interaction; a determination module for determining the emotional characteristics of the user based on a voice command issued by the user and an image of the user's facial expression; a second acquisition module for acquiring a response text based on a speech recognition result of the voice command, and acquiring a response animation segment corresponding to the response text from a preset animation segment library, wherein the preset animation segment library includes multiple preset animation segments corresponding to the target digital human; and a generation module for adjusting the response animation segment based on the emotional characteristics to obtain a target animation segment, and generating a response animation corresponding to the target digital human based on the target animation segment.

[0008] Thirdly, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the interactive control method for a digital human as described in the first aspect by executing the executable instructions.

[0009] Fourthly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the interactive control method for the digital human described in the first aspect.

[0010] By applying the above technical solutions, a target digital human for interaction is obtained; the user's emotional characteristics are determined based on the user's voice commands and facial expressions; response text is obtained based on the speech recognition results of the voice commands, and a corresponding response animation clip is obtained from a preset animation clip library, wherein the preset animation clip library includes multiple preset animation clips corresponding to the target digital human; the response animation clip is adjusted based on the emotional characteristics to obtain a target animation clip, and a response animation corresponding to the target digital human is generated based on the target animation clip. This adjustment of the response animation clip based on the user's emotional characteristics allows the digital human to display response animations that match the user's emotions, achieving more efficient interaction with the digital human and improving the user experience. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating an interactive control method for a digital human according to an embodiment of the present invention is shown.

[0013] Figure 2 A flowchart illustrating an interactive control method for a digital human according to another embodiment of the present invention is shown;

[0014] Figure 3 A flowchart illustrating an interactive control method for a digital human according to another embodiment of the present invention is shown;

[0015] Figure 4 This invention provides a schematic diagram of the structure of an interactive control device for a digital human according to an embodiment of the present invention.

[0016] Figure 5 A schematic diagram of the structure of an electronic device according to an embodiment of the present invention is shown. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0018] It should be noted that other embodiments of this application will readily conceive of by those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this application are indicated in the claims section.

[0019] It should be understood that this application is not limited to the precise structure described below and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

[0020] This application can be used in a wide variety of general-purpose or special-purpose computing environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor devices, distributed computing environments including any of the above devices, etc.

[0021] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0022] The following is combined Figures 1-3 This application describes an interactive control method for digital humans according to exemplary embodiments thereof. It should be noted that the following application scenarios are shown only to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way. Rather, the embodiments of this application can be applied to any applicable scenario.

[0023] This application provides an interactive control method for a digital human, such as... Figure 1 As shown, the method includes the following steps:

[0024] Step S101: Obtain the target digital human for interaction.

[0025] The target digital human can be selected by the user from multiple preset digital humans, or it can be created or uploaded by the user, or it can be assigned to the user according to preset allocation rules, such as assigning a digital human that matches the user's personality traits or user profile data.

[0026] Step S102: Determine the user's emotional characteristics based on the user's voice commands and facial expression images.

[0027] When interaction with the target digital human is needed, the user issues corresponding voice commands. These commands are captured via a pre-set microphone and can include, for example, asking a question, requesting music to be played, or instructing the digital human to perform a certain action. The voice commands can be standard commands conforming to preset command specifications or arbitrary commands issued by the user. Simultaneously, a pre-set camera captures the user's facial image, obtaining their expression. Based on the voice commands and the user's facial expressions, the system determines the user's emotional characteristics. For example, emotional characteristics could include calmness, happiness, depression, anger, sadness, etc.

[0028] Step S103: Obtain the response text based on the speech recognition result of the voice command, and obtain the response animation clip corresponding to the response text from the preset animation clip library.

[0029] In this embodiment, a preset speech recognition algorithm is used to perform speech recognition on the speech command to obtain the speech recognition result. The response text is then obtained based on the speech recognition result. For example, if the speech command is a question, the response text is the answer to that question; or, if the speech command requests the execution of a target action, the response text is a descriptive text corresponding to that target action. A preset animation clip library includes multiple preset animation clips corresponding to the aforementioned target digital human. For each preset animation clip, there is a corresponding response text. The response animation clip is obtained by querying the preset animation clip library based on the response text.

[0030] The preset speech recognition algorithm can be any of the following: Dynamic Time Warping (DTW) algorithm, Vector Quantization (VQ) method based on nonparametric model, Hidden Markov Model (HMM) method based on parametric model, Artificial Neural Network (ANN) and Support Vector Machine.

[0031] In some embodiments of this application, obtaining the response text based on the speech recognition result of the voice command includes:

[0032] The instruction text is obtained based on the speech recognition results;

[0033] If a preset response text corresponding to the instruction text exists in the preset corpus, the preset response text shall be used as the response text.

[0034] If the preset response text is not found in the preset corpus, the instruction text is processed based on the preset language processing model to obtain the response text.

[0035] In this embodiment, the preset corpus includes multiple preset response texts, each corresponding to one or more preset instruction texts. After speech recognition of the voice instruction, the instruction text is obtained based on the speech recognition result. Then, it is determined whether a preset response text corresponding to the instruction text exists in the preset corpus. If it exists, the corresponding preset response text is used as the response text; otherwise, the instruction text is processed based on a preset language processing model to obtain the response text. The preset language processing model is pre-trained using multiple instruction texts and their corresponding response texts. This method of determining the response text based on the judgment result of the preset corpus further improves the accuracy of the response text.

[0036] In some embodiments of this application, each of the preset animation clips is configured with a text tag, and the step of retrieving the response animation clip corresponding to the response text from the preset animation clip library includes:

[0037] The response text is compared with each of the text tags;

[0038] If any of the text tags contains a matching text tag that matches the response text, the preset animation clip corresponding to the matching text tag will be used as the response animation clip.

[0039] If no matching text tag is found among the text tags, a preset default animation clip will be used as the response animation clip.

[0040] In this embodiment, each preset animation segment is equipped with a text tag. The response text is compared with each text tag to determine whether there is a matching text tag that matches the response text. If there is, the preset animation segment corresponding to the matching text tag is used as the response animation segment; otherwise, the preset default animation segment is used as the response animation segment, thereby determining the response animation segment more efficiently. For example, the preset default animation segment could be to make the target digital human show an "sorry" expression and perform a lip-syncing animation of "I'm sorry, I can't answer this question."

[0041] Step S104: Adjust the response animation clip based on the emotional characteristics to obtain a target animation clip, and generate a response animation corresponding to the target digital human based on the target animation clip.

[0042] After acquiring emotional features and response animation clips, the response animation clips are adjusted based on the emotional features to obtain target animation clips that match those features. Then, response animations corresponding to the target digital human are generated based on these target animation clips. For example, the target digital human can be driven by the skeletal and facial expression parameters of the target animation clips to generate corresponding response animations. Thus, response animations that correspond to voice commands and match the user's emotional features are generated based on the target digital human, enabling interaction with the user.

[0043] By applying the above technical solutions, a target digital human for interaction is obtained; the user's emotional characteristics are determined based on the user's voice commands and facial expressions; response text is obtained based on the speech recognition results of the voice commands, and a corresponding response animation clip is obtained from a preset animation clip library, wherein the preset animation clip library includes multiple preset animation clips corresponding to the target digital human; the response animation clip is adjusted based on the emotional characteristics to obtain a target animation clip, and a response animation corresponding to the target digital human is generated based on the target animation clip. This adjustment of the response animation clip based on the user's emotional characteristics allows the digital human to display response animations that match the user's emotions, achieving more efficient interaction with the digital human and improving the user experience.

[0044] This application also proposes an interactive control method for digital humans, such as... Figure 2 As shown, it includes the following steps:

[0045] Step S201: Obtain the target digital human for interaction.

[0046] The target digital human can be selected by the user from multiple preset digital humans, or it can be created or uploaded by the user, or it can be assigned to the user according to preset allocation rules, such as assigning a digital human that matches the user's personality traits or user profile data.

[0047] Step S202: Determine the intensity and pronunciation characteristics of the user's voice based on the voice command.

[0048] Voice commands exhibit specific sound intensity and pronunciation characteristics, allowing us to determine the intensity and pronunciation features of the user's voice. Furthermore, since voice commands may contain environmental noise and background sounds, to improve accuracy, after acquiring the voice command, these noises are identified and removed, leaving only the user's voice in the command and making it more consistent with the user's pronunciation characteristics.

[0049] In some embodiments of this application, determining the intensity and pronunciation features of the user's voice based on the voice command includes: performing spectral analysis on the voice command to obtain spectral intensity, and using the spectral intensity as the intensity feature; identifying the pitch period frequency of the voice command, and determining the pronunciation feature based on the comparison result of the pitch period frequency and a preset period threshold.

[0050] In this embodiment, spectrum analysis processing involves acquiring voice commands in the time domain and then performing a Fourier transform to convert them into frequency domain signals. Pronunciation features include vibrato and steady-state sounds. If the pitch period frequency is less than a preset period threshold, the pronunciation feature is vibrato; otherwise, the pronunciation feature is steady-state sounds. In this way, intensity features and pronunciation features are determined based on spectrum analysis processing and pitch period frequency, improving the accuracy of the user's voice intensity and pronunciation features.

[0051] Step S203: Generate a first emotion recognition result based on the intensity feature, the pronunciation feature, and the preset feature index range.

[0052] In this embodiment, the preset feature index interval can characterize the range of emotion feature indicators corresponding to intensity features and pronunciation features. The first emotion recognition result can be generated based on the intensity features, pronunciation features and the preset feature index interval.

[0053] In some embodiments of this application, generating a first emotion recognition result based on the intensity feature, pronunciation feature, and preset feature index range includes:

[0054] The intensity feature and the pronunciation feature are normalized respectively to obtain the intensity normalization value and the pronunciation normalization value;

[0055] The preset feature index interval is divided based on the preset segmentation value to obtain multiple feature index sub-intervals;

[0056] The first emotion recognition result is generated based on the degree of overlap between the intensity normalization value and the pronunciation normalization value and multiple feature index sub-intervals.

[0057] In this embodiment, the intensity features and pronunciation features are first normalized to obtain normalized intensity values ​​and normalized pronunciation values. Based on a preset segmentation value, the preset feature index interval is divided into multiple feature index sub-intervals. Then, the first emotion recognition result is obtained based on the degree of overlap between the normalized intensity values ​​and pronunciation values ​​and the multiple feature index sub-intervals, thereby further improving the accuracy of the first emotion recognition result. For example, the feature index interval can be [-1, 1], the preset segmentation value is 0.5, and the feature index sub-intervals are [-1, 0.5] and [0.5, 1]. The first emotion recognition result is obtained by judging the degree of overlap between the normalized intensity values ​​and pronunciation values ​​and the multiple feature index sub-intervals. For example, if the normalized intensity value is in the feature index sub-interval [-1, 0.5] and the normalized pronunciation value is in the feature index sub-interval [-1, 0.5], then the corresponding emotion recognition result is determined to be happy. Among them, different emotion recognition results have corresponding feature index sub-intervals, for example, happy: [[0,1],[-0.5,0.5]], sad: [[-1,0],[-1,0]], afraid: [[-1,1],[-1,0]], angry: [[0,1],[-0.5,0.5]].

[0058] Step S204: Perform emotion recognition on the expression image based on the preset expression recognition model to obtain a second emotion recognition result.

[0059] A preset facial expression recognition model is pre-established for emotion recognition. The facial expression image is input into the preset facial expression recognition model to perform emotion recognition and obtain a second emotion recognition result.

[0060] Step S205: Determine the emotion feature based on the first emotion recognition result and the second emotion recognition result.

[0061] The first emotion recognition result represents the emotion feature corresponding to the voice command, and the second emotion recognition result represents the emotion feature corresponding to the facial expression image. By fusing the two, the user's emotion feature can be determined.

[0062] In some embodiments of this application, determining the emotional features based on the first emotion recognition result and the second emotion recognition result includes:

[0063] Identify the same facial expression labels in the first emotion recognition result and the second emotion recognition result, and use the same facial expression labels as target labels;

[0064] The first number of target labels in the first emotion recognition result is multiplied by the first preset weight to obtain the first expression value;

[0065] The second number of target labels in the second emotion recognition result is multiplied by the second preset weight to obtain the second expression value;

[0066] The first expression value and the second expression value are summed to obtain the final expression value. The maximum final expression value among the multiple final expression values ​​is determined, and the emotional feature is determined according to the expression tag corresponding to the maximum final expression value.

[0067] In this embodiment, the first emotion recognition result and the second emotion recognition result each include multiple emotion tags. First, the same expression tags in the first emotion recognition result and the second emotion recognition result are identified, and the same expression tags are used as target tags. For example, target tags may include disgust, anger, contempt, and calmness. Then, the first number of target tags in the first emotion recognition result is multiplied by a first preset weight to obtain a first expression value, and the second number of target tags in the second emotion recognition result is multiplied by a second preset weight to obtain a second expression value. Then, the first expression value and the second expression value are summed to obtain multiple final expression values. Finally, the largest final expression value among the multiple final expression values ​​is determined, and the emotion feature is determined according to the expression tag corresponding to the largest final expression value, thereby improving the accuracy of the emotion feature.

[0068] Step S206: Obtain the response text based on the speech recognition result of the voice command, and obtain the response animation clip corresponding to the response text from the preset animation clip library.

[0069] In this embodiment, a preset speech recognition algorithm is used to perform speech recognition on the speech command to obtain the speech recognition result. The response text is then obtained based on the speech recognition result. For example, if the speech command is a question, the response text is the answer to that question; or, if the speech command requests the execution of a target action, the response text is a descriptive text corresponding to that target action. A preset animation clip library includes multiple preset animation clips corresponding to the target digital human. For each preset animation clip, there is a corresponding response text. The response animation clip is obtained by querying the preset animation clip library based on the response text.

[0070] Step S207: Adjust the response animation clip based on the emotional characteristics to obtain a target animation clip, and generate a response animation corresponding to the target digital human based on the target animation clip.

[0071] After acquiring emotional features and response animation clips, the response animation clips are adjusted based on the emotional features to obtain target animation clips that match those features. Then, response animations corresponding to the target digital human are generated based on these target animation clips. For example, the target digital human can be driven by the skeletal and facial expression parameters of the target animation clips to generate corresponding response animations. Thus, response animations that correspond to voice commands and match the user's emotional features are generated based on the target digital human, enabling interaction with the user.

[0072] By applying the above technical solutions, a target digital human for interaction is obtained; the intensity and pronunciation features of the user's voice are determined according to the voice command; a first emotion recognition result is generated based on the intensity features, the pronunciation features, and a preset feature index range; emotion recognition is performed on the facial expression image based on a preset facial expression recognition model to obtain a second emotion recognition result; the emotion feature is determined based on the first and second emotion recognition results; response text is obtained based on the voice recognition result of the voice command, and a response animation clip corresponding to the response text is obtained from a preset animation clip library; the response animation clip is adjusted based on the emotion feature to obtain a target animation clip, and a response animation corresponding to the target digital human is generated based on the target animation clip. This adjusts the response animation clip based on the user's emotion feature, thereby enabling the digital human to display response animations that match the user's emotions, achieving more efficient interaction with the digital human and improving the user experience.

[0073] This application also proposes an interactive control method for digital humans, such as... Figure 3 As shown, it includes the following steps:

[0074] Step S301: Obtain the target digital human for interaction.

[0075] The target digital human can be selected by the user from multiple preset digital humans, or it can be created or uploaded by the user, or it can be assigned to the user according to preset allocation rules, such as assigning a digital human that matches the user's personality traits or user profile data.

[0076] Step S302: Determine the user's emotional characteristics based on the user's voice commands and facial expression images.

[0077] When interaction with the target digital human is needed, the user issues corresponding voice commands. These commands are captured via a pre-set microphone and can include, for example, asking a question, requesting music to be played, or instructing the digital human to perform a certain action. The voice commands can be standard commands conforming to preset command specifications or arbitrary commands issued by the user. Simultaneously, a pre-set camera captures the user's facial image, obtaining their expression. Based on the voice commands and the user's facial expressions, the system determines the user's emotional characteristics. For example, emotional characteristics could include calmness, happiness, depression, anger, sadness, etc.

[0078] Step S303: Obtain the response text based on the speech recognition result of the voice command, and obtain the response animation clip corresponding to the response text from the preset animation clip library.

[0079] In this embodiment, a preset speech recognition algorithm is used to perform speech recognition on the speech command to obtain the speech recognition result. The response text is then obtained based on the speech recognition result. For example, if the speech command is a question, the response text is the answer to that question; or, if the speech command requests the execution of a target action, the response text is a descriptive text corresponding to that target action. A preset animation clip library includes multiple preset animation clips corresponding to the target digital human. For each preset animation clip, there is a corresponding response text. The response animation clip is obtained by querying the preset animation clip library based on the response text.

[0080] Step S304: Obtain the skeletal motion parameters of the response animation clip.

[0081] Obtain multiple skeletal key points of the target digital human, and determine the skeletal motion parameters based on the motion parameters of each skeletal key point in the response animation clip.

[0082] Step S305: Input the emotional features and the skeletal motion parameters into a preset emotion adjustment model, and obtain the target skeletal motion parameters based on the output of the preset emotion adjustment model.

[0083] Different emotions can influence skeletal motion parameters. For example, in a happy mood, the movement rate and amplitude of key skeletal points are both greater, while in a sad mood, the corresponding movement rate and amplitude are smaller. To accurately determine the impact of different emotions on skeletal motion parameters, a pre-trained emotion adjustment model was developed based on sample data including different emotions and skeletal motion parameters. After acquiring the emotion features and skeletal motion parameters, these features and parameters were input into the pre-trained emotion adjustment model to obtain the adjusted target skeletal motion parameters.

[0084] Step S306: Replace the skeletal motion parameters with the target skeletal motion parameters in the response animation clip to obtain the target animation clip, and generate a response animation corresponding to the target digital human based on the target animation clip.

[0085] The target skeletal motion parameters match the user's emotional characteristics. In the response animation clip, the skeletal motion parameters are replaced with the target skeletal motion parameters to obtain a target animation clip that matches the user's emotional characteristics. The target digital human is driven according to the skeletal motion parameters in the target animation clip, thereby generating a response animation that corresponds to the voice command and matches the user's emotional characteristics based on the target digital human, realizing interaction with the user.

[0086] By applying the above technical solutions, a target digital human for interaction is obtained; the user's emotional characteristics are determined based on the user's voice commands and facial expressions; response text is obtained based on the speech recognition results of the voice commands, and a corresponding response animation clip is obtained from a preset animation clip library; the skeletal motion parameters of the response animation clip are obtained; the emotional characteristics and the skeletal motion parameters are input into a preset emotion adjustment model, and the target skeletal motion parameters are obtained based on the output of the preset emotion adjustment model; the skeletal motion parameters are replaced with the target skeletal motion parameters in the response animation clip to obtain the target animation clip, and a response animation corresponding to the target digital human is generated based on the target animation clip. This adjusts the response animation clip based on the user's emotional characteristics, enabling the digital human to display response animations that match the user's emotions, achieving more efficient interaction with the digital human and improving the user experience.

[0087] This application also proposes an interactive control device for digital humans, such as... Figure 4 As shown, the device includes:

[0088] The first acquisition module 401 is used to acquire a target digital human for interaction; the determination module 402 is used to determine the user's emotional characteristics based on the user's voice command and the user's facial expression image; the second acquisition module 403 is used to acquire response text based on the speech recognition result of the voice command, and acquire a response animation clip corresponding to the response text from a preset animation clip library, wherein the preset animation clip library includes multiple preset animation clips corresponding to the target digital human; the generation module 404 is used to adjust the response animation clip based on the emotional characteristics to obtain a target animation clip, and generate a response animation corresponding to the target digital human based on the target animation clip.

[0089] In a specific application scenario, the determining module 402 is specifically used to: determine the intensity features and pronunciation features of the user's voice according to the voice command; generate a first emotion recognition result based on the intensity features, the pronunciation features, and a preset feature index range; perform emotion recognition on the expression image based on a preset expression recognition model to obtain a second emotion recognition result; and determine the emotion feature based on the first emotion recognition result and the second emotion recognition result.

[0090] In a specific application scenario, the determining module 402 is further configured to: normalize the intensity feature and the pronunciation feature respectively to obtain an intensity normalized value and a pronunciation normalized value; divide the preset feature index interval based on a preset segmentation value to obtain multiple feature index sub-intervals; and generate the first emotion recognition result based on the degree of overlap between the intensity normalized value and the pronunciation normalized value and the multiple feature index sub-intervals.

[0091] In a specific application scenario, the determining module 402 is further configured to: determine the same expression tags in the first emotion recognition result and the second emotion recognition result, and use the same expression tags as target tags; multiply the first number of target tags in the first emotion recognition result by a first preset weight to obtain a first expression value; multiply the second number of target tags in the second emotion recognition result by a second preset weight to obtain a second expression value; sum the first expression value and the second expression value to obtain a final expression value; determine the maximum final expression value among multiple final expression values; and determine the emotion feature based on the expression tag corresponding to the maximum final expression value.

[0092] In a specific application scenario, the generation module 404 is specifically used to: obtain the skeletal motion parameters of the response animation clip; input the emotional features and the skeletal motion parameters into a preset emotion adjustment model, and obtain the target skeletal motion parameters based on the output of the preset emotion adjustment model; and replace the skeletal motion parameters with the target skeletal motion parameters in the response animation clip to obtain the target animation clip.

[0093] In a specific application scenario, the second acquisition module 403 is specifically used for: acquiring instruction text based on the speech recognition result; if a preset response text corresponding to the instruction text exists in the preset corpus, the preset response text is used as the response text; if the preset response text does not exist in the preset corpus, the instruction text is processed based on a preset language processing model to obtain the response text.

[0094] In specific application scenarios, each of the preset animation segments is configured with a text tag. The second acquisition module 403 is further specifically used to: compare the response text with each of the text tags; if there is a matching text tag among the text tags that matches the response text, the preset animation segment corresponding to the matching text tag is used as the response animation segment; if there is no matching text tag among the text tags, the preset default animation segment is used as the response animation segment.

[0095] By applying the above technical solutions, the interactive control device for digital humans includes: a first acquisition module for acquiring a target digital human for interaction; a determination module for determining the user's emotional characteristics based on the user's voice commands and facial expression images; a second acquisition module for acquiring response text based on the speech recognition result of the voice commands, and acquiring a response animation clip corresponding to the response text from a preset animation clip library, wherein the preset animation clip library includes multiple preset animation clips corresponding to the target digital human; and a generation module for adjusting the response animation clip based on the emotional characteristics to obtain a target animation clip, and generating a response animation corresponding to the target digital human based on the target animation clip. This adjustment of the response animation clip based on the user's emotional characteristics allows the digital human to display response animations that match the user's emotions, achieving more efficient interaction with the digital human and improving the user experience.

[0096] This invention also provides an electronic device, such as... Figure 5 As shown, it includes a processor 501, a communication interface 502, a memory 503, and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.

[0097] Memory 503 is used to store the processor's executable instructions;

[0098] Processor 501 is configured to execute the following via executing the executable instructions:

[0099] A target digital human for interaction is acquired; the user's emotional characteristics are determined based on the user's voice commands and facial expressions; response text is acquired based on the speech recognition results of the voice commands, and a response animation clip corresponding to the response text is acquired from a preset animation clip library, wherein the preset animation clip library includes multiple preset animation clips corresponding to the target digital human; the response animation clip is adjusted based on the emotional characteristics to obtain the target animation clip, and a response animation corresponding to the target digital human is generated based on the target animation clip.

[0100] The aforementioned communication bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.

[0101] The communication interface is used for communication between the aforementioned terminal and other devices.

[0102] The memory may include RAM (Random Access Memory) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0103] The processors mentioned above can be general-purpose processors, including CPUs (Central Processing Units), NPs (Network Processors), etc.; they can also be DSPs (Digital Signal Processors), ASICs (Application Specific Integrated Circuits), FPGAs (Field Programmable Gate Arrays), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0104] In another embodiment of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, which, when executed by a processor, implements the interactive control method for a digital human as described above.

[0105] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the interactive control method for a digital human as described above.

[0106] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0107] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0108] The various embodiments in this specification are described in a related manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0109] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A method for interactive control of a digital human, characterized in that, The method includes: Acquire the target digital human for interaction; The user's emotional characteristics are determined based on the user's voice commands and facial expression images; The response text is obtained based on the speech recognition result of the voice command, and the response animation clip corresponding to the response text is obtained from the preset animation clip library, wherein the preset animation clip library includes multiple preset animation clips corresponding to the target digital human. The response animation clip is adjusted based on the emotional characteristics to obtain a target animation clip, and a response animation corresponding to the target digital human is generated based on the target animation clip; Determining the user's emotional characteristics based on the user's voice commands and facial expressions includes: The intensity and pronunciation characteristics of the user's voice are determined based on the voice commands. A first emotion recognition result is generated based on the intensity features, the pronunciation features, and the preset feature index range; Based on a preset facial expression recognition model, emotion recognition is performed on the facial expression image to obtain a second emotion recognition result; The emotion characteristics are determined based on the first emotion recognition result and the second emotion recognition result; The step of adjusting the response animation clip based on the emotional characteristics to obtain the target animation clip includes: Obtain the skeletal motion parameters of the response animation clip; The emotional features and the skeletal motion parameters are input into a preset emotion adjustment model, and the target skeletal motion parameters are obtained based on the output of the preset emotion adjustment model. The target animation clip is obtained by replacing the skeletal motion parameters with the target skeletal motion parameters in the response animation clip.

2. The method as described in claim 1, characterized in that, The step of generating a first emotion recognition result based on the intensity features, pronunciation features, and preset feature index range includes: The intensity feature and the pronunciation feature are normalized respectively to obtain the intensity normalization value and the pronunciation normalization value; The preset feature index interval is divided based on the preset segmentation value to obtain multiple feature index sub-intervals; The first emotion recognition result is generated based on the degree of overlap between the intensity normalization value and the pronunciation normalization value and multiple feature index sub-intervals.

3. The method as described in claim 1, characterized in that, Determining the emotional characteristics based on the first emotion recognition result and the second emotion recognition result includes: Identify the same facial expression labels in the first emotion recognition result and the second emotion recognition result, and use the same facial expression labels as target labels; The first number of target labels in the first emotion recognition result is multiplied by the first preset weight to obtain the first expression value; The second number of target labels in the second emotion recognition result is multiplied by the second preset weight to obtain the second expression value; The first expression value and the second expression value are summed to obtain the final expression value. The maximum final expression value among the multiple final expression values ​​is determined, and the emotional feature is determined according to the expression tag corresponding to the maximum final expression value.

4. The method as described in claim 1, characterized in that, The step of obtaining the response text based on the speech recognition result of the voice command includes: The instruction text is obtained based on the speech recognition results; If a preset response text corresponding to the instruction text exists in the preset corpus, the preset response text shall be used as the response text. If the preset response text is not found in the preset corpus, the instruction text is processed based on the preset language processing model to obtain the response text.

5. The method as described in claim 1, characterized in that, Each of the preset animation clips is configured with a text tag, and the step of retrieving the response animation clip corresponding to the response text from the preset animation clip library includes: The response text is compared with each of the text tags; If any of the text tags contains a matching text tag that matches the response text, the preset animation clip corresponding to the matching text tag will be used as the response animation clip. If no matching text tag is found among the text tags, a preset default animation clip will be used as the response animation clip.

6. An interactive control device for a digital human, characterized in that, The device includes: The first acquisition module is used to acquire the target digital human for interaction. The determination module is used to determine the user's emotional characteristics based on the user's voice commands and the user's facial expression images; The second acquisition module is used to acquire response text based on the speech recognition result of the voice command, and acquire response animation clips corresponding to the response text from a preset animation clip library, wherein the preset animation clip library includes multiple preset animation clips corresponding to the target digital human. The generation module is used to adjust the response animation clip based on the emotional characteristics to obtain a target animation clip, and generate a response animation corresponding to the target digital human based on the target animation clip; Determining the user's emotional characteristics based on the user's voice commands and facial expressions includes: The intensity and pronunciation characteristics of the user's voice are determined based on the voice commands. A first emotion recognition result is generated based on the intensity features, the pronunciation features, and the preset feature index range; Based on a preset facial expression recognition model, emotion recognition is performed on the facial expression image to obtain a second emotion recognition result; The emotion characteristics are determined based on the first emotion recognition result and the second emotion recognition result; The step of adjusting the response animation clip based on the emotional characteristics to obtain the target animation clip includes: Obtain the skeletal motion parameters of the response animation clip; The emotional features and the skeletal motion parameters are input into a preset emotion adjustment model, and the target skeletal motion parameters are obtained based on the output of the preset emotion adjustment model. The target animation clip is obtained by replacing the skeletal motion parameters with the target skeletal motion parameters in the response animation clip.

7. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the interactive control method for a digital human according to any one of claims 1 to 5 by executing the executable instructions.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the interactive control method for digital humans as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Virtual character expression driving method and system

    CN113506360A

  • Method for automatically generating virtual human animation based on text

    CN114419208A