Digital human interaction method and device, electronic equipment and computer storage medium

By pre-building a preset response text library and response audio database, and using question intent recognition to quickly match target text and audio data, the problem of slow interactive response of digital humans is solved, and faster interactive response and a smarter interactive experience are achieved.

CN120669848APending Publication Date: 2025-09-19YOUKU CULTURE TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510571783.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The digital human interaction solutions in the existing technology have a low response speed, resulting in poor smoothness of the interaction process.

Method used

By pre-building a preset response text library and response audio database, question intent recognition is used to quickly match target preset text and audio data, driving the digital human to make interactive responses.

Benefits of technology

It improves the interactive response speed of digital humans, reduces user waiting time, improves the smoothness of interaction, and makes the interactive process more intelligent and in line with user intentions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669848A_ABST
    Figure CN120669848A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a digital human interaction method and device, electronic equipment and a computer storage medium. The digital human interaction method comprises the following steps: acquiring question data to be responded, and performing question intention recognition on the question data to obtain a question intention; in a pre-constructed preset response text library, determining a target preset text matched with the question intention; and obtaining target audio data corresponding to the target preset text, so that the target audio data is adopted to drive the digital human to perform interactive response on the question data. According to the embodiment of the invention, the interaction response speed of the digital human can be remarkably improved, the waiting time of the user is shortened, the interaction fluency of the digital human is effectively improved, and meanwhile, the interaction process is more intelligent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a digital human interaction method, device, electronic device, and computer storage medium. Background Art

[0002] A digital human refers to a digitized or virtualized character or individual that exists in a virtual environment. For example, it can be an animated character, a virtual assistant, or a customized character generated based on artificial intelligence technology.

[0003] Digital human interaction is a new form of interaction that combines artificial intelligence, speech recognition, natural language processing, and digital human generation technologies. Digital human interaction involves users interacting with a digital human through voice or text. The digital human understands the user's input commands or questions and responds in voice. Compared to conventional text-based interaction, this digital human interaction not only integrates voice communication but also incorporates the digital human's visual image, such as mouth movements, facial expressions, and hand gestures, effectively enhancing the realism of the interaction.

[0004] However, in the digital human interaction solutions in related technologies, after the user inputs the interactive content, it usually takes a long time for the digital human to respond, and the interactive response speed is low. Therefore, the smoothness of the interactive process is poor. Summary of the Invention

[0005] In view of this, an embodiment of the present application provides an interactive solution to at least partially solve the above problems.

[0006] According to a first aspect of an embodiment of the present application, a digital human interaction method is provided, comprising:

[0007] Acquire question data to be responded to, and perform question intent recognition on the question data to obtain the question intent;

[0008] Determining a target preset text that matches the question intention in a pre-built preset response text library;

[0009] Target audio data corresponding to the target preset text is acquired, so as to use the target audio data to drive the digital human to interactively respond to the question data.

[0010] According to a second aspect of an embodiment of the present application, a digital human interactive device is provided, comprising:

[0011] Extracting data acquisition module, used to obtain question data to be responded to, and identify the question intention of the question data to obtain the question intention;

[0012] A preset text determination module is used to determine a target preset text that matches the question intention in a pre-built preset response text library;

[0013] The interactive response module is used to obtain target audio data corresponding to the target preset text, so as to use the target audio data to drive the digital human to interactively respond to the question data.

[0014] According to the third aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect.

[0015] According to a fourth aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect is implemented.

[0016] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, which contains computer instructions, and the computer instructions instruct a computing device to perform operations corresponding to the method described in the first aspect.

[0017] According to the digital human interaction solution provided in the embodiment of the present application, after obtaining the question data to be responded to, the question intention is identified for the question data to obtain the question intention; then, in the pre-built preset response text library, the target preset text and the corresponding target audio data that match the question intention are determined, and the target audio data is used to drive the digital human to interactively respond to the question data.

[0018] In the embodiment of the present application, on the one hand, through the pre-loading strategy of the pre-response text, the interactive text corresponding to the question data can be quickly obtained, and then the digital human can be driven to perform an interactive response based on the audio data corresponding to the interactive text, so that the digital human can quickly speak and answer and enter an interactive state with the user; on the other hand, when obtaining the pre-set interactive text, the question intention is identified on the question data, and after identifying the user intention, the target preset text that matches the user intention is used as the interactive text. Therefore, when the digital human performs an interactive response based on the above-mentioned interactive text, the content of the interactive response can be made more in line with the user's question intention.

[0019] In summary, the embodiments of the present application can improve the interactive response speed of digital humans, reduce user waiting time, effectively improve the smoothness of digital human interaction, and at the same time, make the interactive process more intelligent. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0021] Figure 1 A schematic architecture diagram for digital human applications;

[0022] Figure 2 is a flowchart of the steps of a digital human interaction method according to an embodiment of the present application;

[0023] Figure 3 Schematic diagram of the interactive audio acquisition process according to an embodiment of the present application;

[0024] Figure 4 A schematic diagram of an exemplary interactive system applicable to the digital human interaction method according to an embodiment of the present application;

[0025] Figure 5 Based on Figure 4 Schematic diagram of the interaction process of the interactive system shown;

[0026] Figure 6 is a structural block diagram of a digital human interaction device according to an embodiment of the present application;

[0027] Figure 7 Schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present application.

[0029] See also Figure 1 , Figure 1 This is a schematic diagram of the architecture of a digital human application for providing digital human interactive services. Figure 1 An example explanation of the internal architecture of the Digital Human application is given below:

[0030] Digital human applications can refer to applications that provide digital human interactive services. Digital human applications can include the following four layers: hardware layer, computing power layer, and presentation logic layer. Schematically:

[0031] The hardware layer can be configured with a display screen for displaying the digital human image and a pre-set operating system. Furthermore, other accessories can be configured, such as a camera for image acquisition, a microphone for sound signal acquisition, a laser pointer for demonstration assistance, and chips for implementing interactive functions.

[0032] The computing power layer can provide computing power resources for a variety of different types of services (such as portrait recognition, portrait tracking, intelligent voice wake-up, binary classification model prediction, etc.). Furthermore, in order to achieve more efficient, flexible and reliable computing resource utilization, the computing power layer can also be divided into a local computing power layer and a remote computing power layer. Among them, the local computing power layer is used to provide computing power resources to services with low resource consumption, such as: portrait recognition, portrait tracking, intelligent voice wake-up, binary classification model prediction, etc., while the remote computing power layer is used to support services with higher computing power resource requirements, such as: real-time speech recognition, general text vector conversion of recognized text, vector retrieval based on vector data, semantic matching based on recognized text, etc.

[0033] The presentation logic layer may include a state orchestration module, a real-time rendering module, and a model customization module. The state orchestration module is used to orchestrate and update the digital human's state based on the services supported by the computing power layer. For example, the digital human's states may include idle, listening, and greeting states.

[0034] The real-time rendering module can synthesize the digital human's mouth and body movements based on the audio data, generate corresponding digital human image video frames, and use the real-time rendering engine to render the images based on the above video frames to drive the digital human to perform mouth and body movements that match the audio data on the hardware layer display screen.

[0035] The model customization module can provide customized functions based on the neural network model, such as: user emotion prediction function, real-time question and answer function with users, function of conducting multi-round conversations with users, etc.

[0036] Reference Figure 2 , Figure 2 The figure is a flowchart of a digital human interaction method according to an embodiment of the present application. Schematically, the digital human interaction method provided in this embodiment includes the following steps:

[0037] Step 202: Acquire question data to be responded to, and perform question intention recognition on the question data to obtain the question intention.

[0038] Schematically, during the interaction with the Digital Human application, the user can input interactive data into the Digital Human application. This interactive data is the question data to be responded to in the embodiment of this application. In the embodiment of this application, the specific content of the interactive data is not limited and can be any content input by the user. In addition, in the embodiment of this application, the specific method for obtaining the question data is also not limited. For example, the user can provide the question data through text input or send a voice question to the Digital Human application through the microphone in the hardware layer. Furthermore, after the user's question voice is accurately captured, it can be processed by the real-time speech recognition service in the computing power layer to obtain the question data in text format for subsequent processing.

[0039] After obtaining the question data, in this embodiment of the application, question intent recognition is performed on the question data to determine the user's question intent. In this embodiment of the application, question intent can refer to the type of information or purpose of the question that the user desires when asking a specific question to a digital human. Question intent recognition can help understand the core purpose and needs of the user's question.

[0040] Through the above step 202 of the present application, the question data and question intention can be obtained. The obtained question data and question intention have the following differences and relationships: the question data is usually a specific question expressed by the user in the form of language, which is the external manifestation of the user's question intention, while the question intention can be the user's real needs hidden behind the question data. Through in-depth understanding and analysis of the question data, the corresponding question intention can be obtained. For example: for the user's question data "Why is the sky blue?", the corresponding question intention can be: "Common sense question and answer"; for the user's question data "How to choose a suitable loss function in deep learning?", the corresponding question intention can be: "Professional question and answer"; for the user's question data "Can you show me a magic trick?", the corresponding question intention can be: "Talent show"; for the user's question data "I completed a very important project today and feel very happy. Will you be happy for me?", the corresponding question intention can be: "User shares his happiness", etc.

[0041] In an embodiment of the present application, any suitable intent recognition method can be used to obtain the question intention, for example, such as: rule-based intent recognition method (such as keyword matching, regular expression), statistics-based intent recognition method (such as bag-of-words model), and deep learning-based method, etc.

[0042] Step 204: Determine a target preset text that matches the question intention in the pre-built preset response text library.

[0043] Illustratively, a response text library containing a variety of different response texts can be pre-built (for ease of description, in the embodiments of the present application, the pre-built response texts are referred to as preset response texts, and the text library containing the preset response texts is referred to as a preset response text library). After the user's questioning intention is obtained in step 202, a preset response text that is semantically more compatible with the user's questioning intention can be determined from the preset response text library, i.e., a target preset text.

[0044] In the embodiment of the present application, the specific method used to determine the target preset text from the preset response text library is not limited and can be customized according to actual circumstances. For example, keywords in the question intent can be obtained, and then the preset response text containing the question intent keywords can be determined as the target preset text; for another example, the preset response text library can also include question intents corresponding to the preset response texts. After the user's current question intent is obtained through the above step 202, the user's current question intent is matched with each question intent included in the preset response text library, so that the preset response text corresponding to the question intent matched from the preset response text library is determined as the target preset text.

[0045] In order to further improve the interactive response speed of the digital human, the matching mapping relationship between the question intention and the preset response text can also be pre-set. After identifying the user's question intention, the target preset text can be efficiently determined based on the above matching mapping relationship. For example: if the question intention is: "General knowledge Q&A or professional Q&A", the preset response text that matches the question intention can be set to any of the following: "Don't be impatient, wait for me to search my database", "Hey, test me, right?", "Yeah, the words are on the tip of my tongue, let me think about it"; if the question intention is: "Discussion of opinions", the preset response text that matches the question intention can be set to any of the following: "We can have a good chat about this issue", "I need to think about this issue"; if the question intention is: "Talent show", the preset response text that matches the question intention can be set to any of the following: "Hey, is it time for the talent show?", "You are good at asking questions"; if the question intention is: "Users share their happiness", the preset response text that matches the question intention can be set to any of the following: "Wow, I can feel your happiness through the screen!", "Great! Seeing you so happy, I am also infected!"; if the question intention is: "Users share their sadness", the preset response text that matches the question intention can be set to any of the following: "Don't be sad, I will always be with you", "I can imagine that if it were me, I might have similar feelings", etc.

[0046] In addition, since the preset response text is pre-set before the user's question intention is obtained, after the user's question intention is recognized in real time, the following situation may occur: the target preset text that matches the question intention does not exist in the pre-built preset response text library. In response to this situation, in an embodiment of the present application, a universal response text can be further pre-set. When the preset response text library does not contain a preset response text that matches the question intention, the above universal response text is used as the target preset text. Exemplarily, the above universal response text can be "Hmm...", "Wait a moment", "Let me think about it", and the like.

[0047] Step 206: Acquire target audio data corresponding to the target preset text, and use the target audio data to drive the digital human to interactively respond to the question data.

[0048] For example, to enhance the fun and authenticity of interactions, digital human applications typically display mouth movements, facial expressions, or hand gestures that match the spoken content while providing interactive responses via voice. Therefore, after acquiring the target preset text, the target audio data corresponding to the target preset text can be further acquired. Based on this target audio data, the digital human application can not only conduct voice interactions, but also display dynamic visual images that match the spoken content.

[0049] Exemplarily, the process of using target audio data to drive a digital human to interactively respond to question data may include: generating a digital human image video frame that conforms to the situation and current emotional state based on the target audio data through a pre-trained digital human image prediction model; while playing the above-mentioned target audio data, synchronously playing the above-mentioned digital human image video frame, so that the digital human can show expected voice, expression and action, etc.

[0050] In the embodiment of the present application, on the one hand, through the pre-loading strategy of the pre-response text, the interactive text corresponding to the question data can be quickly obtained, and then the digital human can be driven to perform an interactive response based on the audio data corresponding to the interactive text, so that the digital human can quickly speak and answer and enter an interactive state with the user; on the other hand, when obtaining the pre-set interactive text, the question intention is identified on the question data, and after identifying the user intention, the target preset text that matches the user intention is used as the interactive text. Therefore, when the digital human performs an interactive response based on the above-mentioned interactive text, the content of the interactive response can be made more in line with the user's question intention.

[0051] In summary, the embodiments of the present application can improve the interactive response speed of digital humans, reduce user waiting time, effectively improve the smoothness of digital human interaction, and at the same time, make the interactive process more intelligent.

[0052] Optionally, in some of the embodiments, the process of acquiring target audio data corresponding to the target preset text may include: determining the target audio data corresponding to the target preset text in a pre-built response audio database.

[0053] Illustratively, in the above-described embodiment of the present application, in addition to pre-establishing preset response texts, a response audio database is also pre-established. This response audio database contains audio data corresponding to each preset response text. After obtaining the target preset text, the corresponding target audio data can be further selected from the response audio database, eliminating the need for real-time text-to-audio conversion. This can further improve the interactive response speed of the digital human.

[0054] For digital human applications, in order to enhance the fun and authenticity of interaction, in addition to voice output, visual outputs such as mouth movements, facial expressions, and hand movements are usually combined during the interaction with the user. Therefore, further, for each audio data in the response audio database, in the embodiment of the present application, the digital human's mouth and body movements can be synthesized in advance based on the audio data to generate a digital human image video frame corresponding to each audio data. In other words, based on each audio data, a pre-trained digital human image prediction model is used to generate a digital human image video frame that conforms to the context and emotional state corresponding to each audio data. Once the target audio data is determined, the target digital human image video frame corresponding to the target audio data can be efficiently obtained so that the target digital human image video frame can be played synchronously with the target audio data, so that the digital human displays the expected voice, facial expression, and movement.

[0055] In addition to pre-generating audio data, the above embodiment also pre-generates digital human avatar video frames corresponding to each audio data set. This allows the target digital human avatar video frames corresponding to the target audio data to be obtained without the need for real-time video frame generation. This further improves the interactive response speed of the digital human.

[0056] Optionally, in some embodiments, the digital human interaction method may further include:

[0057] If the target audio data corresponding to the target preset text does not exist in the response audio database, generate the target audio data corresponding to the target preset text by using text-to-speech technology;

[0058] Based on the target audio data, the response audio database is updated.

[0059] Schematically, due to reasons such as asynchronous data updates, it may happen that the response audio database does not contain the audio data corresponding to the preset response text. When the response audio database does not contain the target audio data corresponding to the target preset text, the target audio data can be generated in real time through text-to-speech technology, and the newly generated target audio data can be stored in the response audio database to update the corresponding audio database. Through the above method, on the one hand, the smooth execution of this interactive response can be guaranteed; on the other hand, the response audio database is updated based on the target audio data. In this way, when question data with the same question intention is received again, the digital human can be driven based on the newly generated target audio data without having to perform the real-time text-to-speech operation again. Therefore, the interactive response speed of the digital human can be further improved.

[0060] In the embodiment of the present application, when generating target audio data corresponding to the target preset text, any suitable text-to-speech technology may be used, for example, a text-to-speech technology based on deep learning, or related text-to-speech software, etc.

[0061] Optionally, in some embodiments, after obtaining the question data, the digital human interaction method may further include:

[0062] Based on the question data, a formal response text corresponding to the question data is generated; and formal audio data corresponding to the formal response text is generated.

[0063] Correspondingly, the process of using the target audio data to drive the digital human to interactively respond to the question data may also include:

[0064] The target audio data is used to drive the digital human to carry out an interactive response to the question data; after the interactive response is completed, the formal audio data is used to drive the digital human to carry out a formal interactive response to the question data.

[0065] Illustratively, question data can be input into a large language model to generate answer data for the question data. Alternatively, the question data can be combined with the question intent to generate answer data for the question data. Illustratively, after identifying the user's question intent, answer data for the question data, i.e., formal response text, can be generated based on the question data and question intent using a large language model. The generation process can include: performing a search based on the question data and question intent to obtain search results; generating prompt information based on the search results, question data, and question intent; and inputting the prompt information into the large language model to guide the large language model to output the formal response text for the question data through the prompt information.

[0066] After obtaining the formal response text, text-to-speech technology can be used to obtain formal audio data corresponding to the formal response text, so as to drive the digital human to make a formal interactive response based on the formal audio data.

[0067] See also Figure 3 , Figure 3 The following is a schematic diagram of the interactive audio acquisition process according to an embodiment of the present application. Figure 3 The interactive audio acquisition process of the above embodiment of the present application is briefly described as follows:

[0068] After the interaction begins, the question data input by the user is obtained first, and the question data is recognized to obtain the user's question intention; on the one hand, the target preset text that matches the question intention can be determined in the pre-built preset response text library; on the other hand, a query word (Query) can be generated based on the question data and question intention, and the query word can be searched in multiple retrieval databases ( Figure 3 4 retrieval databases are listed as examples, which does not limit the number of retrieval databases) to obtain retrieval results; prompt information is generated according to the retrieval results, question data and question intention, and the large language model is guided by the above prompt information to output the formal response text for the question data; after obtaining the target preset text and the formal response text, the text-to-speech technology is used to obtain interactive audio data: target audio data and formal audio data. At this point, the interactive audio data acquisition process is completed.

[0069] Large language models usually rely on the knowledge learned in the training phase to generate text. Since the knowledge in the training phase is usually static and will not be dynamically updated over time, the final generated text is poor in accuracy and comprehensiveness. In the above embodiment of the present application, query words are first generated based on the question data and the question intention, and a search is performed in the retrieval database based on the query words, and then the large language model is guided to generate a formal response text based on the search results. Through retrieval, new knowledge content related to the query words can be obtained from the external database as a supplement to the existing internal knowledge of the large language model, thereby enabling the large language model to generate more accurate and comprehensive formal response texts.

[0070] In addition, see Figure 3 When the intent recognition model determines that the user's question is meaningless data (for example, the question data consists of a large number of filler words such as "um", "ah", "oh", etc., which have no actual meaning), the process can be ended in advance and wait for the user's next question data.

[0071] In the above embodiment of the present application, in addition to realizing the follow-up interactive response based on the pre-set preset response text, the formal answer content for the user's question data is also generated through the large language model: the formal response text, and then after the above-mentioned follow-up interactive response is completed, the formal interactive response is realized based on the formal audio data corresponding to the above-mentioned formal response text.

[0072] This approach not only improves the digital human's interactive response speed but also allows for sufficient processing time for formal interactive responses. Furthermore, because both the target preset text and the formal response text are derived based on the question's intent, they possess a high degree of semantic relevance. This makes the transition between the follow-up interactive response and the formal interactive response more natural, achieving a seamless connection between the follow-up interactive response and the formal interactive response, allowing the digital human to interact with the user in a more natural and humane manner.

[0073] Optionally, in some embodiments, the question intention is obtained through a pre-trained intention recognition model, and the digital human interaction method may further include:

[0074] Obtain the initial intent recognition model;

[0075] Obtain a fine-tuning sample set, which includes: a fine-tuning question sample and a question intention label; wherein the question intention label satisfies the following conditions: a preset response text matching the question intention label and a semantic relevance between the formal response text of the fine-tuning question sample and the label satisfies a preset condition;

[0076] The fine-tuning sample set is used to adjust the parameters of the initial intent recognition model to obtain the trained intent recognition model.

[0077] Illustratively, the initial intent recognition model may be an intent recognition model trained based on an initial sample set, which may include: initial question samples and question intent labels corresponding to the initial question samples.

[0078] The process of adjusting the parameters of the initial intent recognition model using a fine-tuning sample set may include: inputting the fine-tuning question samples in the fine-tuning sample set into the initial intent recognition model, and outputting the predicted question intent through the initial intent recognition model; calculating the loss value based on the predicted question intent and the question intent label corresponding to the fine-tuning question sample, and adjusting the model parameters in the initial intent recognition model using a preset optimization algorithm (such as a gradient descent algorithm, etc.) according to the loss value; returning to the above step of inputting the fine-tuning question sample into the initial intent recognition model, and repeating the above process (sample input, prediction, calculation of loss value, adjustment of parameters) until a preset fine-tuning stop condition is met. Exemplarily, the fine-tuning stop condition may be reaching a preset number of fine-tuning times, or it may be that the loss value no longer decreases, that is, the model converges.

[0079] The preset condition satisfied by the semantic relevance may be: the value of the semantic relevance is higher than a preset relevance threshold.

[0080] In the above embodiment of the present application, after the initial intent recognition model is obtained, the above initial intent recognition model is also fine-tuned through the fine-tuning sample set. In the fine-tuning sample set, the preset response text that matches the question intention label is semantically associated with the formal response text corresponding to the fine-tuning question sample. Therefore, fine-tuning the initial intent recognition model again based on the fine-tuning sample set can make the question intention predicted by the intent recognition model more accurate, and the target preset text obtained based on the predicted question intention is semantically closer to the formal response text. In summary, the above embodiment of the present application can further make the switching between the follow-up interactive response and the formal interactive response more natural, making the digital human interaction process more natural and humane.

[0081] Optionally, in some embodiments, the process of constructing the fine-tuning sample set may include:

[0082] Obtain fine-tuning question samples and formal response texts corresponding to the fine-tuning question samples;

[0083] In the preset response text library, determining related preset texts that are semantically related to the formal response texts of the fine-tuned question samples;

[0084] According to the matching mapping relationship between the preset question intention and the preset response text, the question intention that matches the associated preset text is determined as the question intention label.

[0085] For example, in the above embodiment, after obtaining the fine-tuning question sample and the corresponding formal response text, the associated preset text is first determined in the preset response text library based on the formal response text. For example, the semantic relevance between the formal response text and each preset response text in the preset response text library can be calculated respectively, and then the above-mentioned associated preset text is determined from each preset response text in the preset response text library according to the calculation result. In the embodiment of the present application, there is no limitation on the calculation method of the semantic relevance, and a suitable method can be customized according to the actual situation. For example, cosine similarity can be used as the semantic relevance, that is, the cosine similarity between the formal response text and the preset response text is calculated as the semantic relevance; Manhattan distance can also be used as the semantic relevance, that is, the Manhattan distance between the formal response text and the preset response text is calculated as the semantic relevance, and so on.

[0086] After determining the associated preset text semantically related to the formal response text of the fine-tuned question sample, the question intent matching the associated preset text is reversely determined based on the matching mapping relationship between the question intent and the preset response text, serving as the question intent label. Constructing the fine-tuning sample set in this manner ensures that the question intent label satisfies the following conditions: the preset response text matching the question intent label meets the pre-set semantic correlation with the formal response text of the fine-tuned question sample, thereby improving the efficiency of constructing the fine-tuning sample set.

[0087] Furthermore, in some of the embodiments, after obtaining the trained intent recognition model, historical data can also be collected regularly during the actual reasoning process of the model. The historical data may include: historical question data, the predicted question intent output by the intent recognition model, the preset response text corresponding to the historical question data, and the formal response text corresponding to the historical question data; then, based on the semantic correlation between the preset response text corresponding to the historical question data and the formal response text, data recall (data with lower correlation) is performed from the collected historical data to determine the target historical question data with lower correlation. For each target historical question data, the question intent of the target historical question data is re-determined, and then a new fine-tuning sample set is generated based on the target historical question data and the re-determined question intent, so that the trained intent recognition model can be continuously optimized through the newly generated fine-tuning sample set.

[0088] In the embodiment of the present application, when re-determining the question intention corresponding to the target historical question data, the specific determination method adopted is not limited and can be customized according to actual conditions.

[0089] For example, a specific method for redetermining the question intent corresponding to the target historical question data may include: determining, from a preset response text library, target associated text semantically related to the formal response text of the target historical question data; and, based on the matching mapping relationship between the question intent and the preset response text, determining the question intent that matches the target associated text as the redetermined question intent. In this method, the question intent is redetermined based on the semantic relevance between the preset corresponding text and the formal response text, and therefore, the accuracy of the determined question intent is generally higher.

[0090] For example, in some other embodiments, the corresponding question intent can be re-determined by displaying the target historical question data and multiple question intents; in response to a selection operation on a question intent, the selected question intent is determined as the question intent corresponding to the target historical question data. In this method, the target historical question data and multiple candidate question intents are displayed to the user through an interface display, so that the question intent can be re-determined based on the user's selection. In this process, there is no need to perform a correlation calculation process, so the question intent determination is generally more efficient.

[0091] The above model optimization method uses a sample set recalled from historical data. Therefore, there's no need to regenerate the formal response text, improving the efficiency of sample set construction. Furthermore, the target historical question data included in this sample set is data from which model inference errors occurred. Therefore, model optimization based on this sample set can quickly produce an intent prediction model that meets the requirements.

[0092] Optionally, in some embodiments, after obtaining the initial intent recognition model and before obtaining the fine-tuning sample set, the digital human interaction method may further include:

[0093] Obtain a test question sample and obtain the predicted question intent of the test question sample based on the initial intent recognition model;

[0094] According to the matching mapping relationship between the preset question intention and the preset response text, a predicted preset text matching the predicted question intention is determined from the preset response text library;

[0095] Calculating semantic relevance between the predicted preset text and the formal response text corresponding to the test question sample to obtain a semantic relevance calculation result;

[0096] If the semantic association calculation result does not meet the preset requirements, the step of obtaining a fine-tuning sample set is executed;

[0097] If the semantic association calculation result meets the preset requirements, the initial intent recognition model is determined as the trained intent recognition model.

[0098] Schematically, in the embodiment of the present application, the specific content of the preset requirements is not limited, and can be customized according to actual conditions. For example, the requirements can be set from the dimension of the specific value of the correlation, such as if the semantic correlation calculation result represents that the semantic correlation between the predicted preset text and the formal response text corresponding to the test question sample is greater than the preset correlation threshold, then the semantic correlation calculation result is determined to meet the preset requirements; otherwise, the semantic correlation calculation result is determined to not meet the preset requirements. Requirements can also be set from the two dimensions of the specific value of the correlation and the amount of sample data at the same time, such as if the test question sample whose semantic correlation between the predicted preset text and the formal response text is greater than the preset correlation threshold is determined as the target test question sample, then when the number of target test question samples is greater than the preset number threshold, then the semantic correlation calculation result is determined to meet the preset requirements; otherwise, the semantic correlation calculation result is determined to not meet the preset requirements.

[0099] In the above embodiment of the present application, after obtaining the initial intent recognition model, the intent recognition capability of the initial intent recognition model is first tested by using test question samples. When the text semantic correlation between the preset response text determined according to the predicted question intent output by the initial intent recognition model and the formal response text corresponding to the test question sample meets the preset requirements, it can be considered that the intent recognition capability of the current initial intent recognition model meets the requirements. At this time, the subsequent model fine-tuning operation can be stopped, and the initial intent recognition model is determined as the trained intent recognition model. Conversely, when the text semantic correlation between the preset response text determined according to the predicted question intent output by the initial intent recognition model and the formal response text corresponding to the test question sample fails to meet the preset requirements, it can be considered that the intent recognition capability of the current initial intent recognition model does not meet the requirements. At this time, the model can be fine-tuned based on the fine-tuning sample set to obtain the final trained intent recognition model.

[0100] Optionally, in some embodiments, the process of performing question intention recognition on question data to obtain the question intention may include:

[0101] Obtain the character attribute information of the digital person; obtain the intention generation prompt word based on the question data and the character attribute information; input the intention generation prompt word into the pre-trained intention recognition model, and obtain the question intent contained in the question data through the intention recognition model.

[0102] Illustratively, the character attribute information of a digital human can be information that characterizes the preset attributes possessed by the digital human. In the embodiments of the present application, the specific content of the preset attributes is not limited and can be customized according to actual circumstances. For example, the preset attributes may include at least one of the following: age attribute (childhood type, adolescent type, elderly type, etc.), personality attribute (such as cheerful type, calm type, gentle type, sweet type, etc.), professional attribute (such as medical and health field, education and counseling field, etc.), language style attribute (formal language style, colloquial language style, classical language style, etc.), etc.

[0103] In the above-mentioned embodiment of the present application, when generating the prompt words used to guide the intent recognition model to output the question intent, in addition to obtaining the question data, the data person's personal attribute information is also obtained. This personal attribute information is provided to the intent recognition model as supplementary background knowledge, thereby making the question intent ultimately predicted by the intent recognition model more accurate and more consistent with the digital person's personal attribute settings. Furthermore, based on this more accurate question intent that is consistent with the digital person's personal attribute settings, the target preset text and formal response text are determined, and then interactive responses are performed based on these texts, which can enhance the realism and intimacy of the interaction and improve the user's interactive experience.

[0104] Furthermore, corresponding to the above-mentioned process of obtaining the question intention, the character attribute information of the digital human can also be introduced in the fine-tuning process of the intention recognition model. Specifically, the fine-tuning process of the intention recognition model can also include:

[0105] Obtain the initial intent recognition model and the digital human's personality attribute information;

[0106] Obtain a fine-tuning sample set, which includes: a fine-tuning question sample and a question intention label; wherein the question intention label satisfies the following conditions: a preset response text matching the question intention label and a semantic relevance between the formal response text of the fine-tuning question sample and the label satisfies a preset condition;

[0107] Based on the fine-tuning question sample and the above-mentioned character attribute information, generate intent generation prompt words; input the intent generation prompt words into the initial intent recognition model, and output the predicted question intent through the initial intent recognition model; based on the above-mentioned predicted question intent and the question intent label corresponding to the fine-tuning question sample, calculate the loss value, and adjust the model parameters in the initial intent recognition model according to the loss value; return to the above step of inputting the fine-tuning question sample into the initial intent recognition model, and repeat the above process until the preset fine-tuning stop condition is met.

[0108] The above-mentioned fine-tuning process of the intent recognition model also obtains the personal attribute information of the data person when generating the intent generation prompt word, and provides the personal attribute information as supplementary background knowledge to the initial intent recognition model, so that the question intent finally predicted by the intent recognition model can be more accurate.

[0109] The digital human interaction method provided in the embodiments of this application can be executed by any suitable device that has a digital human application deployed on it. For example, this can be a server or a PC that has a digital human application deployed on it. Alternatively, in some embodiments, the digital human interaction method provided in the embodiments of this application can also be executed by an interactive system comprising a client device and a server device.

[0110] See also Figure 4 , Figure 4 This is a schematic diagram of an exemplary interactive system applicable to the digital human interactive method of the embodiment of the present application. For ease of understanding, first combine Figure 4 The application scenarios of the digital human interaction method provided in the embodiments of the present application are explained.

[0111] like Figure 4 As shown, the system 400 may include a server 402, a communication network 404 and / or one or more client devices 406. Figure 4 The example in the figure shows multiple client devices.

[0112] The server 402 can be any appropriate device capable of performing tasks such as question intent recognition and preset response text matching, including but not limited to a server, a server cluster, a computing cloud server cluster, etc. In some embodiments, the server 402 can perform any appropriate function. For example, in some embodiments, the server 402 can be used to obtain question data sent by the client device 46, perform question intent recognition on the question data to obtain the question intent, determine the target preset text and target audio data that match the question intent, and enable the client device 406 to use the above-mentioned target audio data to drive the digital human to interactively respond to the question data.

[0113] In some embodiments, the communication network 404 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 404 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communication network. The client device 406 can be connected to the communication network 404 via one or more communication links (e.g., communication link 412), and the communication network 404 can be linked to the server 402 via one or more communication links (e.g., communication link 414). The communication link can be any communication link suitable for transmitting data between the client device 406 and the server 402, such as a network link, a dial-up link, a wireless link, a hard-wired link, any other suitable communication link, or any suitable combination of such links.

[0114] Client device 406 may include any one or more user devices having a display device. In some embodiments, client device 406 may include any suitable type of device. For example, in some embodiments, client device 406 may include a mobile phone, a tablet computer, a foldable screen, and / or any other suitable type of user device.

[0115] Based on the above system, the digital human interaction method provided in the embodiment of the present application can also be implemented by Figure 4 The server 402 in the system shown executes. Schematically:

[0116] Step 202 may include: receiving question data to be responded to from the client device, and performing question intention recognition on the question data to obtain the question intention.

[0117] Correspondingly, step 206 may include: obtaining target audio data corresponding to the target preset text, and generating a digital human image video frame based on the target audio data; returning the target audio data and the digital human image video frame to the client device, so that the client device synchronously displays the above-mentioned digital human image video frame during the playback of the target audio data.

[0118] use Figure 4The interactive system shown executes the digital human interaction solution of the embodiment of the present application, and the client device performs the interactive operation with the user to obtain the question data to be responded to, while the server device performs question intention recognition, obtains the target preset text and target audio data, and generates the digital human image video frame. The server is usually equipped with a high-performance processor, which can efficiently generate the digital human image video frame after obtaining the target audio data. Moreover, the server can provide services to multiple client devices at the same time and can make full use of computing resources. In summary, compared with using a local device to execute the digital human interaction solution of the embodiment of the present application, using Figure 4 The interactive system shown executes the digital human interaction solution of the embodiment of the present application, which can further improve the interactive response speed of the digital human and reduce user waiting time while realizing the centralization of computing resources.

[0119] See also Figure 5 , Figure 5 Based on Figure 4 The following is a schematic diagram of the interactive process of the interactive system. Figure 5 Based on Figure 4 The interactive process of the digital human in the interactive system shown is explained as follows:

[0120] In the first step, the user initiates the interaction. The user can ask a question through the microphone. After receiving the voice data, the client device 406 can convert the voice data into text to obtain the question data in text form, and send the question data to the server 402.

[0121] In the second step, the server receives and establishes a connection. The digital human cloud service module in the server 402 will open a websocket connection to establish a real-time communication channel with the client device 406 to ensure that subsequent data flows can be transmitted smoothly.

[0122] The third step is risk detection. After the connection is established, the Digihuman cloud service module in the server 402 will perform risk detection on the received question data, that is, conduct a security analysis of the question data, and filter out potential sensitive information by analyzing the security of the question content. Once the risk detection passes, the subsequent steps can be executed.

[0123] Step 4: User intent recognition and preset text processing. The Digihuman cloud service module in server 402 analyzes the user's question data using an intent recognition model to determine the user's question intent. After determining the user's question intent, it can further determine target preset text that matches the question intent. This target preset text serves as the initial content of the Digihuman's response. To further accelerate the transition from response text to real-time Digihuman interaction, target audio data corresponding to the target preset text can be preset in advance. This target audio data will be transmitted to the Digihuman synthesis engine in subsequent processes to synthesize the Digihuman image video frames based on the target audio data. Illustratively, commonly used preset audio data can be pre-saved and archived according to the question intent. If target audio data matching the user's question intent exists, the subsequent Digihuman-driven process can be executed. If target audio data matching the user's question intent does not exist, the artificial intelligence algorithm module can be invoked in real time to generate and save the target audio data using a text-to-speech model.

[0124] The fifth step is the synthesis and rendering of digital human image video frames. After obtaining the target audio data, the digital human facial expressions and body movements can be coordinated and synthesized based on the target audio data, and the synthesis results can be pushed to the client device 406, so that the user can see the interactive answers to their questions. This stage is mainly divided into the following steps: 1) Image synthesis and real-time push: The digital human synthesis engine in the server 402 generates digital human image video frames that are consistent with the context and current emotional state based on the target audio data through the audio-driven digital human image prediction model, and pushes them to the real-time rendering engine in the server 402 in the form of a video stream. 2) Real-time rendering: The real-time rendering engine in the server 402 will render the digital human image video frames and push the rendering results to the client device 406, so that the digital human can show the expected voice, expression and movement through the client device 406.

[0125] Step 6: Formal Answer. After receiving the user's question intent, the Digihuman cloud service module in the server 402 will synchronously generate a formal response text corresponding to the question data based on the large language model provided by the AI ​​algorithm module during the preset text processing process. The text-to-speech model in the AI ​​algorithm module will be called in real time to generate formal audio data. Similar to the target audio data, after obtaining the formal audio data, the digital human's facial expressions and body movements can also be synthesized based on the formal audio data, and the synthesis results will be pushed to the client device 406.

[0126] Step 7: End of interaction: After the interaction is completed, the Digital Human Cloud Service Module in the server 402 will close the WebSocket connection and release the communication resources.

[0127] During the aforementioned digital human interaction process, multiple operations can be performed simultaneously to improve interactive response speed. For example, during the preset text processing in step 4 and the synthesis and rendering of the digital human image video frames in step 5, the formal response operation in step 6 can also be performed simultaneously. That is, while the digital human cloud service module in server 402 determines the target preset text and target audio data that match the question intent, and the digital human synthesis engine in server 402 synthesizes the digital human's facial expressions and body movements based on the target audio data, the digital human cloud service module in server 402 can also simultaneously generate the formal response text and formal audio data corresponding to the question data based on the large language model provided by the AI ​​algorithm module. Furthermore, the digital human synthesis engine in server 402 can also synthesize the digital human's facial expressions and body movements based on the formal audio data. The client device can then sequentially play the video frames generated in step 5 first, followed by the video frames generated in step 6. If the video frames generated in step 5 have already been played but step 6 has not yet been completed, the video frames generated in step 5 can be played repeatedly to provide the user with a continuous interactive experience.

[0128] Figure 6 This is a structural block diagram of a digital human interaction device according to an embodiment of the present application. The digital human interaction device provided in this embodiment of the present application includes:

[0129] The question data acquisition module 602 is used to obtain question data to be responded to, and identify the question intention of the question data to obtain the question intention;

[0130] The preset text determination module 604 is used to determine a target preset text that matches the question intention in a pre-built preset response text library;

[0131] The interactive response module 606 is used to obtain target audio data corresponding to the target preset text, so as to use the target audio data to drive the digital human to interactively respond to the question data.

[0132] Optionally, in some embodiments, the interactive response module 606, when executing the step of acquiring target audio data corresponding to the target preset text, is specifically configured to:

[0133] In a pre-built response audio database, target audio data corresponding to the target preset text is determined.

[0134] Optionally, in some embodiments, the interactive response module 606 is further configured to:

[0135] If the target audio data corresponding to the target preset text does not exist in the response audio database, the target audio data corresponding to the target preset text is generated through text-to-speech technology; and the response audio database is updated based on the target audio data.

[0136] Optionally, in some embodiments, the digital human interactive device further comprises:

[0137] The formal text generation module is used to generate a formal response text corresponding to the question data based on the question data and the question intention after obtaining the question data; and generate formal audio data corresponding to the formal response text;

[0138] Correspondingly, when executing the step of using the target audio data to drive the digital human to interactively respond to the question data, the interactive response module 606 is specifically configured to:

[0139] Use target audio data to drive the digital human to interactively respond to question data;

[0140] After the interactive response is completed, the formal audio data is used to drive the digital human to make a formal interactive response to the question data.

[0141] Optionally, in some embodiments, the question intention is obtained through a pre-trained intention recognition model, and the digital human interaction device further includes:

[0142] A training module is used to obtain an initial intent recognition model; obtain a fine-tuning sample set, which includes: fine-tuning question samples and question intent labels; wherein the question intent labels meet the following conditions: the preset response text that matches the question intent label and the semantic correlation between the formal response text of the fine-tuning question sample meet the preset conditions; use the fine-tuning sample set to adjust the parameters of the initial intent recognition model to obtain a trained intent recognition model.

[0143] Optionally, in some embodiments, the digital human interactive device further comprises:

[0144] A fine-tuning sample set construction module is used to obtain fine-tuning question samples and the formal response text corresponding to the fine-tuning question samples; in the preset response text library, the associated preset text that is semantically related to the formal response text of the fine-tuning question sample is determined; based on the matching mapping relationship between the pre-set question intention and the preset response text, the question intention that matches the associated preset text is determined as the question intention label.

[0145] Optionally, in some embodiments, the training module, after obtaining the initial intent recognition model and before obtaining the fine-tuning sample set, is further configured to:

[0146] A test question sample is obtained, and a predicted question intent of the test question sample is obtained based on the initial intent recognition model; based on the matching mapping relationship between the preset question intent and the preset response text, a predicted preset text that matches the predicted question intent is determined from the preset response text library; a semantic association calculation is performed on the predicted preset text and the formal response text corresponding to the test question sample to obtain a semantic association calculation result; if the semantic association calculation result does not meet the preset requirements, the step of obtaining a fine-tuning sample set is executed; if the semantic association calculation result meets the preset requirements, the initial intent recognition model is determined as the trained intent recognition model.

[0147] Optionally, in some embodiments, the question data acquisition module 602, when performing the step of identifying the question intention of the question data to obtain the question intention, is specifically configured to:

[0148] Obtain the character attribute information of the digital person; obtain the intention generation prompt word based on the question data and the character attribute information; input the intention generation prompt word into the pre-trained intention recognition model, and obtain the question intent contained in the question data through the intention recognition model.

[0149] The digital human interactive device of this embodiment is used to implement the corresponding digital human interactive methods of the aforementioned embodiments and has the beneficial effects of the corresponding method embodiments, which will not be described in detail here. Furthermore, the functional implementation of each module in the digital human interactive device of this embodiment can refer to the corresponding descriptions of the aforementioned method embodiments and will not be described in detail here.

[0150] Reference Figure 7 , shows a structural diagram of an electronic device according to an embodiment of the present application. The specific embodiment of the present application does not limit the specific implementation of the electronic device.

[0151] like Figure 7 As shown, the electronic device may include: a processor (processor) 702 , a communication interface (Communication Interface) 704 , a memory (memory) 706 , and a communication bus 708 .

[0152] in:

[0153] The processor 702 , the communication interface 704 , and the memory 706 communicate with each other via a communication bus 708 .

[0154] The communication interface 704 is used to communicate with other electronic devices.

[0155] The processor 702 is used to execute the program 710, and specifically can execute the relevant steps in the above-mentioned embodiment of the digital human interaction method.

[0156] Illustratively, the program 710 may include program code, which includes computer operating instructions.

[0157] Processor 702 may be a CPU, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.

[0158] The memory 706 is used to store the program 710. The memory 706 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0159] The program 710 may include multiple computer instructions. Specifically, the program 710 may enable the processor 702 to execute operations corresponding to the method described in any of the aforementioned method embodiments through the multiple computer instructions.

[0160] The specific implementation of each step in program 710 can refer to the corresponding description of the corresponding steps and units in the above-mentioned method embodiment, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the above-mentioned devices and modules can refer to the corresponding process description in the above-mentioned method embodiment, and will not be repeated here.

[0161] The present application also provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any of the aforementioned method embodiments. The computer storage medium includes, but is not limited to, a compact disc read-only memory (CD-ROM), random access memory (RAM), a floppy disk, a hard disk, or a magneto-optical disk.

[0162] An embodiment of the present application also provides a computer program product, including computer instructions, which instruct a computing device to execute operations corresponding to any one of the above-mentioned multiple method embodiments.

[0163] In addition, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used to train the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0164] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.

[0165] The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded via a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA)). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., random access memory (RAM), read-only memory (ROM), flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown here.

[0166] Those skilled in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of this application.

[0167] The above implementation methods are only used to illustrate the embodiments of the present application, and are not intended to limit the embodiments of the present application. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present application, and the scope of patent protection of the embodiments of the present application should be defined by the claims.

Claims

1. A digital human interaction method, comprising: Acquire question data to be responded to, and perform question intent recognition on the question data to obtain the question intent; Determining a target preset text that matches the question intention in a pre-built preset response text library; Target audio data corresponding to the target preset text is acquired, so as to use the target audio data to drive the digital human to interactively respond to the question data.

2. The method according to claim 1, wherein The acquiring of target audio data corresponding to the target preset text includes: In a pre-built response audio database, target audio data corresponding to the target preset text is determined.

3. The method according to claim 2, wherein: The method further comprises: If the target audio data corresponding to the target preset text does not exist in the response audio database, generating the target audio data corresponding to the target preset text by using text-to-speech technology; The response audio database is updated based on the target audio data.

4. The method according to any one of claims 1 to 3, wherein: After obtaining the question data, the method further includes: Based on the question data, generating a formal response text corresponding to the question data; generating formal audio data corresponding to the formal response text; The step of using the target audio data to drive the digital human to interactively respond to the question data includes: Using the target audio data to drive the digital human to interactively respond to the question data; After the interactive response is completed, the formal audio data is used to drive the digital human to make a formal interactive response to the question data.

5. The method according to claim 1, wherein The question intention is obtained through a pre-trained intention recognition model, and the method further includes: Obtain the initial intent recognition model; Obtain a fine-tuning sample set, the fine-tuning sample set including: a fine-tuning question sample and a question intention label; wherein the question intention label satisfies the following conditions: a preset response text matching the question intention label and a semantic relevance between the formal response text of the fine-tuning question sample and the pre-set condition are satisfied; The fine-tuning sample set is used to adjust the parameters of the initial intent recognition model to obtain a trained intent recognition model.

6. The method according to claim 5, wherein: The process of constructing the fine-tuning sample set includes: Obtaining a sample of a fine-tuned question and a formal response text corresponding to the sample of the fine-tuned question; Determining, in the preset response text library, associated preset texts that are semantically related to the formal response text of the fine-tuning question sample; According to the matching mapping relationship between the preset question intention and the preset response text, the question intention that matches the associated preset text is determined as the question intention label.

7. The method according to claim 5 or 6, wherein: After obtaining the initial intent recognition model and before obtaining the fine-tuning sample set, the method further includes: Acquire a test question sample, and obtain a predicted question intent of the test question sample based on the initial intent recognition model; Determining, from the preset response text library, a predicted preset text that matches the predicted question intention based on a matching mapping relationship between a preset question intention and a preset response text; Calculating semantic relevance between the predicted preset text and the formal response text corresponding to the test question sample to obtain a semantic relevance calculation result; If the semantic association calculation result does not meet the preset requirements, executing the step of obtaining the fine-tuning sample set; If the semantic association calculation result meets the preset requirements, the initial intent recognition model is determined as the trained intent recognition model.

8. The method according to claim 1, wherein The step of identifying the question intention of the question data to obtain the question intention includes: Obtaining character attribute information of the digital human; Based on the question data and the character attribute information, obtaining an intention generation prompt word; The intention generation prompt word is input into a pre-trained intention recognition model, and the question intention contained in the question data is obtained through the intention recognition model.

9. A digital human interactive device, comprising: A question data acquisition module is used to acquire question data to be responded to, and perform question intention recognition on the question data to obtain the question intention; A preset text determination module is used to determine a target preset text that matches the question intention in a pre-built preset response text library; The interactive response module is used to obtain target audio data corresponding to the target preset text, so as to use the target audio data to drive the digital human to interactively respond to the question data.

10. An electronic device comprising: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, where the executable instruction enables the processor to perform an operation corresponding to the method according to any one of claims 1 to 8.

11. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

12. A computer program product comprising computer instructions, wherein the computer instructions instruct a computing device to perform operations corresponding to the method according to any one of claims 1 to 8.