Speech interaction method and apparatus, and device, cluster, storage medium and program product

By combining audio and facial expression data, especially lip-reading data, the problem of digital human interaction systems struggling to accurately identify user questions in noisy environments has been solved, resulting in higher speech recognition accuracy and user experience while reducing hardware costs.

WO2026051429A1PCT designated stage Publication Date: 2026-03-12HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

In high-traffic areas such as exhibition halls and conference centers, the ambient noise is significant, causing errors in how digital human interaction systems understand user questions and affecting the accuracy of their responses. Existing technologies struggle to accurately segment sentences and filter noise.

Method used

By combining audio data and facial expression data, especially lip-reading data, the completeness of user questions is determined. An adaptive threshold algorithm is used to determine whether a question is complete, and the mute and speaking thresholds are dynamically adjusted to reduce reliance on dedicated microphones. A large language model is used to obtain accurate answers.

Benefits of technology

It improves the accuracy of speech recognition in digital human interaction systems, enhances the user interaction experience, reduces hardware costs and latency, and dynamically adjusts thresholds to ensure the best personalized interaction effect for each user.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025095461_12032026_PF_FP_ABST
    Figure CN2025095461_12032026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of artificial intelligence interaction. Disclosed are a speech interaction method and apparatus, and a device, a cluster, a storage medium and a program product. The method comprises: determining an interaction object; acquiring facial expression data and audio data of the interaction object in real time, wherein the facial expression data at least comprises mouth shape data; on the basis of the audio data and the mouth shape data, determining the completeness of a question proposed by the interaction object; at a first moment, determining, on the basis of the audio data and the mouth shape data, that the question is complete, and determining question information of the interaction object up to the first moment; and determining a response corresponding to the question, and generating driving information on the basis of the response; and on the basis of the driving information, outputting media data corresponding to the response. According to the method in the embodiments of the present application, the accuracy of a digital human's responses to questions can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Voice interaction method, device, equipment, cluster, storage medium and program product

[0001] The present application claims priority to the Chinese patent application No. 202411245106.8, filed on September 4, 2024, and entitled "Voice interaction method, device, equipment, cluster, storage medium and program product", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the field of artificial intelligence interaction, in particular to a voice interaction method, device, equipment, cluster, storage medium and program product. BACKGROUND

[0003] With the development of artificial intelligence (AI) technology, AI digital people have gradually become an important part of people's daily life. AI digital people interactive large screens, as a new way of interaction, have been widely used in various occasions, such as exhibition halls, conference centers, enterprises, schools, etc. These digital people interact with the audience in real time through holographic projection technology, providing information transmission, entertainment, education and other services.

[0004] Generally, an AI digital people interactive system is generally composed of audio acquisition devices such as microphones of large screens or separate Bluetooth microphones, large screen devices equipped with cameras, digital people generation services on the cloud, voice recognition services, voice generation services, language model services, etc. The user inputs a wake-up word through the audio acquisition device, wakes up the digital people, and then asks questions, waits for the digital people to analyze and understand the questions, and gives an audio answer and corresponding actions, realizing the experience of interacting with real people.

[0005] However, in places with large flow of people, such as exhibition halls, conference centers, etc., the environmental noise in these places is often large, which makes it necessary to use a special microphone with strong noise reduction function when communicating with digital people in order to improve the accuracy of voice recognition. In addition, due to the differences in the way each person speaks, including speed, volume and pause, etc., these factors increase the difficulty of correct sentence breaking and noise filtering of the device to the user's input sentence in actual application. Therefore, this may cause errors in the understanding of the user's question by the digital people, and further affect the accuracy of the answer. SUMMARY

[0006] To solve the above problems, the present application provides a voice interaction method, device, equipment, cluster, storage medium and program product to solve the problem that electronic devices cannot correctly break sentences and filter noise, resulting in inaccurate answers to user questions.

[0007] In a first aspect, the present application provides a voice interaction method, comprising: determining an interaction object; obtaining facial expression data and audio data of the interaction object, the facial expression data at least including mouth shape data; determining completeness of a question raised by the interaction object based on the audio data and the mouth shape data; if it is determined that the question is complete based on the audio data and the mouth shape data, determining a response corresponding to the question, and generating driving information based on the response; and outputting media data corresponding to the response based on the driving information, such as audio and / or driving a digital person to display animation corresponding to the response.

[0008] The voice interaction method according to the embodiments of the present application adopts the combination of audio data and user facial expression data, which is more conducive to accurately analyzing the question and intention input by the user, and can better determine whether the question raised by the interaction object is complete or the question is complete, thereby improving the experience of digital person large screen interaction. In addition, in the interaction process, the use of a dedicated microphone for input is no longer needed, which can avoid incomplete sentences caused by improper operation and reduce hardware costs.

[0009] As an embodiment of the first aspect of the present application, determining that the question is complete based on the audio data and the mouth shape data comprises: determining that the mouth shape data of the interaction object is less than or equal to a mouth shape speaking threshold value, and the signal value of the audio data is greater than an audio mute threshold value, then determining that the question is complete. Combining the mouth shape speaking threshold value with the audio data can improve the accuracy of the judgment of the completeness of the question, and can better identify noise.

[0010] As an embodiment of the first aspect of the present application, determining that the question is complete based on the audio data and the mouth shape data further comprises: noise reduction processing of the audio data to ensure that the identified question is more accurate.

[0011] As an embodiment of the first aspect of the present application, determining that the question is complete based on the audio data and the mouth shape data comprises: when it is determined that the mouth shape data of the interaction object is less than or equal to a mouth shape speaking threshold value, and the signal value of the audio data is reduced to be less than or equal to an audio mute threshold value, then determining that the question is complete. This judgment method not only can accurately identify the completeness of the question, but also can analyze whether there is noise and timely decide the processing method.

[0012] As an embodiment of the first aspect of the present application, determining the completeness of the question raised by the interaction object based on the audio data and the mouth shape data comprises: if it is determined that the question is incomplete, and no response corresponding to the incomplete question is output within a preset time length, maintaining the state of picking up data. In this way, it can be avoided that the question is answered inaccurately because the question is not answered completely.

[0013] As an embodiment of the first aspect of the present application, the state of maintaining the pickup data comprises: outputting prompt information prompting that the interactive object is picking up data or maintaining the state of the digital human listening, and continuing to pick up the audio data and facial expression data of the interactive object. The user can understand how to operate based on the state of the digital human.

[0014] As an embodiment of the first aspect of the present application, determining that the question is incomplete comprises: determining that the mouth shape data of the interactive object is greater than the mouth shape speaking threshold, and the signal value of the audio data is reduced to less than or equal to the audio mute threshold, or determining that the mouth shape data of the interactive object is greater than the mouth shape speaking threshold, and the signal value of the audio data is greater than the audio mute threshold.

[0015] As an embodiment of the first aspect of the present application, determining that the question is incomplete further comprises: prompting the interactive object to adjust the volume or approach the sound pickup device to ensure that a complete sentence is obtained.

[0016] As an embodiment of the first aspect of the present application, determining the answer corresponding to the question comprises: obtaining the audio data of the complete question; and obtaining the answer corresponding to the question based on the audio data of the question and the large language model. Using a large language model or a knowledge base can obtain a more accurate answer.

[0017] As an embodiment of the first aspect of the present application, obtaining the answer corresponding to the question based on the audio data of the question and the large language model comprises: converting the audio data of the question into text; inputting the text into the large language model to obtain the answer text corresponding to the text based on the large language model; and converting the answer text into an audio answer.

[0018] As an embodiment of the first aspect of the present application, obtaining the answer corresponding to the question based on the audio data of the question and the large language model further comprises: outputting a prompt whether the answer conforms to the intention of the interaction; if a first operation conforming to the intention of the interactive object is received, outputting audio corresponding to the answer and / or driving the digital human to display animation corresponding to the answer; and if a second operation not conforming to the intention of the interactive object is received, adjusting the mouth shape speaking threshold and the audio mute threshold, and prompting the interactive object to ask the question again until the answer conforms to the intention of the interactive object. The accuracy of the question can be confirmed by interacting with the user, and the audio mute threshold and the mouth shape speaking threshold can be dynamically adjusted to match the user, thereby further improving the accuracy of answering the question.

[0019] As an embodiment of the first aspect of the present application, the first operation comprises: a pressing operation on a specific function key, a movement in a specific posture, or inputting a keyword representing conformity to the intention; and the second operation comprises: a pressing operation on a specific function key, a movement in a specific posture, or inputting a keyword representing non-conformity to the intention. These operations can identify the intention of the user.

[0020] As an embodiment of the first aspect of the application, the determining the interaction object comprises: receiving a wake-up word input by the user, and acquiring mouth shape data of the user; and determining, from the user, an interaction object whose wake-up word and mouth shape data match based on the wake-up word and the mouth shape data. The wake-up word and the mouth shape data are combined to confirm the interaction object, which can accurately identify the interaction object, and does not require a dedicated voice collection device, thereby reducing hardware costs and improving user interaction experience.

[0021] As an embodiment of the first aspect of the application, the method further comprises: extracting a voiceprint feature in audio data of the interaction object, and storing the voiceprint feature; and when it is determined that the interaction object is not asking the question for the first time, directly acquiring a mouth shape speaking threshold and an audio silence threshold corresponding to the interaction object based on the voiceprint feature. In this way, when the interaction object asks the question again, the mouth shape speaking threshold and the audio silence threshold that match the interaction object can be quickly obtained, so that an accurate answer can be quickly obtained.

[0022] As an embodiment of the first aspect of the application, the facial expression data further comprises one or more of eyebrow movement data, eye movement data, and nose movement data.

[0023] As an embodiment of the first aspect of the application, the media data comprises animation data of a digital human or a specific shaped object in a display screen, and output audio data.

[0024] In a second aspect, the application further provides a voice interaction device, comprising:

[0025] An object confirmation module configured to determine an interaction object;

[0026] An acquisition module configured to acquire facial expression data and audio data of the interaction object, wherein the facial expression data at least comprises mouth shape data;

[0027] An integrity confirmation module configured to determine the integrity of a question asked by the interaction object based on the audio data and the mouth shape data;

[0028] A generation module configured to, if it is determined that the question is complete based on the audio data and the mouth shape data, determine a response corresponding to the question, and generate driving information based on the response;

[0029] An output module configured to output media data corresponding to the response based on the driving information.

[0030] In a third aspect, the application further provides an electronic device, comprising: a memory configured to store instructions executed by one or more processors of the electronic device, and a processor configured to execute the voice interaction method explained in the first aspect.

[0031] In a fourth aspect, the present application provides a computing device cluster, comprising at least one computing device, each computing device comprising a processor and a memory;

[0032] The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the voice interaction method explained in the first aspect.

[0033] In a fifth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores instructions, and the instructions, when executed on an electronic device, cause the electronic device to perform the voice interaction method explained in the first aspect.

[0034] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when executed on an electronic device, cause the electronic device to perform the voice interaction method explained in the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0035] FIG. 1 is a scenario diagram of a user interacting with a digital person according to an embodiment of the present application;

[0036] FIG. 2 is a data flow diagram of a user interacting with a digital person in some embodiments;

[0037] FIG. 3 is a flow diagram of a user interacting with a digital person large screen device in some embodiments;

[0038] FIG. 4 is a flow diagram of a voice interaction method based on a digital person according to an embodiment of the present application;

[0039] FIG. 5 is a structural diagram of an electronic device according to an embodiment of the present application;

[0040] FIG. 6 is a flow diagram of a voice interaction method according to another embodiment of the present application;

[0041] FIG. 7 is a structural diagram of software modules of a large screen device and a cloud device according to an embodiment of the present application;

[0042] FIG. 8 is a flow diagram of interaction between software modules according to an embodiment of the present application;

[0043] FIG. 9 is a flow diagram of adjusting a preset value according to a detected mouth shape and audio according to an embodiment of the present application;

[0044] FIG. 10 is a structural diagram of a voice interaction apparatus according to an embodiment of the present application;

[0045] FIG. 11 is a structural diagram of a computing device cluster according to an embodiment of the present application;

[0046] FIG. 12 is an interaction diagram between computing device clusters according to an embodiment of the present application;

[0047] FIG. 13 is a schematic diagram of a system on chip according to an embodiment of the present application. DETAILED DESCRIPTION

[0048] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application.

[0049] The illustrative embodiments of the present application include, but are not limited to, a digital person-based voice interaction method, device, equipment, cluster, storage medium and program product.

[0050] In order to facilitate the understanding of the technical solutions of the present application, first, the technical problems to be solved by the present application will be described.

[0051] Referring to FIG. 1, FIG. 1 shows a scene diagram of user interaction with a digital person according to an embodiment of the present application. As shown in FIG. 1, the scene includes a holographic projection large screen device 02 and a cloud device 01 in communication connection with the large screen device. The digital person 03 (media data) is displayed in the large screen device, and the digital person 03 is talking with the user 04. The digital person 03 is a digital virtual image, which is a software system created by using artificial intelligence technology. The digitalized character image close to the human image can perform specific actions, such as imitating the speaking action of a real person, etc. When the user 04 interacts with the digital person 03, the audio of the user is collected by the hardware of the large screen device 02, and then the audio is processed by the audio processing software to become text information that can be recognized by the machine, and the response output for the user 04 interaction is provided based on the recognized text information, for example, the answer to the question raised by the user. The large screen device 02 generates the corresponding voice of the response based on the output, and outputs the corresponding action of the digital person. Through the rich output of the digital person 03, the user can experience the effect close to talking with a real person. For example, the large screen device 02 can include audio collection devices such as microphones, image collection cameras and other hardware, and can also include software services: artificial intelligence generated content (AIGC), i.e. digital person generation service, speech recognition (ASR) service, speech generation (TTS) service, etc. The cloud device 01 can include language large model service, knowledge base system and other services. The user can wake up the digital person through the audio collection device and ask questions, the large screen device 02 uploads the questions to the cloud device 01, and analyzes and understands the questions through the software service system of the cloud device 01, and outputs the audio and action, thereby realizing the experience of interacting with a real person.

[0052] In the above embodiments, the AIGC digital human generation service, the ASR speech recognition service, and the TTS speech generation service are deployed on the large-screen device. In other embodiments, these software modules can also be deployed on an edge computing node, such as a server or gateway close to the large-screen device, to reduce network transmission delay. In addition, these software or part of the services can also be deployed in the cloud, which is not limited in the present application.

[0053] In other embodiments, the language large model service and the knowledge base system can also be deployed on the large-screen device on the edge side.

[0054] It should be noted that in the embodiments of the present application, the digital human is used as the media data. In some embodiments, the media data can also be other objects, such as animals, or robots or objects of specific shapes, which are not limited in the present application.

[0055] Referring to FIG. 2, FIG. 2 shows a data flow diagram of user interaction with a digital human in some embodiments. As shown in FIG. 2, the process of the questioner interacting with the digital human includes the following steps: ① Live questioning. The user can first wake up the digital human, and then ask the digital human using a special microphone with noise reduction function. In a conference center or exhibition hall, a person on the side is usually needed to guide to avoid multiple questioning. ② The digital human large-screen device collects audio through a special microphone. ③ ASR speech recognition. The ASR automatic speech recognition module processes the audio stream, i.e., the input speech is converted into text by the ASR module, and whether the input speech is a complete sentence is detected according to a fixed silence detection time (such as 500 ms). When a complete sentence is recognized, the text is input into a language large model, such as the Pangu large model or a knowledge base question answering system, etc. ④ The language large model (or large language model) or the knowledge base question answering system generates an answer (response), and outputs the answer to the text-to-speech module. ⑤ TTS speech synthesis. The TTS converts the text of the answer into corresponding speech and sends it to the digital human generation service. ⑥ The digital human generation service generates corresponding driving parameters and sends the driving parameters and the response corresponding to the question to the processor of the large-screen device. The processor of the large-screen device drives the digital human to make actions based on the driving parameters, and controls the loudspeaker to play the audio, so that the user can listen to the audio and watch the digital human animation, realizing the experience of interacting with a real person.

[0056] In the above, during the interaction between the user and the digital person, the questioner needs to hold the microphone, and at the same time, human intervention is needed to interfere with the surrounding people, and the user experience is not good. And for the judgment of whether the user's input voice is a complete sentence, it is easy to be disturbed by the user's speaking habits to judge whether the question is finished only from the silence detection. For example, for people who speak slowly or are prone to stuttering during speaking, it is easy to exceed the default threshold of silence detection, resulting in incomplete questions, and the output answer may not be the answer the user wants. In some embodiments, in order to solve the problem of low accuracy of the judgment of the completeness of the question input by the questioner mentioned above, a step of judgment by a language large model is added. For example, after the ASR module speech recognition described in ③ASR speech recognition, the recognized text is first sent to the language large model, and the language large model performs a completeness judgment on the question. If the language large model identifies that the question is not finished, the ASR module continues to process. However, if the language large model is preprocessed every time to judge the completeness, the delay time will be increased, usually >1s, which will affect the user's interactive experience.

[0057] The process of user interaction with the digital person in the above-described embodiments will be further described below in combination with the hardware and software structure of the large-screen device.

[0058] Referring to FIG. 3, FIG. 3 shows a flowchart of the user interacting with the digital person large-screen device in some embodiments. As shown in FIG. 3, it includes a large-screen device and a cloud device. The large-screen device includes dedicated microphones and other hardware, as well as ASR modules, digital person / TTS modules and other software modules, and the cloud device includes language large models (such as the Discuz large model / knowledge base question answering system) and the like. The interaction process shown in FIG. 3 corresponds to the interaction process of steps ①-⑥ in FIG. 2. The interaction process mainly includes two parts, the first part is to wake up the digital person, and the second part is to identify whether the user's question is complete and give an answer.

[0059] The first part, the process of waking up the digital person, includes S10-S13, specifically including: S10, local sound pickup, noise reduction. The sound pickup is the process of obtaining audio data, for example, the user inputs the wake-up word through the dedicated microphone. The dedicated microphone itself can perform noise reduction processing on the audio to ensure the accuracy of the wake-up word. S11, wake-up detection. To determine the accuracy of the wake-up word in the process of multiple people speaking, text detection can be used to determine whether the wake-up word is accurate. S12, the audio module corresponding to the microphone sends a wake-up instruction to the digital person / TTS. For example, after determining the wake-up word, the audio module sends the wake-up word to the digital person / TTS to trigger the digital person to act. S13, the digital person / TTS returns a welcome speech and prepares to answer the user's question.

[0060] The second part is to identify whether the user's question is complete, including two embodiments. The difference between the two embodiments is that the judgment process of the completeness of the user's question is different after the user inputs the voice of the question. As shown in FIG. 3, the first embodiment (solution one) includes steps S14-S18, and the second embodiment (solution two) includes steps S14, S19-S25.

[0061] The interactive processes of the above two embodiments are described below respectively.

[0062] First, solution one is described, as shown in FIG. 3, S14, the audio module corresponding to the microphone sends the question to the ASR. After the digital human is woken up, the user inputs the audio data to the digital human through the special microphone. S15, the ASR module performs the threshold mute detection. The ASR module obtains the audio data (question) input by the user, recognizes the voice and converts it into text, and identifies the audio signal. If the audio signal value in the audio signal is less than or equal to the threshold value, such as 500 ms (threshold mute detection), it is determined that the question has been completed (the sentence has been completed). S16, the ASR module notifies the language large model that the question is complete. The language large model searches based on the question and finds the answer corresponding to the question. S17, the language large model returns the answer to the digital human / TTS module. The TTS module converts the text into voice, and the digital human service module generates the corresponding driving parameters. S18, the digital human answers. The large-screen device can drive the digital human animation through the driving parameters, and output the answer audio corresponding to the question through the loudspeaker. In this interactive process, the person who asks the question needs to hold the microphone, and also needs to manually intervene the interference of the surrounding people, which is not good for the user experience. Moreover, the accuracy of the threshold mute detection method used by the ASR module to judge the completeness of the sentence or the completion of the question is low, which leads to low accuracy of the answered question.

[0063] Scheme two, as shown in FIG. 3, S14, voice question. The user asks the digital person through a dedicated microphone. After the user voice question, the large screen device sends the content of the question to the ASR module. Different from scheme one, S19, the ASR module sends the voice to the language large model, and the language large model judges whether the question is completed. S20, the language large model predicts that the question is completed. That is, the language large model can judge the probability that the user may complete the question based on the input voice, and when the probability exceeds a set value, such as 90%, it is determined that the question is completed. S21, the language large model sends the recognition result to the ASR module. S22, the ASR module continues to detect or returns the text. The ASR module can further perform a configured threshold mute detection, after determining that the question is completed, converting the audio to text, and performing S23, sending the result of recognizing the completion of the question to the language large model / knowledge base system. S24, the language large model / knowledge base system returns the answer to the digital person / TTS module. The TTS module converts the text into voice, and the digital person service module generates the corresponding driving parameters, which is the same as S17. S25, the digital person / TT drives the digital person to answer. The large screen device can drive the digital person animation through the driving parameters, and output the corresponding answer audio through the loudspeaker, which is the same as S18. In this interaction process, although the completeness of the recognized sentence or the accuracy of the completion of the question is improved, the ASR recognition result needs to be judged by the large model whether the question is completed each time, which consumes a lot of time and increases the delay obviously, and reduces the interaction experience.

[0064] Based on the above problems, the present application provides a voice interaction method, which adds mouth shape detection on the basis of mute detection, and through an adaptive algorithm process, the parameters of the two are optimally matched without manual adjustment, so as to more accurately judge the completion of the question and ultimately improve the experience of interaction with the digital person. In addition, the method of the embodiment of the present application can accurately identify the questioner (interaction object), reduce the dependence on the sampling microphone, reduce the possibility of inaccurate recognition caused by improper operation, and reduce the hardware cost.

[0065] The voice interaction method of the embodiment of the present application will be described below with reference to the accompanying drawings.

[0066] Referring to FIG. 4, FIG. 4 shows a flowchart of the voice interaction method of the embodiment of the present application. The method can be executed by a large screen device. In combination with FIG. 4, the interaction method includes S410-S460.

[0067] S410, determine the interaction object. The interaction object is the questioner or questioner determined from the user in front of the large screen based on audio and facial expression data.

[0068] For example, in the confirmation process of the interactive object, the large-screen device can compare the voiceprint input by the user with the stored voiceprint feature, and when it is determined that it is a user with an existing voiceprint, the current interactive object can be directly determined according to the voiceprint feature. In addition, it can also be determined whether it is the object currently speaking by combining the keywords of the voice and the mouth shape of the facial expression.

[0069] In some embodiments, other ways can also be used to determine the interactive object. For example, the user performs a specific action, for example, the user's specific gesture, which can be used for the determination of the interactive object. For another example, a dedicated hardware device can also be used to determine the interactive object, for example, a microphone matched with the large screen, and the user holding the microphone means the determined interactive object. The embodiments of the present application are not limited thereto.

[0070] S420, obtaining facial expression data and audio data of the interactive object.

[0071] After the interactive object is determined, the facial expression data of the interactive object can be directly obtained through the image acquisition device. For example, the facial image of the user is captured through the camera configured on the large-screen device, including the acquisition of the facial key points of the interactive object, such as the points around the eyes, nose, and mouth. The position change of a single key point or the combination of the changes of multiple key points can reflect the change of the expression. For example, the state of the continuously opening and closing of the mouth indicates that the user is speaking, etc. The change of the key point or the combination of the key points can highlight the state when speaking. The audio data can capture the voice signal of the interactive object, and in combination with the facial expression data, it is more conducive to accurately determine whether the question of the interactive object is completed.

[0072] S430, determining whether the question is complete. If, at T1 (the first time), it is determined that the question is complete, S440 is executed. If not, S420 is returned to continue to obtain the facial expression data and audio data of the interactive object until it is determined that the question is complete.

[0073] S440, confirming the corresponding question information based on the information received before T1.

[0074] For example, the time when the interactive object inputs the audio, that is, the time when the question is started, is recorded as T0, and it is determined that the question of the interactive object is complete at T1. The audio input by the interactive object in the period from T0 to T1 is obtained as the audio data of the complete question. Further, based on the recognition result of the audio information, the answer to the question information is determined.

[0075] S450, determining the corresponding answer according to the question information and generating driving information.

[0076] In some embodiments, the question information is sent to the cloud, and relevant answers are searched based on a language large model or a knowledge base system, or a combination of both. The specific process and algorithm for obtaining answers can refer to the prior art, and will not be described in detail in the embodiments of the present application.

[0077] After determining the answer, the large-screen device can generate driving information corresponding to the answer, for example, determining the action sequence that the digital person should perform, such as nodding, waving, mouth opening and closing, etc., and audio playing instructions, etc.

[0078] S460, output audio and digital person animation based on the driving information.

[0079] For example, the large-screen device controls the digital person on the large-screen display to output a waving action based on the driving parameters of the waving, and outputs the answer audio corresponding to the question based on the playing instructions.

[0080] According to the voice interaction method of the embodiments of the present application, the audio data and the user facial expression data are combined, which is more conducive to accurately analyzing the user's input question and intent, and can better determine whether the questioning of the interactive object is complete (whether the question is complete), thereby improving the experience of digital person large-screen interaction. In addition, in this interaction process, it is no longer necessary to rely on a dedicated microphone for input, which can avoid incomplete sentences caused by improper operation and reduce hardware costs.

[0081] The voice interaction method of the embodiments of the present application will be described below in combination with the specific structure of the electronic device.

[0082] Referring to FIG. 5, FIG. 5 shows a structural schematic diagram of the electronic device 100 of the present application. The structure of the electronic device can be the structural schematic diagram of the master device as shown in FIG. 1, or the structural schematic diagram of other slave devices.

[0083] The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) terminal 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyro sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0084] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than illustrated, or combine certain components, or split certain components, or different arrangement of components. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0085] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated in one or more processors.

[0086] The processor 110 can generate operation control signals according to instruction opcodes and timing signals, and complete the control of fetching and executing instructions.

[0087] The processor 110 can also include memory that stores instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can hold instructions or data that the processor 110 has recently used or is likely to use again. If the processor 110 needs to use that instruction or data again, it can be retrieved directly from the memory. This avoids repeated accesses to the main memory, reducing the processor's 110 latency and improving the system's efficiency.

[0088] In some embodiments, the processor 110 receives audio of the user and determines an interaction object, e.g., determines the user's identity using voiceprint features, image features, or sensor data. Facial expression data and audio data of the interaction object are obtained in real time, where the facial expression data can include mouth shape data. Based on the audio data and the mouth shape data, the completeness of a question proposed by the interaction object is determined, e.g., using an adaptive threshold algorithm for audio silence detection and mouth shape silence detection, using existing mouth shape speaking threshold and audio silence threshold, and comparing with the audio data and facial expression data input by the interaction object to determine whether the current question of the interaction object is complete or the question is complete. When T1 (the first time), it is determined that the question is complete, and the question information from the start of the user's question to T1 is obtained. The processor 110 sends the question information to the cloud server, and the cloud server determines the response corresponding to the question based on the question information, and generates driving information based on the response. The processor 110 receives the driving information and controls the loudspeaker 170A to output audio corresponding to the response, and / or controls the digital person of the display screen 194 to display animation corresponding to the response.

[0089] In some embodiments, after obtaining the answer to the question, the processor 110 further outputs confirmation information, for example, asking the interactive object through voice whether the answer is the desired answer. Or displaying the answer on the display screen 194, and confirming whether the answer is accurate by the interactive object. If the interactive object confirms that it is correct, the processor 110 stores the voiceprint information of the interactive object, and the audio mute threshold and the mouth speech threshold corresponding to the interactive object. If the interactive object confirms that it is not correct, the processor 110 controls the audio module 170 to output a prompt to re-input the question, or controls the display screen 194 to display a prompt to re-input the question. The processor 110 dynamically adjusts the audio mute threshold by re-obtaining the audio data and the mouth shape data input by the interactive object. The adjustment of the audio mute threshold refers to the threshold of the duration of silence when the user finishes speaking, and the adjustment of the mouth speech threshold refers to the threshold of the duration of the mouth shape when the user finishes speaking, until the interactive object determines that the answer is accurate. The adjusted audio mute threshold and the mouth speech threshold are stored with the voiceprint feature or the facial feature of the interactive object, so that the corresponding threshold can be directly obtained when the interactive object is asked again, and the response speed is improved. In addition, by dynamically adjusting the audio mute threshold and the mouth speech threshold, each user can be accurately matched, the accuracy of the answer is improved, and the interactive experience is improved.

[0090] The wireless communication function of the electronic device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.

[0091] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: the antenna 1 can be multiplexed as a diversity antenna of a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0092] The mobile communication module 150 can provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transfer the processed signals to the modem processor for demodulation. The mobile communication module 150 can also amplify signals modulated by the modem processor, and radiate the signals as electromagnetic waves through the antenna 1. In some embodiments, at least part of the functional modules of the mobile communication module 150 can be disposed in the processor 110. In some embodiments, at least part of the functional modules of the mobile communication module 150 can be disposed in the same device as at least part of the modules of the processor 110.

[0093] The wireless communication module 160 can provide a solution for wireless communication including wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) network), bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrated with at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering on the electromagnetic wave signals, and transmits the processed signals to the processor 110. The wireless communication module 160 can also receive signals to be transmitted from the processor 110, perform frequency modulation and amplification on the signals, and radiate the signals as electromagnetic waves through the antenna 2.

[0094] In some embodiments, both the mobile communication module 150 and the wireless communication module 160 can enable the electronic device 100 to communicate with the cloud. The processor 110 can control the mobile communication module to transmit the question of the interactive object to the cloud, and the cloud can analyze and understand the question and return an answer.

[0095] The electronic device 100 implements a display function through a GPU, a display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, connecting the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.

[0096] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 100 can include 1 or N display screens 194, N being a positive integer greater than 1.

[0097] In some embodiments, the display screen 194 can display a digital person and information related to the answer or question. The processor 110 can control the digital person on the display screen 194 to make corresponding actions. For example, when the digital person greets the user, the digital person can make a hand waving action. Or, when the electronic device outputs the answer of the interactive object question, the digital person's mouth moves up and down, as if having a real conversation with the user.

[0098] The camera 193 is used to capture still images or videos. An object generates an optical image through a lens and projects it to a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to an ISP to convert it into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into a standard RGB, YUV, etc. format image signal. In some embodiments, the electronic device 100 can include 1 or N cameras 193, N being a positive integer greater than 1.

[0099] In some embodiments, the camera 193 can acquire facial expression data of the interaction object, and body movements of the interaction object. Video data of the entire exhibition hall environment and characters, etc. can also be acquired.

[0100] The electronic device 100 can realize audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor, etc. For example, the microphone can acquire audio data input by the user, etc., and the speaker 170A can output response audio corresponding to the question.

[0101] The voice interaction method provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The method can be applied to an electronic device with a hardware structure as shown in FIG. 5. Or more or less components than the figure, or combination of certain components, or splitting of certain components, or different component arrangement, etc. similar hardware structure and software structure of electronic device.

[0102] The voice interaction method provided by the embodiments of the present application can be applied to an electronic device. For example, the electronic device can be a large-screen device, a smart TV, a vehicle-mounted device, or a mobile phone, a tablet computer, a wearable device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a personal digital assistant (PDA), an artificial intelligence device, etc. The embodiments of the present application do not make any limitation on the specific type of electronic device. In the following embodiments, the electronic device is taken as an example of a large screen.

[0103] Referring to FIG. 6, FIG. 6 shows a flowchart of a voice interaction method of another embodiment of the present application, which is executed by a large-screen device, including S601-S608.

[0104] S601, local wake-up, that is, acquiring the wake-up voice and mouth shape data of the user to wake up the digital person.

[0105] The wake-up voice can be acquired by the microphone of the large-screen device (also called local wake-up). In some embodiments, the wake-up voice can be a specific wake-up word, such as calling the name of the digital person, etc.

[0106] The mouth shape data can be captured by the camera of the large-screen device, for example, standing in front of the large-screen device. The facial expression data of these people will be acquired by the camera.

[0107] S602, match the mouth shape and the wake-up word, and identify the speaker.

[0108] When it is determined that the wake-up word is accurate, the digital person can be started, that is, the digital person is woken up. The wake-up word has a corresponding threshold value of mouth shape opening or closing, for example, the distance between the upper lip and the lower lip. The threshold value is compared with the obtained mouth shape changes of multiple users. The user whose mouth shape change matches the threshold value is the interactive object (the current speaker), so as to complete the speaker identification. In this process, the wake-up word is combined with the mouth shape, so that the interactive object can be accurately determined. Compared with the prior art of using only the keyword to identify the interactive object by audio, the interactive object can be more accurately identified. Moreover, such a determination method can not need a dedicated microphone to ask questions, reducing the dependence on hardware devices and reducing hardware costs.

[0109] S603, obtain audio data and mouth shape data input by the interactive object.

[0110] For example, the interactive object asks the digital person a question, and the microphone of the large-screen device obtains the audio data. At the same time, the camera captures the mouth shape data of the interactive object when speaking, such as the image of the state of the lips opening and closing.

[0111] S604, determine the relationship between the current input audio data and the threshold value. Hereinafter, the audio data is referred to as audio.

[0112] For example, the audio is compared with the audio mute threshold value, and the mouth shape action is compared with the mouth shape speaking threshold value. Based on the comparison result, it is determined whether the question of the interactive object is completed or the question is complete.

[0113] For example, when the silence is identified, it can be preliminarily determined that the question input by the user is completed.

[0114] The audio mute threshold of the embodiments of the present application can use a threshold level (Threshold Level), for example, -60 dBFS (Full Scale), and audio below this level will be considered as mute. In addition, in some embodiments, energy threshold (Energy Threshold), zero crossing rate (Zero Crossing Rate), short-time average magnitude (Short-Time Average Magnitude), and short-time average absolute value (Short-Time Average Absolute Value) can also be used as audio mute threshold, and the embodiments of the present application do not limit this.

[0115] The process of comparing mouth shape action with mouth shape speech threshold adopts mouth shape detection method, which can identify face from image, judge whether mouth shape is speaking, and extract lip shape change feature, which can be input into lip reading model to identify voice pronunciation and other functions. Thus, the lip shape of the wake-up word can be identified, and the interactive object can be determined. In addition, combined with audio mute detection, the interactive object and the question raised by the interactive object can be accurately judged. The following describes several results of judging the current input audio and mouth shape and threshold.

[0116] The mouth shape speech threshold can use the mouth opening degree (Mouth Opening Degree) to measure the degree of mouth opening, for example, the distance between the upper lip and the lower lip, such as 1 cm. When the mouth opening degree exceeds 1 cm, the system considers that the user is speaking, otherwise it considers that the user is not speaking. In addition, in some embodiments, the mouth shape change rate (Mouth Shape Change Rate) can also be used, for example, when the mouth shape change rate is lower than 0.3 times per second, it is considered that the user is not speaking. In some other embodiments, mouth feature vector (Mouth Feature Vector), mouth shape duration (Mouth Shape Duration), and mouth shape similarity (Mouth Shape Similarity) can also be used as mouth shape speech threshold. The setting of these thresholds can be adjusted according to the actual situation to achieve the best recognition effect.

[0117] In some embodiments, in addition to the mouth shape data described above, the audio data can also be combined with other facial expression data, for example, based on the identification of lip opening and closing (mouth shape data), the action of the eyebrows can also be combined to improve the accuracy of determining the question, for example, raising the eyebrows indicates surprise, curiosity or inquiry. Frowning: expressing confusion, disagreement or doubt. In addition, in other embodiments, the facial expression data can also include the action of the cheeks, that is, the audio data is combined with the mouth shape data and the action of the cheeks to determine the accuracy of the question, for example, puffed cheeks may indicate dissatisfaction or anger, relaxed cheeks indicate relaxation or happiness, etc. The combined use of these facial expression data can more accurately identify the user's intention.

[0118] As shown in FIG. 6, the subsequent processing in the step S604 will be based on four different determination results. The four determination results are as follows:

[0119] First, the audio is greater than the audio mute threshold, and the mouth shape action is less than or equal to the mouth shape speaking threshold, it is determined that the question of the interactive object has been completed. In this case, since the mouth shape action understanding is already in the state of stopping speaking, but the audio sound is still very loud, it is judged that there is interference noise, and the audio data is processed for noise reduction. Among them, the specific noise reduction processing method can refer to the existing audio noise reduction method, which will not be described in detail here. After the audio data is noise reduced, the ASR module identifies the audio data to convert the audio to text.

[0120] Second, the audio is less than or equal to the audio mute threshold, and the mouth shape action is greater than the mouth shape speaking threshold, indicating that the interactive object is still speaking, then it is determined that the question of the interactive object is not completed, the question is incomplete, then the interactive object is prompted to adjust the loudness of the sound pickup device, and the user is prompted to approach the digital person or to increase the volume, etc. in order to clearly pick up the audio data of the interactive object, and keep the state of continuing to pick up until it is determined that the question is complete.

[0121] Third, the audio is less than or equal to the audio mute threshold, and the mouth shape action is less than or equal to the mouth shape speaking threshold, indicating that the interactive object has stopped speaking, it is determined that the question of the interactive object has been completed, and the audio data is directly identified by the ASR to convert the audio to text.

[0122] Fourth, the audio is greater than the audio mute threshold, and the mouth shape action is greater than the mouth shape speaking threshold, indicating that the interactive object is still speaking, then the state of picking up is maintained until it is determined that the question of the interactive object has been completed.

[0123] S605, the language large model performs semantic understanding on the text, and confirms the question and intention with the interactive object through the digital person.

[0124] For example, when it is determined that the questioning of the interactive object is completed, the large-screen device sends the text corresponding to the question to the language large model in the cloud, the language large model understands and analyzes the text, and gives the corresponding answer. When the answer is returned to the large-screen device, it is determined whether the answer can accurately answer the question of the interactive object by asking the user. If the user confirms that the answer is accurate, it indicates that the answer can express its own intention. If the interactive object confirms that the answer is not accurate, it indicates that the answer cannot express the intention of the interactive object, and the question needs to be asked again. The specific scheme for obtaining the answer to the question based on the language large model or the knowledge base can refer to the prior art, and the present application does not make a detailed description.

[0125] S606, determining whether to adjust the mouth shape speaking threshold and the audio mute threshold.

[0126] When the large-screen device receives the negative indication input by the interactive object, that is, the interactive object indicates that the answer provided by the digital human is not accurate, it is determined that the audio mute threshold and the mouth shape speaking threshold need to be adjusted, and S607 is performed.

[0127] S607, adjusting the audio mute threshold and the mouth shape speaking threshold.

[0128] For example, when the interactive object first interacts with the digital human, the large-screen device will use the pre-stored audio mute threshold and mouth shape speaking threshold as the initial value for judgment. At this time, the threshold has not been verified, and the threshold is not personalized with respect to the interactive object. Therefore, when giving the answer to the question of the interactive object, it can be further determined whether the answer is accurate. And when the reply of the interactive object is inaccurate, that is, the interactive object performs the corresponding negative action (second operation), for example, the interactive object replies with the words "inaccurate" or "wrong", or clicks the special function key, for example, the interactive object clicks the "no" function key option set on the screen, or a specific gesture, such as waving the hand, shaking the head, etc., to determine that the answer is inaccurate, so that the initial threshold can be adjusted based on the speaking habit of the interactive object. For example, if the question is incomplete, the audio mute threshold and the mouth shape speaking threshold are increased. If the sentence contains a question, but there are other words, the threshold is reduced, and the user is asked to ask again and determine the effect until it is determined that the given answer is accurate.

[0129] When the large-screen device receives the reply of the interactive object, that is, the interactive object performs the correct corresponding action (first operation), for example, the interactive object replies with the words "correct" or "yes", or clicks the special function key, for example, the interactive object clicks the "correct" function key option set on the screen, or a specific gesture, such as the OK gesture, nodding, etc., to determine that the answer is accurate, then the audio mute threshold and the mouth shape speaking threshold are not adjusted, and S608 is performed.

[0130] S608, record the voiceprint information of the interaction object, and store the audio mute threshold and the mouth speech threshold. When the interaction object interacts with the digital person again, the large-screen device can directly obtain the corresponding audio mute threshold and mouth speech threshold according to the voiceprint information, and quickly obtain the question data, thereby improving the interaction efficiency and experience.

[0131] According to the interaction method of the embodiment of the present application, the actual interaction object can be accurately identified by waking up and capturing the mouth shape of the speaker through the camera, and the interaction object does not need to use a special microphone. When the interaction object asks a question, the audio data and the mouth shape data of the interaction object are obtained in real time, and the obtained audio data and mouth shape data are compared with the threshold of the object respectively, and the result of the comparison of the two is combined to judge whether the question of the interaction object is complete, so that the completeness of the question of the interaction object can be accurately judged. And for possible adverse factors, such as low microphone volume, high ambient noise, speech pause and other interference factors, they can also be judged and degraded. In addition, the response to the reply is confirmed with the interaction object, and the mouth speech threshold and the audio mute threshold are further dynamically adjusted, and through multiple rounds of correction, the mouth speech threshold and the audio mute threshold matched with the interaction object are obtained, and personalized communication experience is realized. And the voiceprint of the interaction object and the matched mouth speech threshold and audio mute threshold are stored and saved, so that in subsequent communication between the interaction object and the digital person, the large-screen device can directly obtain the threshold according to the voiceprint information, which can improve the interaction efficiency and experience.

[0132] The voice interaction method of the embodiment of the present application will be described below in combination with the software architecture of the large-screen device and the cloud device.

[0133] Referring to FIG. 7, FIG. 7 shows a software module structure diagram of the large-screen device and the cloud device according to the embodiment of the present application. As shown in FIG. 7, the large-screen device includes a voice wake-up module, a mouth shape detection module, a voiceprint module, a mute detection and mouth shape detection threshold adjustment module, an ASR, a TTS, a digital person management module, a digital person voice driving module, and an audio and video sending module. The cloud device includes a language large model / knowledge base service.

[0134] The modules can be arranged in the software system of the electronic device 100 as shown in FIG. 5. The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture. Embodiments of the present application take an Android system with a layered architecture as an example. The layered architecture divides software into several layers, each layer has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, the application layer, the application framework layer, the Android runtime and the system library, and the kernel layer. Among them, the modules can be arranged in the application layer, the application framework layer, and / or the system library, etc.

[0135] The voice wake-up module can start the device or application by recognizing a specific wake-up word or command. For example, it can be used to detect whether the user has input the voice corresponding to the wake-up word according to the preset wake-up word. If the matching is successful, it indicates that the wake-up is successful, and the digital person can be started. The voice wake-up module can perform the determination of whether the wake-up word is accurate in S602 shown in FIG. 6, and start the digital person if it is accurate.

[0136] The mouth shape detection module is used to obtain the mouth shape data of the interrogator captured by the camera, and detect whether the mouth shape data matches the mouth shape corresponding to the wake-up word. From the mouth shapes of multiple people, the best matching value is identified to determine the interrogator (interaction object). The mouth shape detection module can perform the step of identifying the speaker in S602 shown in FIG. 6.

[0137] The mute detection and mouth shape detection threshold adjustment module is used to dynamically adjust the threshold based on an initial threshold value through actual input to adjust the optimal matching value. This module is used to perform S607 in FIG. 6.

[0138] The ASR module is used to convert the input audio data into text processing. The ASR module can perform the text processing of the audio data in the first and third schemes in S604.

[0139] The digital person management module is responsible for the overall process control of digital person interaction. For example, it sends the text of the question of the interaction object to the language large model or knowledge base, or sends the answer returned by the language large model or knowledge base to the digital person generation module.

[0140] The digital person generation module is used to generate driving parameters corresponding to the answer, drive the digital person to speak through the driving parameters, including mouth shape, expression, and body movement, etc., and send the answer text to TTS.

[0141] A TTS module for converting text to speech, which can be used when the digital person answers questions in the embodiments of the present application. After S605 as shown in FIG. 6, i.e., after determining the answer corresponding to the question, the TTS module converts the text corresponding to the answer into audio to facilitate the large-screen device to output the audio, achieving the realism of the dialogue with the digital person.

[0142] An audio and video sending module, which sends the picture frame generated by the digital person reasoning to the display screen of the large-screen device for display and sends the audio data to the loudspeaker of the large-screen device for audio output.

[0143] A language large model / knowledge base service, which analyzes and understands the question and decides the content of the answer of the digital person. Usually, the combination of the knowledge base and the language large model is used to give the content of the answer, which can improve the accuracy of the content of the answer. The language large model / knowledge base service corresponds to S605 in FIG. 6.

[0144] A voiceprint module, which is used to extract the features of the face and voice of the interactive object after determining the interactive object or determining that the answer to the interactive object meets the intent, and stores the extracted data and the corresponding audio mute threshold and mouth speech threshold in the system database, preparing for the subsequent question matching data. The voiceprint module corresponds to S608 in FIG. 6.

[0145] The voice interaction method of the embodiments of the present application is further described in detail below in combination with the interaction process of each module in FIG. 6.

[0146] Referring to FIG. 8, FIG. 8 shows an interaction flowchart of software modules in the embodiments of the present application. As shown in FIG. 8, the large-screen device includes a voice wake-up module, a mouth shape detection module, a voiceprint module, a mute detection and mouth shape detection threshold adjustment module, an ASR module, a TTS module, a digital person management module, a digital person voice driving module, and an audio and video sending module. The cloud device includes a language large model / knowledge base service module.

[0147] As shown in FIG. 8, the interaction flow between the software modules includes two interaction scenarios.

[0148] Scenario one, first question and answer. In this scenario one, as shown in FIG. 8, the interaction process includes S801-S816.

[0149] S801, the voice wake-up module detects the voice in real time.

[0150] The voice wake-up module has the function of detecting the voice in real time, which can continuously listen to the surrounding sound and respond when hearing the specific wake-up word or command. Among them, the setting of the wake-up word can be multi-language, such as Chinese, English, Japanese, etc., so that users in different language environments can use it.

[0151] In some embodiments, the wake-up word can be customized according to the personal preferences of the user.

[0152] S802, the voice wake-up module informs the lip detection module that the wake-up word has been received.

[0153] For example, when the voice wake-up module obtains the wake-up word, the lip detection module is triggered to start real-time monitoring of the user's lip shape. In some embodiments, the lip detection module can also start without the notification of the voice wake-up module, and always remain in a detection state when the large-screen device is powered on.

[0154] S803, the lip detection module detects the questioner's lip shape data in real time.

[0155] The lip detection module can capture the facial data of the crowd in front of the screen in real time through the camera, locate the face, and then locate the lip area to observe the lip movements in more detail. By continuously tracking the changes in lip movements, the subtle changes in the questioner's speech can be captured more accurately.

[0156] S804, the voice wake-up module determines that the wake-up is successful and informs the lip detection module.

[0157] The voice wake-up module also has the function of wake-up word detection, which can detect keywords, for example, the pre-set wake-up word for the digital person is "Hello, Lili", and the voice wake-up module captures the wake-up word, then determines that the intention of the questioner is to wake up the digital person, and then triggers the next action.

[0158] S805, the lip detection module detects the questioner's lip shape and captures and determines the lip shape.

[0159] The lip detection module identifies specific lip shapes or pronunciations by analyzing the lip movement sequence. The identified lip shape is compared with the wake-up word to determine the interaction object. And according to the lip shape of the interaction object, the audio mute threshold and the lip shape speaking threshold that match it can be determined.

[0160] S806, the mute detection and lip detection dynamic adjustment module adjusts the preset value according to the detected lip shape and audio.

[0161] The preset value refers to a pre-set mute detection threshold and a mouth shape speaking threshold. The process can include two cases. One is that when the first dialogue occurs, the mute detection and mouth shape detection dynamic adjustment module selects an initial value matching the interactive object from the pre-stored audio mute threshold and mouth shape speaking threshold. However, the initial value is not verified and may be inaccurate. Therefore, the second case occurs, that is, after the interactive object answers the question for the first time, the interactive object can be asked whether the answer is accurate. If the answer is accurate, the initial value of the audio mute threshold and the mouth shape speaking threshold does not need to be adjusted. If the interactive object answers inaccurately, the audio mute threshold and the mouth shape speaking threshold are adjusted based on the audio signal value and the mouth shape action of the interactive object until the interactive object determines that the answer is accurate. Thus, the threshold is dynamically adjusted, the completeness of the question input by the interactive object can be more accurately determined, and the interactive experience is improved.

[0162] In addition, in some embodiments, the large-screen device can not need to ask whether the answer to the question is accurate after giving the answer or response, but directly ask the interactive object whether the question itself is accurate, for example, the voice asks, “Is the question you want to ask: the route map of the A mall?” When the user determines that it is, the answer is directly output, and when the answer is not, the audio mute threshold and the mouth shape speaking threshold are adjusted, and the interactive object is reminded to “repeat your question”. Until a positive answer is obtained, the adjustment of the threshold is ended, and is saved. The method can improve the speed of adjusting the threshold.

[0163] The process of adjusting the preset value according to the detected mouth shape and audio will be described below with reference to the accompanying drawings.

[0164] Referring to FIG. 9, FIG. 9 shows a flowchart of adjusting the preset value according to the detected mouth shape and audio in the embodiment of the application. The steps in the flowchart are executed by the mute detection and mouth shape detection threshold adjustment module, and the flowchart corresponds to S604 in FIG. 6.

[0165] As shown in FIG. 9, the flowchart can include S910 and S920. S902 further includes S921-S927.

[0166] S910, the mute detection and mouth shape detection threshold adjustment module obtains audio data and mouth shape data input by the interactive object, wherein the audio data refers to the audio of the question asked by the interactive object. The mouth shape data is the mouth shape data extracted based on the facial expression data, including dynamic data of the mouth shape, and a mouth opening degree value.

[0167] S920, the current input audio and mouth shape are judged.

[0168] When the user is confirmed to be asking a question for the first time, the mute detection and mouth shape detection threshold adjustment module directly uses the initial values of the audio mute threshold and the mouth shape speaking threshold stored in the system. Then the obtained audio data and mouth shape data are compared with the initial values.

[0169] When the audio is greater than the audio mute threshold and the mouth shape action is less than or equal to the mouth shape speaking threshold, it is determined that the question of the interactive object has been completed. S921 determines that there is interference noise, and S922 is performed to perform noise reduction processing on the audio data.

[0170] When the audio is less than or equal to the audio mute threshold and the mouth shape action is greater than the mouth shape speaking threshold, it indicates that the interactive object is still speaking, and it is determined that the question of the interactive object is not completed and the question is not complete. Therefore, S924 is performed to prompt the interactive object to adjust the loudness of the sound pickup device, and S924 is performed to prompt the user to approach the digital person or to increase the volume. S925 continues to pick up sound, reacquires new audio data and mouth shape data for judgment, and continues until it is determined that the question is complete.

[0171] When the audio is less than or equal to the audio mute threshold and the mouth shape action is less than or equal to the mouth shape speaking threshold, it indicates that the interactive object has stopped speaking, and it is determined that the question of the interactive object has been completed. Then S926 is performed to send the audio data to the ASR for recognition, and to convert the audio into text.

[0172] When the audio is greater than the audio mute threshold and the mouth shape action is greater than the mouth shape speaking threshold, it indicates that the interactive object is still speaking, and S927 is performed to continue to pick up sound, reacquire new audio data and mouth shape data for judgment, and continue until it is determined that the question of the interactive object has been completed.

[0173] The above-described four results are all for judging whether the current user's question is completed. When it is determined that the question is completed, it means that the sentence is preliminarily determined to be complete. However, in order to further verify the accuracy of the judgment, S809-S814 in FIG. 8 can be performed, which is a process of further adjusting the audio mute threshold and the mouth shape speaking threshold. It is also a process of determining whether the question of the interactive object is completed again. Through multiple verifications, the best threshold is finally dynamically obtained, thereby improving the accuracy of judging whether the question is completed.

[0174] S807, the mute detection and mouth shape detection dynamic adjustment module sends the voice data of the question to the ASR and notifies the voice-to-text module.

[0175] S808, the ASR module converts the voice into text and notifies the digital person management module.

[0176] In some embodiments, after the ASR module obtains the audio, if it is determined that there is noise, the audio will also be preprocessed, such as noise reduction, gain adjustment, etc., to improve the quality of subsequent audio processing.

[0177] S809, the digital human management module sends the text corresponding to the question to the language large model / knowledge base, and notifies the language large model / knowledge base that it is the first interaction, and enters the question confirmation link.

[0178] S810, the language large model / knowledge base starts the question correction question and answer mode.

[0179] For example, the large screen device can have two modes, such as the correction question and answer mode and the regular mode. In the correction question and answer mode, the question of the interactive object needs to be confirmed. In the regular mode, the interactive object does not need to confirm the question, and directly outputs the answer. Generally, the regular mode corresponds to the scenario of the interactive object asking the question again, that is, the subsequent question and answer after scenario two. The correction question and answer mode can correspond to scenario one, that is, the scenario of the first question and answer.

[0180] S811, the language large model / knowledge base obtains the answer corresponding to the question, and notifies the digital human generation module to drive the digital human to confirm the accuracy of the question with the questioner.

[0181] In this process, the digital human generation module can generate driving parameters of the driven digital human based on the answer, such as audio, “Do you think the answer of XX is satisfactory?”, and action parameters, such as data of mouth opening, posture and body action of the digital human, etc.

[0182] S812, the digital human generation module notifies the TTS module to generate voice from the answer text, and controls the audio and video sending module to output the audio corresponding to the answer, and displays the animation of the digital human.

[0183] After the TTS module generates voice (audio) from the answer text, the audio and video sending module sends the corresponding audio to the loudspeaker, and outputs the audio. After the audio and video sending module receives the driving digital human action instruction, it drives the digital human action according to the driving parameters, so that the digital human action is synchronized with the audio, to achieve the effect of simulating human conversation.

[0184] In addition, the digital human management module also issues a confirmation instruction, which can be output in the form of audio through the loudspeaker, for example, after the digital human outputs the answer audio, it asks the questioner whether the answer is satisfactory. Or output in the form of text displayed on the display screen, ask the questioner whether the current answer is satisfactory. And after obtaining the result input by the questioner, feedback the result.

[0185] S813, the language large model / knowledge base module obtains the confirmation result of the questioner, and adjusts the corresponding threshold according to the confirmation result.

[0186] When the language large model / knowledge base module confirms that the result is inaccurate, the adjustment threshold is confirmed. For example, the question is incomplete, indicating that the threshold setting threshold is too high, then the audio mute threshold and mouth speech threshold are increased. If there is other text in the question, indicating that the threshold setting threshold is too low, then the threshold is reduced. After adjustment, the user is asked again and the effect is determined, through multiple rounds of correction, the optimal audio mute threshold and mouth speech threshold for the questioner are obtained, that is, S814 is executed, and the correction is completed.

[0187] When the confirmation result is accurate, or after multiple rounds of correction, the language large model / knowledge base can execute S815.

[0188] S815, the language large model / knowledge base module informs the digital human management module that the correction is completed.

[0189] S816, the digital human management module informs the voiceprint module to obtain voiceprint information and keep the threshold.

[0190] The voiceprint module obtains voiceprint information from the obtained audio data, and stores the voiceprint information and the audio mute threshold and the mouth speech threshold, so that when the questioner asks again in the future, the audio mute threshold and the mouth speech threshold corresponding to the questioner can be quickly obtained based on the voiceprint information, and the interaction efficiency can be improved.

[0191] Scenario two, follow-up question and answer (ask again). In this scenario two, the interaction process includes S817-S819.

[0192] S817, the voice wake-up module detects voice information and informs the mute detection and mouth detection dynamic adjustment module that it is a non-wake-up situation.

[0193] For example, the voice wake-up module obtains audio data of the question input by the user, and identifies the identity of the user based on voiceprint information in the audio data.

[0194] Among them, the non-wake-up situation indicates that the user is not asking for the first time.

[0195] S818, the mute detection and mouth detection dynamic adjustment module informs the voiceprint module and triggers the voiceprint module to identify voiceprint information and obtain previously saved audio mute threshold and mouth speech threshold. The mute detection and mouth detection dynamic adjustment module updates the system set audio mute threshold and mouth speech threshold to the threshold corresponding to the voiceprint information.

[0196] Subsequently, when obtaining a question from the same questioner, the large-screen device can directly convert the voice data to text after obtaining the voice data and the mouth shape data, and then give an answer to the question based on the text and the mouth shape data by the language large model / knowledge base module, and output the corresponding audio and animation of the answer. In this process, there is no need to perform threshold correction confirmation again, that is, S819, to determine to follow the normal process, and the mute detection and mouth shape detection dynamic adjustment module does not need to perform correction confirmation again. The normal process can be a regular mode, that is, a process without correction confirmation.

[0197] In some embodiments, in the scenario of subsequent interaction, the processing of the accuracy of the matching of the question and the answer can also be increased. For example, if the internal matching answer is not ideal and the matching degree is lower than 70%, it is determined that the recognized question cannot find an answer, and the above steps S809-S814 of threshold adjustment process can be performed again, and the interactive object is asked to input the question again. The digital human management module can also perform S820 to start the question and answer stability detection, and if the question cannot be matched for multiple times, the threshold adjustment is restarted to improve the overall experience of digital human interaction. The adjustment process can refer to the above steps S809-S814.

[0198] In addition, considering the changes of environmental or personal factors, for example, the interval time between the first question and the second question is relatively long, such as half a year, 1 year, etc., the user's volume or speaking speed may change. Therefore, the storage time of the audio mute threshold and the mouth shape speaking threshold of the user who has asked the question for the first time can be set to a period of time, such as 1 year, and when the period of 1 year is exceeded, the automatic threshold adjustment needs to be restarted. In addition, whether to adjust the threshold can also be combined with the user's own physical changes, for example, the user has just come back from running, and at this time, the speaking speed and volume of the question may be different from the speaking speed and volume in the normal state. Therefore, when the question is asked again, the threshold determined by the first question will also affect the accuracy of the judgment. At this time, factors such as heart rate can be combined into the judgment, and when the heart rate value is within the normal range, the threshold of the first question is used. When the heart rate value exceeds the normal range in one of the questions, the threshold adjustment is restarted. By combining these factors, the accuracy of the answer can be further improved, thereby improving the interactive experience.

[0199] It should be noted that in the above embodiments, the facial expression data is taken as an example of the mouth shape data for description. In some embodiments, the facial expression data can also include one or more of eyebrow movement data, eye movement data, and nose movement data. These facial expression data can be identified by the end-side small model. By identifying the expression of the interactive object, the emotional parameter is increased and combined into the above judgment of whether the user's question is completed. The user's intention can be more accurately judged, and the user's question can be more accurately judged as incomplete. The digital person can make more accurate question answers based on understanding the text and combining emotional data, gradually approaching the experience of communicating with a real person.

[0200] In addition, the facial expression of the digital person can also be correspondingly increased in the driving parameter of the digital person, so that the whole interaction process is more vivid and realistic.

[0201] In summary, based on the above scheme, the voice interaction method of the embodiments of the present application accurately determines the interactive object by combining mouth shape detection and wake-up word recognition. Then, by combining voice mute detection and mouth shape detection, and with the help of continuous interaction between the digital person and the user, the most suitable mute and mouth shape detection parameters for the user are gradually learned and adjusted to meet the individual needs. In addition, voiceprint technology is also used to store personal interaction data, so that these data can be directly called in future interaction, thereby improving the overall experience of AI digital person interaction.

[0202] As shown in FIG. 10, the present application also provides a voice interaction device 1000, comprising:

[0203] The object confirmation module 1010 is configured to determine the interactive object.

[0204] The acquisition module 1020 is configured to acquire facial expression data and audio data of the interactive object, wherein the facial expression data at least includes mouth shape data.

[0205] The completeness confirmation module 1030 is configured to determine the completeness of the question raised by the interactive object based on the audio data and the mouth shape data.

[0206] The generation module 1040 is configured to determine the response corresponding to the question based on the audio data and the mouth shape data in the case that the question has been completed, and generate driving information based on the response.

[0207] The output module 1050 is configured to output the media data corresponding to the response based on the driving information.

[0208] The object confirmation module 1010, the obtaining module 1020, the integrity confirmation module 1030, the generating module 1040, and the output module 1050 can be implemented by software or by hardware. For example, the implementation of the object confirmation module 1010 is described below. The implementation of the obtaining module 1020, the integrity confirmation module 1030, the generating module 1040, and the output module 1050 can be similar to the implementation of the object confirmation module 1010.

[0209] As an example of a software functional unit, the object confirmation module 1010 can include code running on a computing instance. The computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the computing instance can be one or more. For example, the object confirmation module 1010 can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers running the code can be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers running the code can be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple data centers in close geographical proximity. Generally, one region can include multiple AZs.

[0210] Similarly, the multiple hosts / virtual machines / containers running the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Generally, one VPC is set in one region, and a communication gateway needs to be set in each VPC for cross-region communication between two VPCs in the same region or between VPCs in different regions, and the interconnection between VPCs is realized through the communication gateway.

[0211] As an example of a hardware functional unit, the object confirmation module 1010 can include at least one computing device, such as a server or the like. Alternatively, the object confirmation module 1010 can also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), and the like. The PLD can be implemented by a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0212] The plurality of computing devices included in the object confirmation module 1010 can be distributed in the same region or in different regions. The plurality of computing devices included in the object confirmation module 1010 can be distributed in the same AZ or in different AZs. Similarly, the plurality of computing devices included in the object confirmation module 1010 can be distributed in the same VPC or in multiple VPCs. The plurality of computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0213] It should be noted that in other embodiments, the object confirmation module 1010 can be configured to perform any of the steps of the voice interaction method, the acquisition module 1020 can be configured to perform any of the steps of the voice interaction method, the integrity confirmation module 1030 can be configured to perform any of the steps of the voice interaction method, the generation module 1040 can be configured to perform any of the steps of the voice interaction method, and the output module 1050 can be configured to perform any of the steps of the voice interaction method. The steps implemented by the object confirmation module 1010, the acquisition module 1020, the integrity confirmation module 1030, the generation module 1040, and the output module 1050 can be specified as needed, and the voice interaction device 1000 can be implemented by the object confirmation module 1010, the acquisition module 1020, the integrity confirmation module 1030, the generation module 1040, and the output module 1050 implementing different steps of the voice interaction method.

[0214] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0215] As shown in Figure 11, the computing device cluster includes at least one computing device 1100. The memory 1106 of one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the voice interaction methods explained in Figures 4-9 of the above embodiments.

[0216] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the voice interaction methods explained in Figures 4-9 of the above embodiments. In other words, a combination of one or more computing devices 1100 can jointly execute instructions for executing the voice interaction methods explained in Figures 4-9 of the above embodiments.

[0217] It should be noted that the memory 1106 in different computing devices 1100 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the voice interaction device. That is, the instructions stored in the memory 1106 of different computing devices 1100 can implement the functions of one or more modules among the object verification module, acquisition module, integrity verification module, generation module, and output module.

[0218] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 12 illustrates one possible implementation. As shown in Figure 12, two computing devices 1100A and 1100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 1106 in computing device 1100A stores instructions for performing the functions of the integrity verification module. Simultaneously, the memory 1106 in computing device 1100B stores instructions for performing the functions of the object verification module, the acquisition module, the generation module, and the output module.

[0219] The connection method between the computing device clusters shown in Figure 12 can be such that, considering the voice interaction method provided in this application needs to obtain a large amount of audio data and facial expression data of the interactive object (e.g., a large amount of stored data) and to judge and recognize the completeness of the user's sentences through a large language model, the functions implemented by the integrity confirmation module are handed over to the computing device 1100A to perform, taking into account processing speed and storage capacity.

[0220] It should be understood that the functions of the computing device 1100B shown in FIG. 12 can also be completed by a plurality of computing devices 1100. Similarly, the functions of the computing device 1100A can also be completed by a plurality of computing devices 1100.

[0221] The embodiments of the present application also provide another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to the connection mode of the computing device cluster described with reference to FIG. 11 and FIG. 12. The difference is that the memory 1106 in one or more computing devices 1100 in the computing device cluster can store the same instructions for performing the voice interaction method explained in FIG. 4-FIG. 9 in the above embodiments.

[0222] In some possible implementation manners, the memory 1106 of one or more computing devices 1100 in the computing device cluster can also respectively store part of the instructions for performing the voice interaction method explained in FIG. 4-FIG. 9 in the above embodiments. In other words, the combination of one or more computing devices 1100 can collectively execute the instructions for performing the voice interaction method explained in FIG. 4-FIG. 9 in the above embodiments. The present application also provides an electronic device, comprising:

[0223] a memory for storing instructions executed by one or more processors of the device, and

[0224] a processor for executing the method explained in conjunction with FIG. 4 to FIG. 9 in the above embodiments.

[0225] The present application also provides a computer readable storage medium, which can be any available medium or data storage device such as a data center containing one or more available media that can be stored by a computing device. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk) and the like. The computer readable storage medium includes instructions, which instruct the computing device to execute the method explained in FIG. 4 to FIG. 9 in the above embodiments, or instruct the computing device to execute the method explained in FIG. 4 to FIG. 9 in the above embodiments.

[0226] The present application also provides a computer program product containing instructions, which can be software or program products containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it makes the at least one computing device execute the method explained in FIG. 4 to FIG. 9 in the above embodiments.

[0227] Referring now to FIG. 13, shown is a block diagram of a SoC (System on a Chip) 1400 in accordance with an embodiment of the present application. In FIG. 13, like elements have the same reference numbers as in FIG. 12. Additionally, dashed lined boxes are optional features on more advanced SoCs. In FIG. 13, the SoC 1400 includes an interconnect unit 1450 coupled to an application processor 1410; a system agent unit 1480; a bus controller unit 1490; an integrated memory controller unit 1440; a set or one or more coprocessors 1420A-N which can include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random access memory (SRAM) unit 1430; a direct memory access (DMA) unit 1460. In one embodiment, the coprocessors 1420A-N can also include a dedicated processor, such as for example a network or communication processor, compression engine, cryptography processor, GPGPU, a high- throughput MIC processor, embedded processor, etc.

[0228] The SRAM unit 1430 can include one or more computer-readable media for storing data and / or instructions. The computer-readable storage media can store instructions, specifically, temporary and permanent copies of the instructions. The instructions can include those that, when executed by at least one of the processors, cause the SoC 1400 to perform a method of processing according to the embodiments described above, specifically, the voice interaction method explained with reference to the embodiments of FIG. 4 and FIG. 9, which will not be repeated here.

[0229] Embodiments of the mechanisms disclosed herein can be implemented in hardware, software, firmware, or any combination thereof. Embodiments of the application can be implemented as computer programs or program code executing on programmable systems comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0230] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices, in known fashion. For purposes of this application, a processing system includes any system that has a processor, such as a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0231] The program code can be implemented in a high level of programming language or a object-oriented programming language to communicate with a processing system. When necessary, the program code can also be implemented in assembly or machine language. In fact, the mechanisms described in this application are not limited to any particular programming language. In any case, the language can be a compiled or interpreted language.

[0232] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried by or stored on a transitory or non-transitory machine-readable (e.g., computer-readable) medium, which can be read and executed by one or more processors. For example, the instructions can be distributed over the network or by other computer readable media. Thus, a machine-readable medium can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including without limitation floppy disks, optical disks, optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or any other suitable device. Accordingly, a machine-readable medium includes any medium that is suitable for storing or transmitting the instructions for use by or in connection with an instruction execution system, apparatus, or device.

[0233] In the drawings, some of the structural or methodological features can be shown in particular arrangements and / or orders. However, it should be understood that such particular arrangements and / or orders can not be required. Instead, these features can be arranged in a different manner and / or order in some embodiments from that shown in the illustrative drawings. Additionally, inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and these features can be excluded or combined with other features in some embodiments.

[0234] While the present application has been illustrated and described with reference to certain preferred embodiments thereof, it should be understood that various changes in form and detail can be made therein without departing from the spirit and scope of the application.

Claims

1. A voice interaction method, characterized in that, The method comprises: determining an interactive object; acquiring facial expression data and audio data of the interactive object, the facial expression data at least including mouth shape data; determining completeness of a question raised by the interactive object based on the audio data and the mouth shape data; if it is determined that the question is complete based on the audio data and the mouth shape data, determining a response corresponding to the question and generating driving information based on the response; outputting media data corresponding to the response based on the driving information.

2. The method of claim 1, wherein, The determination of the completeness of the question based on the audio data and the mouth shape data comprises: determining that the mouth shape data of the interactive object is less than or equal to a mouth shape speaking threshold value and a signal value of the audio data is greater than an audio mute threshold value, and then determining that the question is complete.

3. The method of claim 2, wherein, The determination of the completeness of the question based on the audio data and the mouth shape data further comprises noise reduction processing of the audio data.

4. The method of claim 2, wherein, The determination of the completeness of the question based on the audio data and the mouth shape data comprises: determining that the mouth shape data of the interactive object is less than or equal to the mouth shape speaking threshold value and the signal value of the audio data decreases to be less than or equal to the audio mute threshold value, and then determining that the question is complete.

5. The method of claim 1, wherein, The determination of the completeness of the question raised by the interactive object based on the audio data and the mouth shape data comprises: if it is determined that the question is not complete and no response corresponding to the incomplete question is outputted within a preset time length, maintaining a state of picking up data.

6. The method of claim 5, wherein, The maintaining of the state of picking up data comprises: outputting prompt information prompting that the interactive object is picking up data or maintaining a state of a digital human listening, and continuing to pick up audio data and facial expression data of the interactive object.

7. The method according to claim 5 or 6, characterized in that, The determination of the incompleteness of the question comprises: determining that the mouth shape data of the interactive object is greater than the mouth shape speaking threshold value and the signal value of the audio data decreases to be less than or equal to the audio mute threshold value, or determining that the mouth shape data of the interactive object is greater than the mouth shape speaking threshold value and the signal value of the audio data is greater than the audio mute threshold value.

8. The method of claim 7, wherein, The determination of the incompleteness of the question further comprises: prompting the interactive object to adjust the volume or approach the sound pickup device.

9. The method of claim 1, wherein, The determination of the response corresponding to the question comprises: acquiring audio data of the complete question; obtaining the response corresponding to the question based on the audio data of the question and a language large model.

10. The method of claim 9, wherein, The obtaining of the response corresponding to the question based on the audio data of the question and the language large model comprises: converting the audio data of the question into text; inputting the text into the language large model to obtain response text corresponding to the text based on the language large model; converting the response text into the audio response.

11. The method of claim 9, wherein, The obtaining of the response corresponding to the question based on the audio data of the question and the language large model further comprises: outputting a prompt of whether the response conforms to an intention of the interactive object; if a first operation conforming to the intention of the interactive object is received, outputting media data corresponding to the response. If a second operation that does not conform to the intention of the interactive object is received, the mouth shape speaking threshold and the audio mute threshold are adjusted, and the interactive object is prompted to ask again until the response conforms to the intention of the interactive object.

12. The method of claim 11, wherein, The first operation includes a pressing operation on a specific function key, a movement in a specific posture, or input of a keyword representing conformity to the intention. The second operation includes a pressing operation on a specific function key, a movement in a specific posture, or input of a keyword representing non-conformity to the intention.

13. The method of claim 1, wherein, The determination of the interactive object includes: Receiving a wake-up word input by a user, and obtaining mouth shape data of the user; Based on the wake-up word and the mouth shape data, the interactive object whose wake-up word and mouth shape data match is determined from the user.

14. The method of claim 12, wherein, Also includes: Extracting a voiceprint feature in the audio data of the interactive object, and storing; When it is determined that the interactive object is not asking for the first time, the mouth shape speaking threshold and the audio mute threshold corresponding to the interactive object are directly obtained based on the voiceprint feature.

15. The method of claim 1, wherein, The facial expression data further includes one or more of eyebrow movement data, eye movement data, and nose movement data.

16. The method of claim 1, wherein, The media data includes animation data of a digital human or a specific shaped object in a display screen, and output audio data.

17. A voice interactive device, characterized by It includes: An object confirmation module for determining an interactive object; An acquisition module for acquiring facial expression data and audio data of the interactive object, the facial expression data including at least mouth shape data; An integrity confirmation module for determining the integrity of a question asked by the interactive object based on the audio data and the mouth shape data; A generation module for determining that the question is complete based on the audio data and the mouth shape data, determining a response corresponding to the question, and generating driving information based on the response; An output module for outputting media data corresponding to the response based on the driving information.

18. An electronic device, comprising: It includes: A memory for storing instructions executed by one or more processors of an electronic device, and a processor for executing the voice interaction method of any one of claims 1 to 16.

19. A cluster of computing devices, characterized in that, It includes at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the computing device cluster to perform the voice interaction method of any one of claims 1 to 16.

20. A computer-readable storage medium, characterized in that, The computer readable storage medium has instructions stored thereon, which when executed on an electronic device, cause the electronic device to perform the voice interaction method of any one of claims 1 to 16.

21. A computer program product comprising instructions, wherein: When the computer program product is running on an electronic device, it causes the electronic device to perform the voice interaction method of any one of claims 1 to 16.

Citation Information

Patent Citations

  • A search method based on a tutor machine and a tutor machine

    CN109145088A

  • Voice recognition method and device, electronic equipment and storage medium

    CN110534109A

  • Virtual character driving method, system and equipment based on multi-modal data

    CN114840090A

  • Voice interaction method, electronic equipment and medium

    CN115691498A

  • Method and system for authenticating user of a mobile device via hybrid biometics information

    US20130227678A1