Digital human voice interaction processing method, device, electronic device and medium

By performing noise reduction processing based on voiceprint data during the digital human voice interaction process, the problem of noise interference in voice interaction is solved, and the recognition accuracy of user interruption commands is improved, especially when the similarity between the user and the digital human voiceprint is high.

CN119495298BActive Publication Date: 2025-09-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411613352.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-09-23
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

When a user interacts with a digital human through voice, there is noise in the user's voice audio collected during the voice interaction process, which affects the recognition accuracy of the user's interruption command.

Method used

During the voice interaction process, the voiceprint data currently used by the digital human for voice broadcasting is obtained, and noise reduction processing is performed to improve the clarity of the user's voice. Targeted noise reduction is performed on the digital human's own broadcast voice based on the voiceprint data to identify the user's interruption commands.

Benefits of technology

The recognition accuracy of user interruption commands has been improved, especially when the similarity between the user's voiceprint and the digital human's broadcast voice is high. The noise-reduced audio data can more accurately recognize the user's commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119495298B_ABST
    Figure CN119495298B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, device, electronic device, and medium for processing digital human voice interaction, relating to the fields of natural language processing technology, and in particular, speech recognition, semantic recognition, intelligent agents, and generative search technology. The implementation scheme is as follows: in response to receiving first audio data, obtaining first voiceprint data of the digital human, wherein the first audio data indicates that the user has initiated a voice interaction request to the digital human, and the first voiceprint data indicates the voiceprint data currently used by the digital human for voice broadcasting; in response to receiving second audio data, performing first noise reduction processing on the second audio data based on the first voiceprint data to obtain third audio data, wherein the time when the second audio data is received is after the time when the first audio data is received; and in response to determining, based on the third audio data, that the user has issued a first instruction to interrupt the voice interaction, generating a stop instruction and sending it to the digital human to control the digital human to stop voice broadcasting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of natural language processing technology, in particular to the field of speech recognition, semantic recognition, intelligent agents and generative search technology, and specifically to a method, device, electronic device, computer-readable storage medium and computer program product for processing digital human voice interaction. Background Art

[0002] When the user is interacting with the digital human through voice, he or she may issue voice commands or click on the display screen to interrupt the current voice interaction, causing the digital human to stop the voice broadcast and restart the next voice interaction.

[0003] The approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, the issues raised in this section should not be considered as having been recognized in any prior art. Summary of the Invention

[0004] The present disclosure provides a method, device, electronic device, computer-readable storage medium, and computer program product for processing digital human voice interaction.

[0005] According to one aspect of the present disclosure, a method for processing digital human voice interaction is provided, comprising: in response to receiving first audio data, obtaining first voiceprint data of the digital human, wherein the first audio data indicates that a user initiates a voice interaction request to the digital human, and the first voiceprint data indicates the voiceprint data currently used by the digital human for voice broadcasting; in response to receiving second audio data, performing first noise reduction processing on the second audio data based on the first voiceprint data to obtain third audio data, wherein the time when the second audio data is received is after the time when the first audio data is received; and in response to determining, based on the third audio data, that the user issues a first instruction to interrupt the voice interaction, generating a stop instruction and sending it to the digital human to control the digital human to stop voice broadcasting.

[0006] According to another aspect of the present disclosure, a processing device for digital human voice interaction is provided, comprising: an acquisition module, configured to, in response to receiving first audio data, acquire first voiceprint data of the digital human, wherein the first audio data indicates that a user initiates a voice interaction request to the digital human, and the first voiceprint data indicates the voiceprint data currently used by the digital human for voice broadcasting; a first noise reduction module, configured to, in response to receiving second audio data, perform first noise reduction processing on the second audio data based on the first voiceprint data to obtain third audio data, wherein the second audio data is received after the first audio data; and a processing module, configured to, in response to determining, based on the third audio data, that the user issues a first instruction to interrupt the voice interaction, generate a stop instruction and send it to the digital human to control the digital human to stop voice broadcasting.

[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above method.

[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the above method.

[0009] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the above method when executed by a processor.

[0010] According to one or more embodiments of the present disclosure, a method for processing digital human voice interaction is provided. When a user issues a voice interaction request, the voiceprint data currently used by the digital human for voice broadcast is obtained, and during the voice interaction process, noise reduction processing is performed on the digital human's own broadcast voice based on the voiceprint data, thereby making the user's voice in the audio data used for recognition clearer and improving the accuracy of identifying user interruption commands during the interaction process.

[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings illustrate exemplary embodiments and constitute a part of the specification. Together with the description of the specification, they serve to explain exemplary implementation of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals designate similar, but not necessarily identical, elements.

[0013] Figure 1 is a schematic diagram illustrating an example system in which the various methods described herein may be implemented, according to an exemplary embodiment;

[0014] Figure 2 A flowchart of a method for processing digital human voice interaction according to an embodiment of the present disclosure is shown;

[0015] Figure 3 A partial flow chart of another method for processing digital human voice interaction according to an embodiment of the present disclosure is shown;

[0016] Figure 4 A partial flow chart of another method for processing digital human voice interaction according to an embodiment of the present disclosure is shown;

[0017] Figure 5 A partial flow chart of another method for processing digital human voice interaction according to an embodiment of the present disclosure is shown;

[0018] Figure 6 A partial flow chart of another method for processing digital human voice interaction according to an embodiment of the present disclosure is shown;

[0019] Figure 7 A partial flow chart of another method for processing digital human voice interaction according to an embodiment of the present disclosure is shown;

[0020] Figure 8 A structural block diagram of a device for processing digital human voice interaction according to an embodiment of the present disclosure is shown; and

[0021] Figure 9 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0022] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0023] In this disclosure, unless otherwise specified, the use of terms such as "first" and "second" to describe various elements is not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, while in some cases, based on the context of the description, they may also refer to different instances.

[0024] The terms used in the descriptions of the various examples described in this disclosure are for the purpose of describing specific examples only and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in this disclosure encompasses any one and all possible combinations of the listed items.

[0025] In related technologies, when interacting with a digital human through voice, a user might issue a voice command to interrupt the current interaction, causing the digital human to stop broadcasting and begin recognizing and answering the user's next topic for voice interaction. However, the user's voice audio collected during the voice interaction contains a lot of noise, which significantly affects the accuracy of the user's interruption command.

[0026] To solve the above problems, the present disclosure provides a method for processing digital human voice interaction. When a user issues a voice interaction request, the voiceprint data currently used by the digital human for voice broadcast is obtained, and during the voice interaction process, noise reduction processing is performed on the digital human's own broadcast voice based on the voiceprint data, thereby making the user's voice in the audio data used for recognition clearer and improving the accuracy of identifying user interruption commands during the interaction process.

[0027] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0028] Figure 1 FIG2 is a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein may be implemented according to an embodiment of the present disclosure. Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.

[0029] In an embodiment of the present disclosure, the server 120 may run one or more services or software applications that enable execution of a processing method for digital human voice interaction.

[0030] In some embodiments, server 120 may also provide other services or software applications that may include non-virtualized environments and virtualized environments. In some embodiments, these services may be provided as web-based services or cloud services, such as provided to users of client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.

[0031] exist Figure 1 In the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may, in turn, utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may differ from the system 100. Therefore, Figure 1 is an example of a system for implementing the processing method of digital human voice interaction described herein and is not intended to be limiting.

[0032] The user can use client devices 101, 102, 103, 104, 105 and / or 106 to perform the digital human voice interaction processing method. The client device can provide an interface that enables the user of the client device to interact with the client device. The client device can also output information to the user via the interface. Figure 1 Only six client devices are depicted, but one skilled in the art will appreciate that the present disclosure can support any number of client devices.

[0033] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors or other sensing devices, etc. These computer devices may run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablet computers, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing a variety of different applications, such as various internet-related applications, communication applications (such as email applications), and short message service (SMS) applications, and may use various communication protocols.

[0034] The network 110 may be any type of network known to those skilled in the art that can support data communications using any of a variety of available protocols, including but not limited to TCP / IP, SNA, IPX, etc. By way of example only, the one or more networks 110 may be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0035] Server 120 may include one or more general-purpose computers, specialized server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running virtual operating systems, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that may be virtualized to maintain a server's virtual storage device). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0036] The computing units in the server 120 may run one or more operating systems including any of the operating systems described above as well as any commercially available server operating systems. The server 120 may also run any of a variety of additional server applications and / or middle-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, and the like.

[0037] In some implementations, server 120 may include one or more applications to analyze and consolidate data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display the data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.

[0038] In some embodiments, server 120 may be a distributed system server or a server integrated with blockchain. Server 120 may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host equipped with artificial intelligence technology. A cloud server is a host product within the cloud computing service system that addresses the management difficulties and poor scalability of traditional physical hosts and virtual private servers (VPS) services.

[0039] The system 100 may also include one or more databases 130. In some embodiments, these databases can be used to store audio data and other information. For example, one or more of the databases 130 can be used to store audio data such as that collected by a microphone. The databases 130 can reside in various locations. For example, the database used by the server 120 can be local to the server 120, or can be remote from the server 120 and can communicate with the server 120 via a network-based or dedicated connection. The databases 130 can be of different types. In some embodiments, the databases used by the server 120 can be, for example, relational databases. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.

[0040] In some embodiments, one or more of the databases 130 may also be used by applications to store application data. The databases used by the applications may be different types of databases, such as a key-value store, an object store, or a conventional store backed by a file system.

[0041] Figure 1 The system 100 may be configured and operated in various ways to enable application of the various methods and apparatuses described in accordance with the present disclosure.

[0042] Figure 2 A flowchart of a method for processing digital human voice interaction according to an embodiment of the present disclosure is shown.

[0043] like Figure 2 As shown, the method 200 for processing digital human voice interaction includes:

[0044] Step 210: In response to receiving the first audio data, obtaining the first voiceprint data of the digital human, wherein the first audio data indicates that the user has initiated a voice interaction request to the digital human, and the first voiceprint data indicates the voiceprint data currently used by the digital human for voice broadcasting;

[0045] Step 220: In response to receiving the second audio data, perform a first noise reduction process on the second audio data based on the first voiceprint data to obtain third audio data, wherein the second audio data is received after the first audio data is received; and

[0046] Step 230: In response to determining, based on the third audio data, that the user issues a first instruction to interrupt the voice interaction, a stop instruction is generated and sent to the digital human to control the digital human to stop the voice broadcast.

[0047] During the digital human's voice interaction, users often interrupt the digital human while it's speaking, allowing them to restart the interaction with the next topic. In this case, the digital human's speech becomes the primary noise in the audio data captured by the microphone. Furthermore, the digital human's speech is highly correlated with the current topic, interfering with the recognition of the user's interruption command.

[0048] Therefore, by performing targeted noise reduction processing based on the voiceprint data of the digital human's own voice during voice interaction, the user's voice can be made clearer in the audio data used for recognition, improving the accuracy of identifying user interrupt commands during the interaction process. In particular, when the user's voiceprint and the digital human's own voiceprint are highly similar, this method of user command recognition based on noise-reduced audio data is more effective.

[0049] In step 210, the first audio data may be, for example, audio data collected by the digital human's microphone, and the first voiceprint data may be obtained directly from the digital human. In one example, the digital human's default voiceprint data may be directly obtained. In another example, if the user has set a specific voice for the digital human's announcement, the set voiceprint data of the digital human may be obtained as the first voiceprint data.

[0050] Figure 3A partial flow chart of another method for processing digital human voice interaction according to an embodiment of the present disclosure is shown.

[0051] According to some embodiments, Figure 3 As shown, step 210 includes:

[0052] Step 311: In response to receiving the first audio data, perform speech recognition on the first audio data using a first speech recognition model to obtain a first recognized text; and

[0053] Step 312: In response to determining, based on the first recognition text, that the user initiates a voice interaction request to the digital human, obtaining first voiceprint data.

[0054] In step 311, the first speech recognition model can be constructed based on a neural network such as a DNN (Deep Neural Network), a CNN (Convolutional Neural Network), an LSTM (Long Short-Term Memory), a Conformer (a hybrid network structure), or a TDNN (Time-Delay Neural Network). It is understandable that the first speech recognition model can adopt other types of network structures, which are not limited here.

[0055] In step 312, the judgment condition for determining whether the user initiates a voice interaction request may be that the first recognition text includes specific keywords indicating the initiation of a voice interaction request, such as the name of the digital person, etc.; or, the judgment condition may also be that the number of words in the recognized first recognition text exceeds the target threshold.

[0056] Therefore, it is possible to convert the first audio data into text based on voice recognition to determine whether the user has initiated a voice interaction request. The implementation is simple and has high accuracy.

[0057] In step 220, the second audio data may be, for example, audio data collected by the microphone of the digital human. In this example, all audio data collected by the microphone after the user initiates a voice interaction request may be subjected to noise reduction processing to avoid missing the user's interruption instruction due to noise interference.

[0058] Figure 4 A partial flow chart of another method for processing digital human voice interaction according to an embodiment of the present disclosure is shown.

[0059] According to some embodiments, Figure 4 As shown, the second audio data at least includes the user's voice audio. In addition to steps 210 to 230, method 200 further includes:

[0060] Step 410: Determine the user's second voiceprint data based on the first audio data; and

[0061] Step 420: Determine whether the second audio data is received according to the second voiceprint information.

[0062] Therefore, only when the user's voice data is recognized will it be used as the second audio data for noise reduction processing, which can effectively save resource overhead.

[0063] In the example, if this is not the first time the user interacts with the digital human and the digital human or server has stored the user's voiceprint data, the stored user voiceprint data can be directly called as the second audio data to further reduce the processing difficulty.

[0064] Figure 5 A partial flow chart of another method for processing digital human voice interaction according to an embodiment of the present disclosure is shown.

[0065] According to some embodiments, Figure 5 As shown, step 230 includes:

[0066] Step 531: Use the second speech recognition model to perform speech recognition on the third audio data to obtain a second recognized text;

[0067] Step 532: In response to determining that the user has issued a first instruction based on the second recognized text, a stop instruction is generated and sent to the digital human to control the digital human to stop voice broadcasting.

[0068] In step 531, the second speech recognition model can be constructed based on a neural network such as DNN, CNN, LSTM, Conformer, TDNN, etc. It is understandable that the second speech recognition model can adopt other types of network structures, which are not limited here.

[0069] In step 532, the judgment condition for determining that the user has issued the first instruction can be that the second recognized text contains specific keywords indicating a request to interrupt voice interaction, such as "Please wait," "Next question," or "Excuse me." Alternatively, the judgment condition can be that the number of words in the recognized second recognized text exceeds a target threshold. For example, if the recognized user voice only contains interjections such as "hmm," "ah," or "oh," it can be ignored as invalid audio data.

[0070] Therefore, it is possible to convert the second audio data into text based on voice recognition to determine whether the user has initiated a request to interrupt the voice interaction. The implementation is simple and has high accuracy.

[0071] Figure 6A partial flow chart of another method for processing digital human voice interaction according to an embodiment of the present disclosure is shown.

[0072] According to some embodiments, before the user issues an instruction to interrupt the voice interaction, the number of interactions between the two may be greater than one. Based on this, Figure 5 As shown, in addition to steps 210 to 230, method 200 further includes:

[0073] Step 610: In response to determining, based on the second recognized text, that the user has issued a second instruction to continue voice interaction, processing the second recognized text using a deep learning network model to obtain a target output text, wherein the target output text indicates the digital human's reply content to the second instruction; and

[0074] Step 620: Convert the target output text into fourth audio data for the digital human to perform voice broadcasting.

[0075] In step 610 , the deep learning network model may be, for example, a large language model, to generate incremental reply content as target output text based on the user's second instruction.

[0076] In step 620, illustratively, the large language model usually outputs the target output text in the form of segments, so when the digital human performs voice broadcasting, it can immediately convert each segment of the target output subtext into audio data for broadcasting when it receives the segment of the target output subtext, or it can simultaneously convert multiple target output subtexts into audio data for broadcasting when the number or size of the received target output subtexts is greater than a specific threshold, and it can also convert the completed target output text into audio data for broadcasting when it is received.

[0077] According to some embodiments, method 200 further includes sending the target output text to a display device for display to the user. For example, the display device may be a display screen of the digital human. Therefore, even if the user issues an interrupt command causing the digital human to stop the voice broadcast, the complete text content corresponding to the voice broadcast will still be displayed to the user in text form, providing a better interactive experience for the user.

[0078] Figure 7 A partial flow chart of another method for processing digital human voice interaction according to an embodiment of the present disclosure is shown.

[0079] According to some embodiments, Figure 7 Before step 230, the method 200 further includes:

[0080] Step 710: Perform environmental noise reduction processing on the third audio data to update the third audio data; and

[0081] Step 720: Determine whether the user issues a first instruction to interrupt the voice interaction based on the updated third audio data.

[0082] For example, the environmental noise reduction processing in step 710 may target noises such as other human voices, traffic noise, wind and rain noises, etc., other than the voice of the user performing voice interaction. This can further reduce noise interference in the audio data used for user interruption command recognition, further improving the accuracy of user command recognition.

[0083] Figure 8 The structural block diagram of the processing device of digital human voice interaction according to an embodiment of the present disclosure is shown.

[0084] According to another aspect of the present disclosure, Figure 8 As shown, a processing device 800 for digital human voice interaction is provided, comprising: an acquisition module 810, configured to acquire first voiceprint data of the digital human in response to receiving first audio data, wherein the first audio data indicates that the user has initiated a voice interaction request to the digital human, and the first voiceprint data indicates the voiceprint data currently used by the digital human for voice broadcasting; a first noise reduction module 820, configured to perform first noise reduction processing on the second audio data based on the first voiceprint data in response to receiving second audio data, to obtain third audio data, wherein the time when the second audio data is received is after the time when the first audio data is received; and a processing module 830, configured to generate a stop instruction in response to determining, based on the third audio data, that the user has issued a first instruction to interrupt the voice interaction, and send the stop instruction to the digital human to control the digital human to stop voice broadcasting.

[0085] According to some embodiments, the acquisition module 810 includes: a first recognition module 811, configured to perform speech recognition on the first audio data using a first speech recognition model in response to receiving the first audio data, to obtain a first recognition text; and an acquisition sub-module 812, configured to obtain the first voiceprint data in response to determining, based on the first recognition text, that the user initiates a voice interaction request to the digital human.

[0086] According to some embodiments, the second audio data includes at least the user's voice audio, and the device 800 also includes: a first determination module 840, configured to determine the user's second voiceprint data based on the first audio data; and a second determination module 850, determining whether the second audio data is received based on the second voiceprint information.

[0087] According to some embodiments, the processing module 830 includes: a second recognition module 831, configured to use a second speech recognition model to perform speech recognition on the third audio data to obtain a second recognition text; and a generation submodule 832, configured to generate the stop instruction and send it to the digital human to control the digital human to stop voice broadcasting in response to determining that the user has issued the first instruction based on the second recognition text.

[0088] According to some embodiments, the device 800 also includes: an output module 860, which is configured to process the second recognition text using a deep learning network model in response to determining that the user has issued a second instruction to continue the voice interaction based on the second recognition text, to obtain a target output text, wherein the target output text indicates the digital human's reply content to the second instruction; and a conversion module 870, which is configured to convert the target output text into fourth audio data for the digital human to perform voice broadcasting.

[0089] According to some embodiments, the apparatus 800 further includes: a display module 880 configured to send the target output text to a display device for displaying to a user.

[0090] According to some embodiments, the device 800 also includes: a second noise reduction module 890, configured to perform environmental noise reduction processing on the third audio data to update the third audio data; and a third determination module 8100, configured to determine whether the user issues a first instruction to interrupt the voice interaction based on the updated third audio data.

[0091] According to another aspect of the present disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the aforementioned method.

[0092] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is further provided, wherein the computer instructions are used to enable the computer to execute the aforementioned method.

[0093] According to another aspect of the present disclosure, a computer program product is further provided, including a computer program, wherein the computer program implements the aforementioned method when executed by a processor.

[0094] like Figure 9As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0095] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, an output unit 907, a storage unit 908, and a communication unit 909. The input unit 906 can be any type of device that can input information to the electronic device 900. The input unit 906 can receive input digital or character information, and generate key signal input related to user settings and / or function control of the electronic device, and can include but is not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone and / or a remote control. The output unit 907 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator and / or a printer. The storage unit 908 can include but is not limited to a magnetic disk, an optical disk. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver and / or a chipset, such as Bluetooth TM devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0096] The computing unit 901 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as the GPU-based matrix calculation method. For example, in some embodiments, the GPU-based matrix calculation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the GPU-based matrix calculation method described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the GPU-based matrix calculation method by any other appropriate means (e.g., by means of firmware).

[0097] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0098] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0099] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0100] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0101] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0102] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0103] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0104] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-mentioned methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only limited by the claims after authorization and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. In addition, the steps may be performed in an order different from that described in this disclosure. Further, the various elements in the embodiments or examples may be combined in various ways. It is important that as technology evolves, many of the elements described herein may be replaced by equivalent elements that appear after this disclosure.

Claims

1. A method for processing digital human voice interaction, comprising: In response to receiving the first audio data, obtaining first voiceprint data of the digital human, wherein the first audio data indicates that the user initiates a voice interaction request to the digital human, and the first voiceprint data indicates the voiceprint data currently used by the digital human for voice broadcasting; In response to receiving second audio data, performing a first noise reduction process on the second audio data based on the first voiceprint data to obtain third audio data, wherein the second audio data is received after the first audio data is received; and In response to determining, based on the third audio data, that the user issues a first instruction to interrupt the voice interaction, generating a stop instruction and sending it to the digital human to control the digital human to stop the voice broadcast, The second audio data at least includes the user's voice audio, and the method further includes: Determining second voiceprint data of the user based on the first audio data; and determining whether the second audio data is received according to the second voiceprint data, And wherein, the operation of performing the first noise reduction process on the second audio data is performed after determining that the second audio data including the user's voice audio is received.

2. The method according to claim 1, wherein The step of obtaining first voiceprint data of the digital person in response to receiving the first audio data includes: In response to receiving the first audio data, performing speech recognition on the first audio data using a first speech recognition model to obtain a first recognized text; and In response to determining, based on the first recognition text, that the user initiates a voice interaction request to the digital human, the first voiceprint data is acquired.

3. The method according to claim 1, wherein In response to determining, based on the third audio data, that the user issues a first instruction to interrupt the voice interaction, generating a stop instruction and sending it to the digital human to control the digital human to stop the voice broadcast, the method includes: Performing speech recognition on the third audio data using a second speech recognition model to obtain a second recognized text; and In response to determining, based on the second recognition text, that the user has issued the first instruction, a stop instruction is generated and sent to the digital human to control the digital human to stop voice broadcasting.

4. The method according to claim 3, further comprising: In response to determining, based on the second recognized text, that the user has issued a second instruction to continue voice interaction, processing the second recognized text using a deep learning network model to obtain a target output text, wherein the target output text indicates the digital human's reply content to the second instruction; and The target output text is converted into fourth audio data for the digital human to perform voice broadcasting.

5. The method according to claim 4, further comprising: The target output text is sent to a display device to be displayed to the user.

6. The method according to any one of claims 1 to 5, wherein Before the step of determining, based on the third audio data, that the user issues a first instruction to interrupt the voice interaction, the method further includes: performing environmental noise reduction processing on the third audio data to update the third audio data; and Determine whether the user issues the first instruction to interrupt the voice interaction based on the updated third audio data.

7. A digital human voice interaction processing device, comprising: an acquisition module configured to acquire first voiceprint data of the digital human in response to receiving first audio data, wherein the first audio data indicates that the user initiates a voice interaction request to the digital human, and the first voiceprint data indicates the voiceprint data currently used by the digital human for voice broadcasting; a first noise reduction module configured to, in response to receiving second audio data, perform a first noise reduction process on the second audio data based on the first voiceprint data to obtain third audio data, wherein the second audio data is received after the first audio data is received; and a processing module configured to generate a stop instruction and send it to the digital human to control the digital human to stop the voice broadcast in response to determining, based on the third audio data, that the user issues a first instruction to interrupt the voice interaction; The second audio data at least includes the user's voice audio, and the device further includes: a first determining module, configured to determine second voiceprint data of the user according to the first audio data; and A second determining module is configured to determine whether the second audio data is received according to the second voiceprint data, And wherein, the operation of performing the first noise reduction process on the second audio data is performed after determining that the second audio data including the user's voice audio is received.

8. The device according to claim 7, wherein The acquisition module includes: a first recognition module configured to, in response to receiving the first audio data, perform speech recognition on the first audio data using a first speech recognition model to obtain a first recognized text; and The acquisition submodule is configured to acquire the first voiceprint data in response to determining, based on the first recognition text, that the user initiates a voice interaction request to the digital human.

9. The device according to claim 7, wherein The processing module includes: a second recognition module configured to perform speech recognition on the third audio data using a second speech recognition model to obtain a second recognized text; and The generating submodule is configured to generate the stop instruction and send it to the digital human to control the digital human to stop voice broadcasting in response to determining that the user has issued the first instruction based on the second recognition text.

10. The apparatus according to claim 9, further comprising: an output module configured to, in response to determining, based on the second recognized text, that the user has issued a second instruction to continue voice interaction, process the second recognized text using a deep learning network model to obtain a target output text, wherein the target output text indicates the digital human's reply content to the second instruction; and The conversion module is configured to convert the target output text into fourth audio data for the digital human to perform voice broadcasting.

11. The apparatus according to claim 10, further comprising: The display module is configured to send the target output text to a display device for displaying to the user.

12. The apparatus according to any one of claims 7 to 11, further comprising: a second noise reduction module, configured to perform environmental noise reduction processing on the third audio data to update the third audio data; as well as The third determination module is configured to determine whether the user issues the first instruction to interrupt the voice interaction based on the updated third audio data.

13. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.

15. A computer program product comprising a computer program, wherein When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Method for realizing voice control, robot and computer-readable medium

    CN107610698A

  • Method and device for man-machine conversation

    CN113779208A