Control method, device and electronic equipment for human-computer voice dialogue

By obtaining the voice flow state characteristics in the human-machine voice dialogue and selecting control instructions to control the machine-side response, the problem of high response delay in the existing technology is solved, the voice duplex dialogue mode is realized, and the user experience is improved.

CN114999470BActive Publication Date: 2025-10-03ALIBABA INNOVATION PRIVATE LIMITED
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110229744.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-02
Publication Date
2025-10-03
Estimated Expiration
2041-03-02

AI Technical Summary

Technical Problem

Existing human-computer voice dialogue methods have the problem of not being able to provide timely and accurate feedback during the dialogue process, resulting in high response delays and poor user experience.

Method used

By receiving the voice streams from the user and machine ends, obtaining the state characteristics of the voice streams in the time slice, selecting the corresponding control instructions to control the response of the machine end, and realizing the voice duplex dialogue mode, it ensures that the machine end responds to the user's voice content in a timely and accurate manner at any time.

Benefits of technology

It reduces response delay, improves user experience, and makes human-computer voice dialogue smoother and more efficient.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114999470B_ABST
    Figure CN114999470B_ABST
Patent Text Reader

Abstract

The present application discloses a method for controlling human-machine voice dialogue, including: receiving a first voice stream for human-machine voice dialogue on a user side and a second voice stream for human-machine voice dialogue on a monitoring machine side; obtaining a first state feature of the first voice stream at a first time slice and a second state feature of the second voice stream at the first time slice; selecting corresponding control instructions from a set control instruction set based on the first state feature and the second state feature; wherein the control instruction set includes an instruction for controlling the machine side to broadcast and an instruction for controlling the machine side to mute; after the first time slice, controlling the machine side to conduct the human-machine voice dialogue according to the matching control instructions. This method enables an electronic device to timely and accurately control the machine side to respond to the voice stream sent by the user at any time, thereby reducing response delay and improving user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and more specifically, to a method and device for controlling human-computer voice dialogue, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the continuous development of computer technology, human-computer dialogue technology, especially human-computer voice dialogue technology, has been widely used in various fields, greatly facilitating people's lives.

[0003] Currently, when implementing human-computer voice dialogue, it is usually based on the human-computer text dialogue mode. For example, when a user interacts with a smart speaker, the user needs to wake up the device first, then speak, and then the smart speaker responds based on the user's voice; after that, the user needs to wake up the device again and speak again before the device responds again.

[0004] In the process of implementing this application, the inventor discovered that, unlike text conversations, voice conversations are often characterized by continuity and exclusivity. When one party conveys voice information, the other party can simultaneously understand the information and interrupt to make a timely response. The current human-computer voice dialogue method is based on the human-computer text dialogue mode, and there is a problem of not being able to provide timely and accurate dialogue feedback during the dialogue process. Summary of the Invention

[0005] One purpose of the embodiments of the present disclosure is to provide a new technical solution for controlling human-computer voice dialogue.

[0006] In a first aspect of the present disclosure, a method for controlling human-computer voice dialogue is provided, the method comprising:

[0007] Receive a first voice stream of a human-machine voice dialogue at a user end and a second voice stream of the human-machine voice dialogue at a monitoring machine end;

[0008] Acquire a first state feature of the first voice stream at a first time slice and a second state feature of the second voice stream at the first time slice;

[0009] Selecting corresponding control instructions from a set of control instructions based on the first state characteristic and the second state characteristic; wherein the control instruction set includes an instruction for controlling the machine end to broadcast and an instruction for controlling the machine end to mute;

[0010] After the first time slice, the machine end is controlled according to the control instruction to perform the human-machine voice dialogue.

[0011] Optionally, the instructions for controlling the machine-side broadcast include at least one of a first control instruction for continuing the current broadcast, a second control instruction for starting a new broadcast, a third control instruction for broadcasting the set sentence continuation content, a fourth control instruction for broadcasting the set first round of question and answer content, and a fifth control instruction for broadcasting the set mute prompt content; and / or,

[0012] The instruction for controlling the machine end to be muted includes at least one of a sixth control instruction for stopping the current broadcast and a seventh control instruction for keeping the machine end muted.

[0013] Optionally, obtaining a first state feature of the first voice stream at a first time slice and a second state feature of the second voice stream at the first time slice includes:

[0014] Detecting the occurrence of a triggering event;

[0015] According to the detected trigger event, a first state feature of the first voice stream in a first time slice before the trigger event is detected and a second state feature of the second voice stream in the first time slice are acquired.

[0016] Optionally, the trigger event includes at least one of an event of starting the human-computer voice dialogue, an event of a non-silent segment appearing in the first voice stream, an event of a silent segment appearing in the first voice stream, an event of a non-silent segment appearing in the second voice stream, an event of a silent segment appearing in the second voice stream, and the arrival of a set trigger time.

[0017] Optionally, the triggering event includes an event in which a non-silent segment appears in the first voice stream, and the step of detecting the event in which a non-silent segment appears in the first voice stream includes:

[0018] Splitting the first voice stream to obtain a first silent segment and a second silent segment that are adjacent to each other, wherein the first silent segment is earlier than the second silent segment;

[0019] When the timing of the first silent segment and the second silent segment are not connected, the voice segment between the first silent segment and the second silent segment is extracted as a non-silent segment, and an event in which a non-silent segment occurs in the first voice stream is determined.

[0020] Optionally, the selecting, from a set control instruction set, control instructions corresponding to the first state characteristic and the second state characteristic includes:

[0021] Determining, based on the first state feature and the second state feature, whether the machine terminal has a speaking right after the first time slice, and obtaining a determination result;

[0022] According to the judgment result, a control instruction corresponding to the first state feature and the second state feature is selected from the control instruction set.

[0023] Optionally, determining whether the machine terminal has the speaking right after the first time slice based on the first state feature and the second state feature, and obtaining a determination result, includes:

[0024] If the second state characteristic is that the machine end remains non-muted or the machine end changes from being muted to being non-muted, determining that the machine end has the speaking right after the first time slice;

[0025] The instruction for controlling the broadcast of the machine end includes a first control instruction for continuing the current broadcast, and selecting, from the control instruction set according to the judgment result, a control instruction corresponding to the first state feature and the second state feature, includes:

[0026] When the machine end has the right to speak, the first control instruction is selected as the corresponding control instruction.

[0027] Optionally, determining whether the machine terminal has the speaking right after the first time slice based on the first state feature and the second state feature, and obtaining a determination result, includes:

[0028] If the first state feature indicates that a non-silent segment appears in the first voice stream, and the second state feature indicates that the machine terminal remains non-silent or changes from being silent to being non-silent, determining that the machine terminal does not have a speaking right after the first time slice;

[0029] The instruction for controlling the machine end to mute includes a sixth control instruction for stopping the current broadcast, and selecting, based on the judgment result, the control instruction corresponding to the first state feature and the second state feature in the control instruction set includes:

[0030] When the machine end does not have the right to speak, the sixth control instruction is selected as the corresponding control instruction.

[0031] Optionally, the first state feature indicating that a non-silent segment appears in the first voice stream includes: the user terminal changes from silent to non-silent and / or the user terminal changes from non-silent to silent.

[0032] Optionally, determining whether the machine terminal has the speaking right after the first time slice based on the first state feature and the second state feature, and obtaining a determination result, includes:

[0033] When both the first state feature and the second state feature are in a conversation start state, determining that the machine end has the speaking right after the first time slice;

[0034] The control machine end broadcasting instruction includes a fourth control instruction for broadcasting the set first round of question and answer content, and the control instruction corresponding to the first state feature and the second state feature in the control instruction set according to the judgment result includes:

[0035] When the machine end has the right to speak, the fourth control instruction is selected as the corresponding control instruction.

[0036] Optionally, determining whether the machine terminal has the speaking right after the first time slice based on the first state feature and the second state feature, and obtaining a determination result, includes:

[0037] When the first state characteristic is that the user terminal changes from being non-mute to being mute and the second state characteristic is that the machine terminal remains mute, determining that the machine terminal has the speaking right after the first time slice;

[0038] The instructions for controlling the machine-side broadcast include a second control instruction for starting a new broadcast and / or a third control instruction for broadcasting the continuation content of a set sentence, and selecting, from the control instruction set based on the judgment result, the control instructions corresponding to the first state feature and the second state feature, includes:

[0039] When the machine end has the right to speak, the second control instruction or the third control instruction is selected as the corresponding control instruction.

[0040] Optionally, determining whether the machine terminal has the speaking right after the first time slice based on the first state feature and the second state feature, and obtaining a determination result, includes:

[0041] When the first state characteristic is that the user terminal changes from being non-mute to being mute and the second state characteristic is that the machine terminal remains mute, determining that the machine terminal does not have a speaking right after the first time slice;

[0042] The instruction for controlling the device to be muted includes a seventh control instruction for keeping the device muted. Selecting, from the control instruction set, the control instructions corresponding to the first state characteristic and the second state characteristic based on the judgment result includes:

[0043] When the machine end does not have the right to speak, the seventh control instruction is selected as the corresponding control instruction.

[0044] Optionally, determining whether the machine terminal has the speaking right after the first time slice based on the first state feature and the second state feature, and obtaining a determination result, includes:

[0045] When the first state characteristic is that the user terminal remains silent and the second state characteristic is that the machine terminal changes from being non-silent to being silent, determining that the machine terminal has the speaking right after the first time slice;

[0046] The instruction for controlling the machine end to broadcast includes a fifth control instruction for broadcasting a set mute prompt content, and the control instruction corresponding to the first state feature and the second state feature in the control instruction set is selected according to the judgment result, including:

[0047] When the machine end has the right to speak, the fifth control instruction is selected as the corresponding control instruction.

[0048] Optionally, determining whether the machine terminal has the speaking right after the first time slice based on the first state feature and the second state feature, and obtaining a determination result, includes:

[0049] When the first state characteristic is that the user terminal remains silent and the second state characteristic is that the machine terminal changes from being non-silent to being silent, determining that the machine terminal does not have a speaking right after the first time slice;

[0050] The instruction for controlling the device to be muted includes a seventh control instruction for keeping the device muted. Selecting, from the control instruction set, the control instructions corresponding to the first state characteristic and the second state characteristic based on the judgment result includes:

[0051] When the machine end does not have the right to speak, the seventh control instruction is selected as the corresponding control instruction.

[0052] Optionally, controlling the machine end to perform the human-machine voice dialogue according to the control instruction includes:

[0053] The control instruction is sent to the machine end, so that the machine end performs the human-machine voice dialogue according to the control instruction.

[0054] Optionally, the machine side performs the human-machine voice dialogue according to the corresponding control instruction, including:

[0055] The machine end obtains response information corresponding to the corresponding control instruction according to pre-stored mapping data, wherein the mapping data reflects the corresponding relationship between each control instruction in the control instruction set and each set response information;

[0056] Conduct human-computer voice dialogue based on the obtained response information.

[0057] A second aspect of the present disclosure further provides a control device for human-machine voice dialogue, comprising:

[0058] A voice stream receiving module, configured to receive a first voice stream of a user terminal for human-computer voice dialogue;

[0059] The voice stream monitoring module is used to monitor the machine segment and the second voice stream of the reverse human-machine voice dialogue.

[0060] A state acquisition module, configured to acquire a first state feature of the first voice stream at a first time slice and a second state feature of the second voice stream at the first time slice;

[0061] a decision module, configured to select control instructions corresponding to the first state characteristic and the second state characteristic from a set control instruction set; wherein the control instruction set includes an instruction for controlling the machine end to broadcast and an instruction for controlling the machine end to mute; and

[0062] An execution module is used to control the machine end to perform the human-machine voice dialogue according to the control instruction after the first time slice.

[0063] According to a third aspect of the present disclosure, an electronic device is further provided, including the apparatus according to the second aspect of the present disclosure; or including:

[0064] a memory for storing executable instructions;

[0065] The processor is configured to operate the electronic device to execute the method according to the first aspect of the present disclosure under the control of the executable instructions.

[0066] According to a fourth aspect of the present disclosure, a computer-readable storage medium is further provided, wherein the computer-readable storage medium stores a computer program that can be read and executed by a computer, and the computer program is used to execute the method according to the first aspect of the present disclosure when read and executed by the computer.

[0067] According to an embodiment of the present disclosure, during a human-machine voice conversation, an electronic device receives a first voice stream from a user end for the human-machine voice conversation and monitors a second voice stream from a machine end for the human-machine voice conversation. This method eliminates the need to wait for the user to issue a voice call before controlling the machine end to respond. Instead, during the process of receiving the first voice stream, the electronic device obtains a first state feature of the first voice stream at a first time slice and a second state feature of the second voice stream at the first time slice. Based on the first state feature and the second state feature, the electronic device selects a control instruction for controlling the machine end to conduct the human-machine voice conversation after the first time slice. The electronic device then controls the machine end to respond promptly and accurately to the user end's output voice after the first time slice according to the control instruction. This method allows the human-machine voice conversation to be conducted in voice duplex mode, allowing the electronic device to promptly and accurately control the machine end to respond to the user's voice stream at any time, reducing response delays and improving user experience.

[0068] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0070] Figure 1 This is a schematic diagram of data processing in the existing human-computer voice dialogue process provided by an embodiment of the present disclosure.

[0071] Figure 2 This is a scenario diagram of the method for controlling human-computer voice dialogue provided by an embodiment of the present disclosure.

[0072] Figure 3 The figure is a hardware configuration diagram of a human-machine voice dialogue control system that can be used to implement the human-machine voice dialogue control method of the embodiment of the present disclosure.

[0073] Figure 4 This is a flowchart of a method for controlling human-computer voice dialogue provided by an embodiment of the present disclosure.

[0074] Figure 5 It is a schematic diagram of obtaining non-silent segments provided by an embodiment of the present disclosure.

[0075] Figure 6 This is a schematic diagram of the architecture for controlling human-computer voice dialogue provided by an embodiment of the present disclosure.

[0076] Figure 7 This is a schematic principle block diagram of a control device for human-machine voice dialogue provided by an embodiment of the present disclosure.

[0077] Figure 8a It is a schematic principle block diagram of an electronic device according to an embodiment of the present disclosure.

[0078] Figure 8b is a schematic principle block diagram of an electronic device according to another embodiment of the present disclosure. DETAILED DESCRIPTION

[0079] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure.

[0080] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.

[0081] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0082] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0083] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0084] In the process of realizing this application, the inventors found that the current implementation principle of controlling human-computer voice dialogue is generally based on the existing text dialogue, turn-based conversation interaction (TBC) mode, by adding automatic speech recognition technology (ASR, Automatic Speech Recognition) and text-to-speech synthesis (TTS, Text-to-speech) technology. Figure 1As shown, the existing method is generally that an electronic device uses built-in automatic speech recognition technology to obtain the text of a round of speech uttered by a user; obtains the semantic information of the text through natural language understanding (NLU); obtains the response semantic information of the semantic information through a dialogue control module; obtains the response text corresponding to the response semantic information through natural language generation (NLG) technology; and then synthesizes and controls the output of the response speech for the response text through text-to-speech synthesis technology.

[0085] As described in the background technology, the current methods for controlling human-computer voice dialogues do not address the continuous and exclusive characteristics of voice dialogues. Therefore, during the human-computer dialogue process, the machine side may not be able to provide timely and accurate dialogue feedback, resulting in high response delays and a poor user experience.

[0086] In order to solve the above problems, the present invention provides a method for controlling human-machine voice dialogue that can realize voice duplex, that is, a method for enabling two or more parties to conduct a voice dialogue without blocking each other during the dialogue process, so that the dialogue can proceed at any time. Figure 2 , which is a scene diagram of the control method of human-computer voice dialogue provided by the embodiment of the present disclosure. In practice, the method provided by this embodiment can be applied to the voice dialogue robot of an enterprise, for example, in the hotline intelligent customer service, wherein the language dialogue robot can be an electronic device that provides pre-sale and after-sale services to users through voice, for example, Figure 2 The server 1100 shown. In a specific implementation, the user establishes a connection with the server 1100 through the terminal device 1200 and sends a first voice stream to the server 1100. For example, a voice stream can be sent to inquire about product usage issues. The server 1100 can receive the first voice stream and monitor and obtain the second voice stream for the human-computer voice dialogue. Thereafter, by respectively obtaining the first state feature of the first voice stream at the first time slice and the second state feature of the second voice stream at the first time slice, a control instruction for controlling the human-computer voice dialogue after the first time slice is selected according to the first state feature and the second state feature. Then, after the first time slice, response information for the human-computer voice dialogue is obtained according to the control instruction, and according to the response information, the second voice stream is continuously output to respond to the voice content in the first voice stream sent by the user in a timely and accurate manner.

[0087] For example, when the user is describing a long question and the server 1100 remains silent, at a certain time slice, the server 1100 detects that a silent segment appears in the voice segment of the first voice stream, that is, within the time slice, the first state feature of the first voice stream indicates that the user end changes from a non-silent state to a silent state, and the second state feature of the second voice stream indicates that the machine end remains silent. The control instruction after the time slice can be selected as the control instruction for controlling the content to be continued in the sentence set for broadcast, so as to control the content to be continued in the sentence to be broadcast according to the control instruction, that is, within the time slice. After that, the server 1100 can issue a mid-sentence connection voice, such as "um", "yes", "please continue", etc., to indicate that the current dialogue connection is normal; or, within the time slice, if the server 1100 can understand the user's question based on the voice issued by the user before the time slice, it can determine that the control instruction after the time slice is the control instruction to start a new broadcast, so as to interrupt the user's current voice after the time slice according to the control instruction, and control the broadcast of the response voice to the user's question, so that the user's question can be responded to in time, reduce the response delay, and improve the user experience.

[0088] It should be noted that the above is a scenario in which the method can be implemented. When it is implemented specifically, the method can also be applied to other scenarios. For example, in the field of Internet of Things (IOT), the method can be used in intelligent voice interaction devices so that the intelligent interaction devices can interact with users in a timely and accurate manner. For example, for smart speakers, different from the turn-based dialogue interaction mode, the smart speaker can receive the first voice stream sent by the user after receiving the user's wake-up word, and monitor the second voice stream sent by itself. In this process, according to the first state characteristics of the first voice stream in a certain time slice and the second state characteristics of the second voice stream in the time slice, the corresponding control instruction is selected to interact with the user according to the control instruction. For example, when a user asks the smart speaker "What's the weather like today?", the smart speaker can interrupt the user as soon as the user utters "today's weather" and play the current weather information to the user, thereby reducing response delay and improving user experience; of course, if the user is not asking about today's weather information, but "Should I go traveling based on today's weather?", the user can continue to utter the voice "Should I go traveling" while the smart speaker is playing the current weather information based on the voice "today's weather", and the smart speaker can interrupt the currently playing weather information based on the real-time voice and output a response voice like "Today's weather is suitable for visiting indoor attractions. Based on the distance, we recommend going to xxx museum."

[0089] Figure 3The figure is a hardware configuration diagram of a control system for human-machine voice dialogue that can be used to implement the control method for human-machine voice dialogue according to an embodiment of the present disclosure.

[0090] like Figure 3 As shown, the control system 1000 for human-machine voice dialogue in this embodiment includes a server 1100 , a terminal device 1200 and a communication network 1300 .

[0091] The server 1100 may be, for example, a blade server, a rack server, etc. The server 1100 may also be a server cluster deployed in the cloud, which is not limited here.

[0092] like Figure 3 As shown, server 1100 may include a processor 1110, a memory 1120, an interface device 1130, a communication device 1140, a display device 1150, and an input device 1160. Processor 1110 may be, for example, a central processing unit (CPU). Memory 1120 may include, for example, ROM (read-only memory), RAM (random access memory), or a non-volatile memory such as a hard disk. Interface device 1130 may include, for example, a USB interface or a serial interface. Communication device 1140 may be capable of wired or wireless communication. Display device 1150 may be, for example, a liquid crystal display. Input device 1160 may include, for example, a touch screen or a keyboard.

[0093] In this embodiment, the server 1100 may be used to participate in implementing the method according to any embodiment of the present disclosure.

[0094] As used in the embodiments of the present disclosure, the memory 1120 of the server 1100 is used to store instructions for controlling the processor 1110 to perform operations to support implementation of the methods according to any embodiment of the present disclosure. A skilled person can design instructions based on the solutions disclosed herein. How instructions control processor operations is well known in the art and will not be described in detail here.

[0095] It should be understood by those skilled in the art that although Figure 3 , multiple devices of the server 1100 are shown; however, the server 1100 of the embodiment of the present disclosure may only involve some of the devices, for example, only the processor 1110 and the memory 1120.

[0096] like Figure 3As shown, terminal device 1200 may include a processor 1210, a memory 1220, an interface device 1230, a communication device 1240, a display device 1250, an input device 1260, an audio output device 1270, an audio input device 1280, and the like. Processor 1210 may be a central processing unit (CPU), a microprocessor (MCU), or the like. Memory 1220 may include, for example, ROM (read-only memory), RAM (random access memory), or a non-volatile memory such as a hard disk. Interface device 1230 may include, for example, a USB interface or a headphone jack. Communication device 1240 may be capable of wired or wireless communication. Display device 1250 may be, for example, an LCD display or a touchscreen display. Input device 1260 may include, for example, a touchscreen or a keyboard. Terminal device 1200 may output audio information via audio output device 1270, which may include, for example, a speaker. Terminal device 1200 may also capture user input voice information via audio pickup device 1280, which may include, for example, a microphone.

[0097] The terminal device 1200 can be a smart phone, a laptop computer, a desktop computer, a tablet computer, a wearable device, etc.

[0098] It should be understood by those skilled in the art that although Figure 3 , multiple devices of the terminal device 1200 are shown, however, the terminal device 1200 of the embodiment of the present disclosure may only involve some of the devices, for example, only the processor 1210, the memory 1220, etc.

[0099] The communication network 1300 may be a wireless network or a wired network, a local area network or a wide area network. The terminal device 1200 may communicate with the server 1100 via the communication network 1300 .

[0100] Figure 3 The control system 1000 for human-machine speech dialogue is illustrative only and is in no way intended to limit the present disclosure, its application, or use. Figure 3 Only one server 1100 and one terminal device 1200 are shown, but this does not mean to limit the respective numbers. The system 1000 may include multiple servers 1100 and / or multiple terminal devices 1200.

[0101] It should be noted that the method provided in any embodiment of the present disclosure can be used in the server 1100. Of course, in specific implementation, the method can also be applied to the terminal device 1200 as needed, and there is no special limitation here.

[0102] Figure 4The method provided in this embodiment can be applied to electronic devices, for example, Figure 3 In the server 1100 shown. In addition, in this embodiment, unless otherwise specified, a scenario in which a user interacts with a server through a terminal device via a human-computer voice dialogue is used as an example for description, that is, the user sends a first voice stream to the server through the terminal device, and the server generates response information in real time based on the first voice stream, and sends a second voice stream to the terminal device based on the response information.

[0103] like Figure 4 As shown, the method for controlling human-machine voice dialogue in this embodiment may include the following steps S4100-S4400, which are described in detail below.

[0104] Step S4100: receiving a first voice stream of a human-machine voice dialogue at a user end and a second voice stream of the human-machine voice dialogue at a monitoring machine end.

[0105] In this embodiment, the first voice stream is a data stream formed by the voice uttered by the user during a human-computer voice dialogue at the user end. The first voice stream can be generated by an audio pickup device of the user terminal device, such as a microphone, to collect the voice uttered by the user.

[0106] The second voice stream is a data stream formed by the voice emitted by the machine end during the human-machine voice dialogue on the machine end, that is, a data stream formed by the machine end responding to the voice content in the first voice stream; in a specific implementation, the voice in the voice stream can, for example, be obtained by the server by recognizing the voice in the first voice stream, obtaining the corresponding response semantics, and obtaining the corresponding response text through natural language generation technology, and then obtaining the corresponding response voice through text-to-speech synthesis technology, and forming the voice in the voice stream by outputting the response voice.

[0107] Step S4200: Acquire a first state feature of the first voice stream at a first time slice and a second state feature of the second voice stream at the first time slice.

[0108] In practice, when starting or conducting a human-computer voice dialogue, there are often short periods of silence between words, sentences, or paragraphs when the user speaks. For example, when a user asks a hotline intelligent robot about a product issue, when they say "Hello, I want to...", after the user speaks the corresponding voice of "Hello", there is usually a period of silence, such as about 300ms, before continuing to speak.

[0109] Therefore, in order to enable the machine end to respond to the real-time voice content emitted by the user end in a timely and accurate manner, the electronic device can trigger the control machine end to respond by detecting a specific trigger event. In this embodiment, the trigger event can be, for example: an event of starting a human-computer voice dialogue, an event of a non-silent segment appearing in the first voice stream, an event of a silent segment appearing in the first voice stream, an event of a non-silent segment appearing in the second voice stream, an event of a silent segment appearing in the second voice stream, and at least one of the following: an event of starting a human-computer voice dialogue, an event of a non-silent segment appearing in the first voice stream, an event of a silent segment appearing in the second voice stream, and an event of reaching a set trigger time.

[0110] Among them, the silent segment can be a voice segment with a silent duration within a preset duration, and the preset duration can be, for example, 200ms-400ms. Of course, the preset duration can also be set as needed, and there is no special limitation here; correspondingly, the non-silent segment can be a voice segment containing voice content.

[0111] Specifically, the obtaining of the first state characteristics of the first voice stream in the first time slice and the second state characteristics of the second voice stream in the first time slice includes: detecting the occurrence of a trigger event; and according to the detected trigger event, obtaining the first state characteristics of the first voice stream in the first time slice before the trigger event is detected and the second state characteristics of the second voice stream in the first time slice.

[0112] In this embodiment, a time slice can be a slice whose start and end times are the occurrence times of two adjacent triggering events. For example, in a voice stream, event 1 is a non-silent segment, event 2 is a silent segment, and events 1 and 2 are temporally connected. When event 2 is detected, a time slice can be obtained with the start time of event 1 and the start time of event 2 as the start and end times, respectively. This time slice can be considered the time slice before event 2. Of course, in specific implementations, time slices can be divided as needed, and this is not specifically limited here.

[0113] The first state feature may be a feature representing the state change of the first voice stream at the user end within a time slice, for example, characterizing the first voice stream changing from silent to non-silent, non-silent to silent, continuous silent, and continuous non-silent, etc. Correspondingly, the second state feature may be a feature representing the state change of the second voice stream within the corresponding time slice.

[0114] In one embodiment, when the triggering event includes the time when a non-silent segment appears in the first voice stream, the step of detecting the event of the non-silent segment appearing in the first voice stream includes: splitting the first voice stream to obtain adjacent first silent segments and second silent segments, wherein the first silent segment is earlier than the second silent segment; when the timing of the first silent segment and the second silent segment are not connected, extracting the voice segment between the first silent segment and the second silent segment as a non-silent segment, and determining the event of the non-silent segment appearing in the first voice stream.

[0115] That is, in a specific implementation, the silent segments in the voice stream emitted by the user end can be identified, and the silent segments can be used as the voice boundaries in the voice stream, and the voice segments between two adjacent silent segments that are not connected in time sequence can be used as non-silent segments, and then the events of the occurrence of non-silent segments in the voice stream can be determined. Among them, the silent segments in the voice stream can be detected by voice activity detection (VAD, Voice Activity Detection) technology, and its detailed processing process will not be repeated here.

[0116] Please see Figure 5 , which is a schematic diagram of obtaining non-silent segments provided by an embodiment of the present disclosure. Figure 5 As shown, in a specific implementation, the silent segments in the voice stream can be identified as voice boundaries, and the voice segments between two adjacent voice boundaries that are not connected in time sequence can be regarded as non-silent segments, that is, the non-silent segments can be regarded as the smallest processing unit in the human-computer voice dialogue, that is, micro-turn.

[0117] To sum up, after detecting the occurrence of a trigger event and obtaining the first state characteristics of the first voice stream in the first time slice before the trigger event and the second state characteristics of the second voice stream in the first time slice, the control instructions for controlling the machine end to respond after the time slice can be determined based on the real-time dialogue state of the current human-computer voice dialogue represented by the first state characteristics and the second state characteristics, which are explained in detail below.

[0118] Step S4300: Select corresponding control instructions from a set control instruction set according to the first state characteristics and the second state characteristics; wherein the control instruction set includes instructions for controlling the machine end to broadcast and instructions for controlling the machine end to mute.

[0119] During the human-machine voice dialogue, when a trigger event is detected, the machine side may make different responses to different state changes of the voice stream in the human-machine voice dialogue. For example, when the user side makes a voice for a long time, in order to indicate that the call connection is normal, the machine side can timely broadcast sentence transitions such as "um", "right", etc., or it can remain silent; or, it can also interrupt the current speech of the user side according to the situation and choose to directly broadcast the response voice; or, in the process of the machine side broadcasting the voice, the user side may interrupt the conversation, that is, the user side may interrupt the current broadcast of the machine side. In this case, the machine side can choose to continue broadcasting the voice, or choose to remain silent according to the situation, or choose to re-broadcast a new response voice according to the voice uttered by the user when the machine side interrupted.

[0120] Therefore, in this embodiment, in order to facilitate the electronic device to accurately generate control instructions based on the state changes of the voice streams respectively sent by the user end and the machine end in the human-machine voice dialogue, in this embodiment, the control instruction set set may include at least one of the following Table 1:

[0121] Table 1:

[0122]

[0123] In this embodiment, the instructions for controlling the broadcast on the machine side may include at least one of a first control instruction for continuing the current broadcast, a second control instruction for starting a new broadcast, a third control instruction for broadcasting the set content in the sentence, a fourth control instruction for broadcasting the set first round of question and answer content, and a fifth control instruction for broadcasting the set mute prompt content; and / or, the instructions for controlling the mute of the machine side include at least one of a sixth control instruction for stopping the current broadcast and a seventh control instruction for keeping the machine side silent; of course, in specific implementation, other control instructions may also be set as needed to control the machine side to conduct human-computer voice dialogue, which will not be repeated here.

[0124] In one embodiment, the selecting of control instructions corresponding to the first state feature and the second state feature in the set control instruction set includes: judging whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature, and obtaining a judgment result; and selecting, in the control instruction set, control instructions corresponding to the first state feature and the second state feature based on the judgment result.

[0125] In specific implementation, based on the voice changes of the first voice stream and the second voice stream in the first time slice respectively represented by the first state characteristics and the second state characteristics, it can be decided first whether the machine end has the right to speak after the first time slice, that is, whether a response voice needs to be issued; if it has the right to speak, it can be further decided whether the response voice should maintain the current broadcast or interrupt the current broadcast and make a new broadcast for the user voice, etc., and if it does not have the right to speak, it can be selected according to the state represented by the second state characteristics to control the machine end to interrupt the current broadcast and remain silent, or continue to remain silent, etc.

[0126] That is, in this embodiment, the control instructions corresponding to the first state characteristics and the second state characteristics are selected in the set control instruction set, including: judging whether the machine end has the right to speak after the first time slice based on the first state characteristics and the second state characteristics, and obtaining a judgment result; based on the judgment result, selecting the control instructions corresponding to the first state characteristics and the second state characteristics in the control instruction set. The following describes in detail different situations.

[0127] For ease of explanation, please refer to Table 2, which is a schematic table for judging whether the machine end has the right to speak and selects control instructions after the first time slice based on the first state characteristics and the second state characteristics. The following is a description of each embodiment in conjunction with Table 2, where S_1 represents the first state characteristic and S_2 represents the second state characteristic.

[0128] Table 2:

[0129]

[0130]

[0131] As shown in Table 2, in one embodiment, the determining whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature to obtain a judgment result includes: when the second state feature is that the machine end remains non-silent or the machine end changes from silent to non-silent, determining that the machine end has the right to speak after the first time slice; the instruction to control the machine end broadcast includes a first control instruction to continue the current broadcast, and selecting the control instruction corresponding to the first state feature and the second state feature in the control instruction set based on the judgment result includes: when the machine end has the right to speak, selecting the first control instruction as the corresponding control instruction.

[0132] That is, if the second state feature of the second voice stream emitted by the machine end in the first time slice indicates that the machine end has been continuously emitting response voice before the current moment, or the machine end has begun to respond to the user voice, then at the next moment, it can be determined that the machine end has the right to speak, that is, the machine end can be controlled to continue to emit the current response voice.

[0133] Please continue to refer to Table 2. In one embodiment, the determining, based on the first state feature and the second state feature, whether the machine end has the right to speak after the first time slice, and obtaining a determination result, includes: when the first state feature indicates that a non-silent segment appears in the first voice stream, and the second state feature indicates that the machine end remains non-silent or the machine end changes from silent to non-silent, determining that the machine end does not have the right to speak after the first time slice; the instruction for controlling the machine end to be muted includes a sixth control instruction for stopping the current broadcast; and selecting, based on the determination result, the control instruction corresponding to the first state feature and the second state feature in the control instruction set includes: when the machine end does not have the right to speak, selecting the sixth control instruction as the corresponding control instruction, wherein the first state feature indicating that a non-silent segment appears in the first voice stream includes: the user end changes from silent to non-silent and / or the user end changes from non-silent to silent.

[0134] That is, when both the machine side and the user side are speaking, there is a possibility that the current speech content of the machine side is incorrect and the user side is re-describing the problem. Therefore, when this situation exists, it can be determined that the machine side does not have the right to speak after this time slice, that is, by sending the sixth control instruction to stop the current broadcast to the machine side, the machine side is controlled to stop speaking after this time slice and remain silent so as to re-understand the content of the user's speech.

[0135] Please continue to refer to Table 2. In one embodiment, the determining whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature to obtain a judgment result includes: when the first state feature and the second state feature are both in the dialogue starting state, determining that the machine end has the right to speak after the first time slice; the instruction controlling the machine end to broadcast includes a fourth control instruction for broadcasting the set first round of question and answer content, and selecting the control instruction corresponding to the first state feature and the second state feature in the control instruction set based on the judgment result includes: when the machine end has the right to speak, selecting the fourth control instruction as the corresponding control instruction.

[0136] That is, at the beginning of the human-computer voice dialogue, in order to improve the user experience, the machine side can be controlled to speak first to determine what questions the user wants to ask. For example, based on the user's latest order, the user can be asked first whether he wants to inquire about the use of the product or whether he needs to return the product, etc.

[0137] Please continue to refer to Table 2. In one embodiment, the determining whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature to obtain a judgment result includes: when the first state feature is that the user end changes from non-silent to silent and the second state feature is that the machine end remains silent, determining that the machine end has the right to speak after the first time slice; the instruction for controlling the machine end to broadcast includes a second control instruction to start a new broadcast and / or a third control instruction for continuing the content in the sentence set for the broadcast, and selecting the control instruction corresponding to the first state feature and the second state feature in the control instruction set based on the judgment result includes: when the machine end has the right to speak, selecting the second control instruction or the third control instruction as the corresponding control instruction.

[0138] That is, when the first state feature indicates that the user has stopped speaking, it may be that the user has finished describing the problem and is waiting for the machine to respond; or the problem is long and the user has taken a short break; to indicate that the connection status of the current human-computer voice dialogue is normal, it can be determined that the machine has the right to speak after this time slice, so that the machine either starts to respond directly to the user's speech content, or issues a sentence connection to indicate that the machine is continuing to listen.

[0139] Please continue to refer to Table 2. In one embodiment, the determining whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature to obtain a judgment result includes: when the first state feature is that the user end changes from non-silent to silent and the second state feature is that the machine end remains silent, determining that the machine end does not have the right to speak after the first time slice; the instruction for controlling the machine end to be silent includes a seventh control instruction for the machine end to remain silent, and selecting the control instruction corresponding to the first state feature and the second state feature in the control instruction set based on the judgment result includes: when the machine end does not have the right to speak, selecting the seventh control instruction as the corresponding control instruction.

[0140] That is, when the user end stops speaking and turns silent, since the user may not have finished describing the problem, the electronic device can determine whether the user's problem has been described based on the context content of the user end. If it has not been described, the machine end needs to be silent so that the user can continue to describe the problem. At this time, it can be determined that the machine end has no right to speak after the first time slice, and the seventh control instruction to remain silent is sent to the machine end.

[0141] Please continue to refer to Table 2. In one embodiment, the determining whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature to obtain a judgment result includes: when the first state feature is that the user end remains silent and the second state feature is that the machine end changes from non-silent to silent, determining that the machine end has the right to speak after the first time slice; the instruction controlling the machine end to broadcast includes a fifth control instruction for broadcasting the set mute prompt content, and selecting the control instruction corresponding to the first state feature and the second state feature in the control instruction set based on the judgment result includes: when the machine end has the right to speak, selecting the fifth control instruction as the corresponding control instruction.

[0142] That is, when the user side remains silent and the machine side plays the response voice to the user's speech and also switches to silent state, it can be determined that the machine side is ready to release the right to speak and wait for the user side to take over the right to speak again to ask new questions. At this time, the electronic device can determine that the machine side has the right to speak and send a control instruction to the machine side to play the set silent prompt content, where the set silent prompt content can be, for example, "Do you have any other questions?", which is not specially limited here.

[0143] Please continue to refer to Table 2. In one embodiment, the determining whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature to obtain a judgment result includes: when the first state feature is that the user end remains silent and the second state feature is that the machine end changes from non-silent to silent, determining that the machine end does not have the right to speak after the first time slice; the instruction for controlling the machine end to be silent includes a seventh control instruction for the machine end to remain silent, and selecting the control instruction corresponding to the first state feature and the second state feature in the control instruction set based on the judgment result includes: when the machine end does not have the right to speak, selecting the seventh control instruction as the corresponding control instruction.

[0144] That is, when the user side remains silent and the machine side plays the response voice to the user's speech and also switches to silent state, it can also be determined that the machine side has released the right to speak and is waiting for the user side to regain the right to speak to ask new questions. At this time, the electronic device can determine that the machine side no longer has the right to speak and send a control instruction to the machine side to remain silent.

[0145] The above describes in detail how the electronic device selects the corresponding control instruction based on the first state characteristics and the second state characteristics. After determining the control instruction, the machine end can be controlled to continue the human-computer voice dialogue according to the control instruction after the first time slice.

[0146] It should be noted that, in a specific implementation, after obtaining the above-mentioned first state feature and second state feature, the first state feature and the second state feature can be input into a pre-trained instruction decision model to decide the corresponding control instruction. When training the decision model, at least one of the total duration of the real-time response voice on the machine side (currentTtsDuration), the playback duration of the real-time response voice (currentTtsPlayStartTime), the real-time total duration of the first voice stream (currentSayTime), the total silence duration on the user side (silenceTime), the total duration of the conversation (sessionTime), the conversation history context information (microTurnContext) and the user query text (Query) can be obtained as feature information to improve the accuracy of the decision model. The training process of the model will not be repeated here.

[0147] After step S4300, step S4400 is executed to control the machine end to perform the human-machine voice dialogue according to the control instruction after the first time slice.

[0148] In a specific implementation, controlling the machine end to perform the human-machine voice dialogue according to the control instruction includes: sending the control instruction to the machine end, so that the machine end performs the human-machine voice dialogue according to the control instruction.

[0149] Among them, the machine end conducts the human-computer voice dialogue according to the corresponding control instruction, including: the machine end obtains response information corresponding to the corresponding control instruction based on pre-stored mapping data, wherein the mapping data reflects the correspondence between each control instruction in the control instruction set and each set response information; and conducts human-computer voice dialogue according to the obtained response information.

[0150] Please refer to Table 3, which is a table showing the correspondence between control instructions and different response information:

[0151] Table 3:

[0152]

[0153] Please see Figure 6 , which is a schematic diagram of the architecture of human-computer voice dialogue provided by the embodiment of the present disclosure. Figure 6 As shown, for the first voice stream emitted by the user end and the second voice stream obtained by the monitoring machine end, trigger event detection processing is used to detect whether a specific trigger event occurs; if the occurrence of a trigger event is detected, the first state characteristics of the first voice stream and the second state characteristics of the second voice stream in the first time slice before the trigger event are obtained respectively; then the corresponding control instruction is obtained through the instruction decision model, and according to the instruction type of the control instruction, it is determined whether the control machine end stops the current broadcast, or obtains the response text through the question and answer processing module, and then the response text is processed by text-to-speech synthesis to control the machine end to continuously send out the response voice for the real-time voice in the first voice stream.

[0154] In summary, the method for controlling human-machine voice dialogue provided in this embodiment allows, during a human-machine voice dialogue, an electronic device, by receiving a first voice stream from a user end for a human-machine voice dialogue and monitoring a second voice stream from a machine end for the human-machine voice dialogue, to control the machine end to respond without waiting for the user to speak a round of voice before controlling the machine end to respond. Instead, during the process of receiving the first voice stream, the electronic device obtains a first state feature of the first voice stream at a first time slice and a second state feature of the second voice stream at the first time slice. Based on the first state feature and the second state feature, the electronic device selects a control instruction for controlling the machine end to conduct the human-machine voice dialogue after the first time slice. The electronic device then controls the machine end to respond promptly and accurately to the user end's output voice after the first time slice according to the control instruction. This method enables the human-machine voice dialogue to be conducted in voice duplex mode, allowing the machine end to respond promptly and accurately to the voice stream emitted by the user end at any time, reducing response delay and improving user experience.

[0155] Corresponding to the above embodiment, this embodiment also provides a control device for human-machine voice dialogue, such as Figure 7 As shown, it is a schematic principle block diagram of the control device for human-computer voice dialogue provided by an embodiment of the present disclosure.

[0156] according to Figure 7 As shown, the control device 7000 for human-machine voice dialogue in this embodiment includes a voice stream receiving module 7100 , a voice stream monitoring module 7200 , a decision module 7300 and an execution module 7400 .

[0157] The voice stream receiving module 7100 is used to receive a first voice stream of a human-computer voice dialogue performed by a user terminal.

[0158] The voice stream monitoring module 7200 is used to monitor the second voice stream of the human-machine voice dialogue on the machine side.

[0159] In one embodiment, when the voice stream monitoring module 7200 obtains the first state characteristics of the first voice stream in the first time slice and the second state characteristics of the second voice stream in the first time slice, it can be used to: detect the occurrence of a trigger event; and based on the detected trigger event, obtain the first state characteristics of the first voice stream in the first time slice before the trigger event is detected and the second state characteristics of the second voice stream in the first time slice.

[0160] In one embodiment, the triggering event includes an event in which a non-silent segment appears in the first voice stream. When detecting the event in which a non-silent segment appears in the first voice stream, the voice stream monitoring module 7200 can be used to: split the first voice stream to obtain adjacent first silent segments and second silent segments, wherein the first silent segment is earlier than the second silent segment; when the timing of the first silent segment and the second silent segment are not connected, extract the voice segment between the first silent segment and the second silent segment as a non-silent segment, and determine the event in which a non-silent segment appears in the first voice stream.

[0161] The state acquisition module 7300 is used to select control instructions corresponding to the first state characteristics and the second state characteristics from a set control instruction set; wherein the control instruction set includes instructions for controlling the machine end to broadcast and instructions for controlling the machine end to mute.

[0162] In one embodiment, when the state acquisition module 7300 selects the control instructions corresponding to the first state feature and the second state feature in the set control instruction set, it can be used to include: judging whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature, and obtaining a judgment result; based on the judgment result, selecting the control instructions corresponding to the first state feature and the second state feature in the control instruction set.

[0163] In one embodiment, the instruction for controlling the broadcast of the machine end includes a first control instruction to continue the current broadcast. The state acquisition module 7300 determines whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature. When the judgment result is obtained, it can be used to: when the second state feature is that the machine end remains non-silent or the machine end changes from silent to non-silent, determine that the machine end has the right to speak after the first time slice; when the state acquisition module 7300 selects the control instruction corresponding to the first state feature and the second state feature in the control instruction set, it can be used to: when the machine end has the right to speak, select the first control instruction as the corresponding control instruction.

[0164] In one embodiment, the instruction for controlling the machine end to mute includes a sixth control instruction to stop the current broadcast. The state acquisition module 7300 determines whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature. When obtaining the judgment result, it can be used to: when the first state feature indicates that a non-silent segment appears in the first voice stream, and the second state feature is that the machine end remains non-silent or the machine end changes from silent to non-silent, determine that the machine end does not have the right to speak after the first time slice; when the state acquisition module 7300 selects the control instruction corresponding to the first state feature and the second state feature in the control instruction set, it can be used to: when the machine end does not have the right to speak, select the sixth control instruction as the corresponding control instruction.

[0165] In one embodiment, the instructions broadcast by the control machine end include a fourth control instruction for broadcasting the set first round of question and answer content. The state acquisition module 7300 determines whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature. When the judgment result is obtained, it can be used to: when the first state feature and the second state feature are both in the dialogue starting state, determine that the machine end has the right to speak after the first time slice; when the state acquisition module 7300 selects the control instruction corresponding to the first state feature and the second state feature in the control instruction set based on the judgment result, it can be used to: when the machine end has the right to speak, select the fourth control instruction as the corresponding control instruction.

[0166] In one embodiment, the instruction for controlling the broadcast of the machine end includes a second control instruction for starting a new broadcast and / or a third control instruction for continuing the content in the sentence set for the broadcast. The state acquisition module 7300 determines whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature. When the judgment result is obtained, it can be used to: when the first state feature is that the user end changes from non-mute to mute and the second state feature is that the machine end remains silent, determine that the machine end has the right to speak after the first time slice; when the state acquisition module 7300 selects the control instruction corresponding to the first state feature and the second state feature in the control instruction set based on the judgment result, it can be used to: when the machine end has the right to speak, select the second control instruction or the third control instruction as the corresponding control instruction.

[0167] In one embodiment, the instruction for controlling the machine end to be muted includes a seventh control instruction for the machine end to remain muted. When the state acquisition module 7300 determines whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature, and obtains the judgment result, it can be used to: when the first state feature is that the user end changes from non-muted to muted and the second state feature is that the machine end remains muted, determine that the machine end does not have the right to speak after the first time slice; when the state acquisition module 7300 selects the control instruction corresponding to the first state feature and the second state feature in the control instruction set based on the judgment result, it can be used to: when the machine end does not have the right to speak, select the seventh control instruction as the corresponding control instruction.

[0168] In one embodiment, the instructions for controlling the broadcast of the machine end include a fifth control instruction for broadcasting the set mute prompt content. The state acquisition module 7300 determines whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature. When the judgment result is obtained, it can be used to: when the first state feature is that the user end remains silent and the second state feature is that the machine end changes from non-silent to silent, determine that the machine end has the right to speak after the first time slice; when the state acquisition module 7300 selects the control instruction corresponding to the first state feature and the second state feature in the control instruction set based on the judgment result, it can be used to: when the machine end has the right to speak, select the fifth control instruction as the corresponding control instruction.

[0169] In one embodiment, the instruction for controlling the machine end to be muted includes a seventh control instruction for the machine end to remain muted. When the state acquisition module 7300 determines whether the machine end has the right to speak after the first time slice based on the first state feature and the second state feature, and obtains the judgment result, it can be used to: when the first state feature is that the user end remains muted and the second state feature is that the machine end changes from non-muted to muted, determine that the machine end does not have the right to speak after the first time slice; when the state acquisition module 7300 selects the control instruction corresponding to the first state feature and the second state feature in the control instruction set based on the judgment result, it can be used to: when the machine end does not have the right to speak, select the seventh control instruction as the corresponding control instruction.

[0170] The execution module 7400 is used to control the machine end to perform the human-machine voice dialogue according to the control instruction after the first time slice.

[0171] In one embodiment, when the execution module 7400 controls the machine end to perform the human-machine voice dialogue according to the control instruction, it can be used to: send the control instruction to the machine end so that the machine end performs the human-machine voice dialogue according to the control instruction.

[0172] In this embodiment, when the human-computer voice dialogue is conducted according to the corresponding control instruction, the execution module 7400 can be used to: the machine end obtains response information corresponding to the corresponding control instruction based on pre-stored mapping data, wherein the mapping data reflects the correspondence between each control instruction in the control instruction set and each set response information; and conducts human-computer voice dialogue based on the obtained response information.

[0173] Corresponding to the above embodiment, this embodiment provides an electronic device, such as Figure 8a As shown, the electronic device 100 includes a control device 7000 for human-computer voice dialogue according to any embodiment of the present disclosure.

[0174] In another embodiment, Figure 8b As shown, the electronic device 100 may include a memory 110 and a processor 120, wherein the memory 110 is used to store executable instructions; and the processor 120 is used to execute a method as any method embodiment of the present disclosure under the control of the executable instructions.

[0175] Corresponding to the above embodiments, in this embodiment, a computer-readable storage medium is also provided, which stores a computer program that can be read and executed by a computer, and the computer program is used to execute the method described in any of the above embodiments of the present disclosure when read and executed by the computer.

[0176] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0177] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0178] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0179] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0180] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0181] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0182] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0183] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of an instruction, and the module, program segment or part of the instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are all equivalent.

[0184] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, their practical applications, or technical improvements in the marketplace, or to enable other persons skilled in the art to understand the embodiments disclosed herein. The scope of the present disclosure is defined by the appended claims.

Claims

1. A method for controlling human-computer voice dialogue, comprising: Receive a first voice stream of a human-machine voice dialogue at a user end and a second voice stream of the human-machine voice dialogue at a monitoring machine end; Obtaining a first state feature of the first voice stream at a first time slice and a second state feature of the second voice stream at the first time slice, wherein the first time slice includes a slice with occurrence times of two adjacent trigger events as start and end times, and the trigger event includes an event for triggering the machine end to respond; Selecting a corresponding control instruction from a set of control instructions based on the first state characteristic and the second state characteristic; wherein the control instruction set includes an instruction for controlling the machine end to broadcast and an instruction for controlling the machine end to mute, and the instruction for controlling the machine end to broadcast includes an instruction for interrupting the user end's speech and selecting to directly broadcast a response voice; After the first time slice, the machine end is controlled according to the control instruction to perform the human-machine voice dialogue.

2. The method according to claim 1, wherein The instructions for controlling the machine-side broadcast include at least one of a first control instruction for continuing the current broadcast, a second control instruction for starting a new broadcast, a third control instruction for broadcasting the set sentence continuation content, a fourth control instruction for broadcasting the set first round of question and answer content, and a fifth control instruction for broadcasting the set mute prompt content; and / or, The instruction for controlling the machine end to be muted includes at least one of a sixth control instruction for stopping the current broadcast and a seventh control instruction for keeping the machine end muted.

3. The method according to claim 1, wherein The acquiring a first state feature of the first voice stream at a first time slice and a second state feature of the second voice stream at the first time slice includes: Detecting the occurrence of a triggering event; According to the detected trigger event, a first state feature of the first voice stream in a first time slice before the trigger event is detected and a second state feature of the second voice stream in the first time slice are acquired.

4. The method according to claim 3, wherein: The trigger event includes at least one of the event of starting the human-computer voice dialogue, the event of a non-silent segment appearing in the first voice stream, the event of a silent segment appearing in the first voice stream, the event of a non-silent segment appearing in the second voice stream, the event of a silent segment appearing in the second voice stream, and the arrival of a set trigger time.

5. The method according to claim 3, wherein The triggering event includes an event in which a non-silent segment appears in the first voice stream, and the step of detecting the event in which a non-silent segment appears in the first voice stream includes: Splitting the first voice stream to obtain a first silent segment and a second silent segment that are adjacent to each other, wherein the first silent segment is earlier than the second silent segment; When the timing of the first silent segment and the second silent segment are not connected, the voice segment between the first silent segment and the second silent segment is extracted as a non-silent segment, and an event in which a non-silent segment occurs in the first voice stream is determined.

6. The method according to claim 1, wherein The selecting, from a set control instruction set, a control instruction corresponding to the first state characteristic and the second state characteristic comprises: Determining, based on the first state feature and the second state feature, whether the machine terminal has a speaking right after the first time slice, and obtaining a determination result; According to the judgment result, a control instruction corresponding to the first state feature and the second state feature is selected from the control instruction set.

7. The method according to claim 6, wherein: The determining, based on the first state feature and the second state feature, whether the machine terminal has a speaking right after the first time slice, and obtaining a determination result includes: If the second state characteristic is that the machine end remains non-muted or the machine end changes from being muted to being non-muted, determining that the machine end has the speaking right after the first time slice; The instruction for controlling the broadcast of the machine end includes a first control instruction for continuing the current broadcast, and selecting, from the control instruction set according to the judgment result, a control instruction corresponding to the first state feature and the second state feature, includes: When the machine end has the right to speak, the first control instruction is selected as the corresponding control instruction.

8. The method according to claim 6, wherein: The determining, based on the first state feature and the second state feature, whether the machine terminal has a speaking right after the first time slice, and obtaining a determination result includes: If the first state feature indicates that a non-silent segment appears in the first voice stream, and the second state feature indicates that the machine terminal remains non-silent or changes from being silent to being non-silent, determining that the machine terminal does not have a speaking right after the first time slice; The instruction for controlling the machine end to mute includes a sixth control instruction for stopping the current broadcast, and selecting, based on the judgment result, the control instruction corresponding to the first state feature and the second state feature in the control instruction set includes: When the machine end does not have the right to speak, the sixth control instruction is selected as the corresponding control instruction.

9. The method according to claim 8, wherein The first state feature indicates that a non-silent segment appears in the first voice stream, including: the user terminal changes from silent to non-silent and / or the user terminal changes from non-silent to silent.

10. The method according to claim 6, wherein: The determining, based on the first state feature and the second state feature, whether the machine terminal has a speaking right after the first time slice, and obtaining a determination result includes: When both the first state feature and the second state feature are in a conversation start state, determining that the machine end has the speaking right after the first time slice; The control machine end broadcasting instruction includes a fourth control instruction for broadcasting the set first round of question and answer content, and the control instruction corresponding to the first state feature and the second state feature in the control instruction set according to the judgment result includes: When the machine end has the right to speak, the fourth control instruction is selected as the corresponding control instruction.

11. The method according to claim 6, wherein: The determining, based on the first state feature and the second state feature, whether the machine terminal has a speaking right after the first time slice, and obtaining a determination result includes: When the first state characteristic is that the user terminal changes from being non-mute to being mute and the second state characteristic is that the machine terminal remains mute, determining that the machine terminal has the speaking right after the first time slice; The instructions for controlling the machine-side broadcast include a second control instruction for starting a new broadcast and / or a third control instruction for broadcasting the continuation content of a set sentence, and selecting, from the control instruction set based on the judgment result, the control instructions corresponding to the first state feature and the second state feature, includes: When the machine end has the right to speak, the second control instruction or the third control instruction is selected as the corresponding control instruction.

12. The method according to claim 6, wherein: The determining, based on the first state feature and the second state feature, whether the machine terminal has a speaking right after the first time slice, and obtaining a determination result includes: When the first state characteristic is that the user terminal changes from being non-mute to being mute and the second state characteristic is that the machine terminal remains mute, determining that the machine terminal does not have a speaking right after the first time slice; The instruction for controlling the device to be muted includes a seventh control instruction for keeping the device muted. Selecting, from the control instruction set, the control instructions corresponding to the first state characteristic and the second state characteristic based on the judgment result includes: When the machine end does not have the right to speak, the seventh control instruction is selected as the corresponding control instruction.

13. The method according to claim 6, wherein: The determining, based on the first state feature and the second state feature, whether the machine terminal has a speaking right after the first time slice, and obtaining a determination result includes: When the first state characteristic is that the user terminal remains silent and the second state characteristic is that the machine terminal changes from being non-silent to being silent, determining that the machine terminal has the speaking right after the first time slice; The instruction for controlling the machine end to broadcast includes a fifth control instruction for broadcasting a set mute prompt content, and the control instruction corresponding to the first state feature and the second state feature in the control instruction set is selected according to the judgment result, including: When the machine end has the right to speak, the fifth control instruction is selected as the corresponding control instruction.

14. The method according to claim 6, wherein The determining, based on the first state feature and the second state feature, whether the machine terminal has a speaking right after the first time slice, and obtaining a determination result includes: When the first state characteristic is that the user terminal remains silent and the second state characteristic is that the machine terminal changes from being non-silent to being silent, determining that the machine terminal does not have a speaking right after the first time slice; The instruction for controlling the device to be muted includes a seventh control instruction for keeping the device muted. Selecting, from the control instruction set, the control instructions corresponding to the first state characteristic and the second state characteristic based on the judgment result includes: When the machine end does not have the right to speak, the seventh control instruction is selected as the corresponding control instruction.

15. The method according to claim 1, wherein The controlling the machine end to perform the human-machine voice dialogue according to the control instruction includes: The control instruction is sent to the machine end, so that the machine end performs the human-machine voice dialogue according to the control instruction.

16. The method according to claim 15, wherein The machine side performs the human-machine voice dialogue according to the corresponding control instruction, including: The machine end obtains response information corresponding to the corresponding control instruction according to pre-stored mapping data, wherein the mapping data reflects the corresponding relationship between each control instruction in the control instruction set and each set response information; Conduct human-computer voice dialogue based on the obtained response information.

17. A control device for human-computer voice dialogue, comprising: A voice stream receiving module, configured to receive a first voice stream of a user terminal for human-computer voice dialogue; A voice stream monitoring module, configured to monitor a second voice stream of the human-machine voice dialogue on the machine side; a state acquisition module, configured to acquire a first state feature of the first voice stream at a first time slice and a second state feature of the second voice stream at the first time slice, wherein the first time slice includes a slice with the occurrence time of two adjacent trigger events as a start and end time, and the trigger event includes an event for triggering the control of the machine end to respond; a decision module, configured to select control instructions corresponding to the first state characteristic and the second state characteristic from a set control instruction set; wherein the control instruction set includes an instruction for controlling the machine end to broadcast and an instruction for controlling the machine end to mute, and the instruction for controlling the machine end to broadcast includes an instruction for interrupting the user end's speech and selecting to directly broadcast a response voice; and An execution module is used to control the machine end to perform the human-machine voice dialogue according to the control instruction after the first time slice.

18. An electronic device comprising the control device according to claim 17; or comprising: a memory for storing executable instructions; The processor is configured to operate the electronic device to execute the method according to any one of claims 1 to 16 under the control of the executable instructions.

19. A computer-readable storage medium, wherein: The computer-readable storage medium stores a computer program that can be read and executed by a computer, and the computer program is used to execute the method according to any one of claims 1 to 16 when read and executed by the computer.

Citation Information

Patent Citations

  • Voice interruption method and device

    CN111540349A