Voice control method, device, equipment, storage medium
By obtaining call hang-up or interrupt commands in the intelligent outbound call system and generating pause commands to control the processing of voice conversion and synthesis devices, the problem of low flexibility in voice control in the prior art is solved, and more efficient resource utilization and user experience improvement is achieved.
Patent Information
- Application Number
- CN202110963640.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-20
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-08-20
AI Technical Summary
The existing intelligent outgoing call system continues to process unfinished voice information when the user hangs up a session or interrupts a conversation, resulting in lower flexibility in voice control.
The control device obtains call hang-up or interrupt commands, generates pause commands, and sends pause commands to the voice conversion, answering and voice synthesis devices to pause processing tasks. Use websocket connection to transmit control commands and multimedia information to avoid HTTP protocol settings.
It improves the flexibility and user experience of voice control, saves voice system resources, and reduces the complexity of voice conversion and synthesis equipment.
Smart Images

Figure CN113611302B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent voice technology, and particularly to a voice control method, device, equipment, storage medium and program. Background Art
[0002] An intelligent outbound call system can automatically call a user's client device and conduct simple voice communication with the user through an intelligent robot.
[0003] Currently, when the intelligent outbound call system receives a user's voice message, it can convert the voice message into a text message, determine a corresponding response text according to the text message, and synthesize a response voice corresponding to the response text. The intelligent outbound call system sends the response voice to the user to communicate with the user. However, the scenarios of communication between the intelligent outbound call system and the user are variable (such as the user hanging up the conversation, interrupting the conversation, etc.). When the user hangs up the conversation or interrupts the conversation, the intelligent outbound call system still processes the unprocessed voice message and continues to generate the corresponding voice, resulting in a low flexibility of voice control of the intelligent outbound call system. Summary of the Invention
[0004] The main object of the present invention is to provide a voice control method, device, equipment, storage medium and program, aiming to solve the technical problem of low flexibility of voice control in the prior art.
[0005] In a first aspect, an embodiment of the present application provides a voice control method applied to a voice system. The voice system includes a control device, a call center device, a voice conversion device, a response device and a voice synthesis device, and there is a call connection between the call center device and the user device. The method includes:
[0006] The control device obtains a first instruction, where the first instruction is a call hang-up instruction or an interruption instruction for the call connection, and the interruption instruction is used to indicate that the user voice interrupts the voice playback of the call center device.
[0007] The control device generates a pause instruction according to the first instruction, and the pause instruction includes the call identifier of the voice call.
[0008] Send the pause instruction to the voice conversion device, the response device and the voice synthesis device, and the pause instruction is used to instruct the voice conversion device, the response device and the voice synthesis device to pause the processing tasks corresponding to the call connection.
[0009] In a possible implementation manner, the control device obtaining the first instruction includes:
[0010] The control device receives the user hang-up instruction sent by the call center device and determines the user hang-up instruction as the first instruction; or,
[0011] The control device receives the user voice sent by the call center device and determines the first instruction according to the user voice, where the first instruction is the interruption instruction or the system hang-up instruction; the call hang-up instruction includes the user hang-up instruction and the system hang-up instruction.
[0012] In a possible implementation manner, the first instruction is the interruption instruction; determining the first instruction according to the user voice information includes:
[0013] The control device sends the user voice to the voice conversion device;
[0014] The control device receives the user text corresponding to the user voice sent by the voice conversion device;
[0015] The control device obtains the number of characters included in the user text and the time information of the first response voice sent by the control device to the call center device last time, where the time information includes the first sending moment and the first duration of the first response voice;
[0016] Determine the interruption instruction according to the number of characters included in the user text and the time information.
[0017] In a possible implementation manner, determining the interruption instruction according to the number of characters included in the user text and the time information includes:
[0018] Determine whether the call center device is playing the first response voice according to the current moment, the first sending moment and the first duration of the first response voice;
[0019] If so, generate the interruption instruction when the number of characters included in the user text is greater than or equal to a preset threshold.
[0020] In a possible implementation manner, the first instruction is the system hang-up instruction; determining the first instruction according to the user voice information includes:
[0021] The control device sends the user voice to the voice conversion device;
[0022] The control device receives the user text corresponding to the user voice sent by the voice conversion device;
[0023] The control device sends the user text to the response device;
[0024] The control device receives the system hang-up instruction sent by the answering device;
[0025] The control device determines the system hang-up instruction as the first instruction.
[0026] In a possible implementation manner, the method further includes:
[0027] After the control device sends the second response voice to the call center device, the control device obtains the second sending moment and the second duration of the second response voice;
[0028] If the control device does not receive the user voice within a third duration after the second sending moment, the control device performs a preset operation, and the preset operation includes: the control device sends a preset voice to the call center device, and the control device sends a system hang-up instruction to the call center device, and the third duration is greater than or equal to the sum of the second duration and the preset duration.
[0029] In a possible implementation manner, the control device is respectively connected to the call center device, the voice conversion device, and the voice synthesis device through a websocket; the websocket connection is used to transmit control instructions and multimedia information, the control instructions include the first instruction and the pause instruction, and the multimedia information includes text information and voice information.
[0030] In a second aspect, an embodiment of the present application provides a voice system, including a control device, a call center device, a voice conversion device, an answering device, and a voice synthesis device, where,
[0031] The call center device is configured to send user voice or a user hang-up instruction to the control device through a websocket connection;
[0032] The control device is configured to send user voice to the voice conversion device through a websocket connection;
[0033] The voice conversion device is configured to convert the user voice into user text and send the user text to the control device through a websocket connection;
[0034] The control device is further configured to determine an interruption instruction according to the user text;
[0035] The control device is further configured to send the user text to the answering device;
[0036] The answering device is configured to determine a system hang-up instruction according to the user text and send the system hang-up instruction to the control device;
[0037] The control device is further configured to determine the interruption instruction, the system hang-up instruction, or the user hang-up instruction as a first instruction, and generate a pause instruction according to the first instruction;
[0038] The control device is further configured to send the pause instruction to the voice conversion device, the answering device, and the voice synthesis device.
[0039] In a possible implementation manner, the control device is further configured to execute the method described in the first aspect.
[0040] In a third aspect, an embodiment of the present application provides a voice control device, which is applied to a voice system. The voice system includes a control device, a call center device, a voice conversion device, an answering device, and a voice synthesis device. There is a call connection between the call center device and a user device; the voice control device includes a first acquisition module, a generation module, and a sending module, where:
[0041] The first acquisition module is configured to acquire a first instruction, where the first instruction is a call hang-up instruction or an interruption instruction for the call connection, and the interruption instruction is used to instruct a user voice to interrupt the voice playback of the call center device;
[0042] The generation module is configured to generate a pause instruction according to the first instruction, where the pause instruction includes a call identifier of the voice call;
[0043] The sending module is configured to send the pause instruction to the voice conversion device, the answering device, and the voice synthesis device, and the pause instruction is used to instruct the voice conversion device, the answering device, and the voice synthesis device to pause the processing tasks corresponding to the call connection.
[0044] In a possible implementation manner, the first acquisition module is specifically configured to:
[0045] The control device receives the user hang-up instruction sent by the call center device, and determines the user hang-up instruction as the first instruction; or,
[0046] The control device receives the user voice sent by the call center device, and determines the first instruction according to the user voice, where the first instruction is the interruption instruction or the system hang-up instruction; the call hang-up instruction includes the user hang-up instruction and the system hang-up instruction.
[0047] In a possible implementation manner, the first acquisition module is specifically configured to:
[0048] The control device sends the user voice to the voice conversion device;
[0049] The control device receives the user text corresponding to the user voice sent by the voice conversion device;
[0050] The control device obtains the number of characters included in the user text, and the time information of the first response voice sent by the control device to the call center device last time. The time information includes the first sending moment and the first duration of the first response voice;
[0051] Determine the interruption instruction according to the number of characters included in the user text and the time information.
[0052] In a possible implementation manner, the first obtaining module is specifically configured to:
[0053] Determine whether the call center device is playing the first response voice according to the current moment, the first sending moment and the first duration of the first response voice;
[0054] If so, generate the interruption instruction when the number of characters included in the user text is greater than or equal to a preset threshold.
[0055] In a possible implementation manner, the first obtaining module is specifically configured to:
[0056] The control device sends the user voice to the voice conversion device;
[0057] The control device receives the user text corresponding to the user voice sent by the voice conversion device;
[0058] The control device sends the user text to the response device;
[0059] The control device receives the system hang-up instruction sent by the response device;
[0060] The control device determines the system hang-up instruction as the first instruction.
[0061] In another possible implementation manner, the voice control device further includes a second obtaining module, and the second obtaining module is used for:
[0062] After the control device sends a second response voice to the call center device, the control device obtains the second sending moment and the second duration of the second response voice;
[0063] If the control device does not receive the user voice within a third time period after the second sending moment, the control device performs a preset operation, where the preset operation includes: the control device sending a preset voice to the call center device, and the control device sending a system hang-up instruction to the call center device, and the third time period is greater than or equal to the sum of the second time period and a preset time period.
[0064] In a possible implementation manner, the control device is respectively connected to the call center device, the voice conversion device, and the voice synthesis device through websocket connections; the websocket connections are used to transmit control instructions and multimedia information, the control instructions include the first instruction and the pause instruction, and the multimedia information includes text information and voice information.
[0065] In a fourth aspect, an embodiment of the present application provides a voice control device, including a processor and a memory;
[0066] The memory stores computer execution instructions;
[0067] The processor executes the computer execution instructions stored in the memory, so that the processor executes the voice control method as described in the first aspect.
[0068] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, where computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the voice control method as described in the first aspect.
[0069] In a sixth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the voice control method as described in the first aspect.
[0070] An embodiment of the present invention provides a voice control method, device, equipment, storage medium and program, which are applied to a voice system. The voice system includes a control device, a call center device, a voice conversion device, a response device and a voice synthesis device. There is a call connection between the call center device and the user. The control device obtains a first instruction, where the first instruction is a call hang-up instruction or an interruption instruction for the call connection, and the interruption instruction is used to instruct the user's voice to interrupt the voice playback of the call center device. The control device generates a pause instruction according to the first instruction and sends the pause instruction to the voice conversion device, the response device and the voice synthesis device, so that the voice conversion device, the response device and the voice synthesis device pause the processing tasks corresponding to the call connection. According to the above method, when the control device determines that the user interrupts or hangs up the conversation, the control device can timely generate a pause instruction and control the voice conversion device, the response device and the voice synthesis device to pause the processing tasks corresponding to the call connection through the pause instruction, which can not only save the resources of the voice system, but also flexibly formulate pause strategies, thereby improving the flexibility of voice control. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0072] Figure 1 FIG. is a schematic structural diagram of a voice system provided by an embodiment of the present application;
[0073] Figure 2 FIG. is a schematic flowchart of a voice control method provided by an embodiment of the present application;
[0074] Figure 3 FIG. is a schematic diagram of the process of generating an interruption instruction provided by an embodiment of the present application;
[0075] Figure 4 FIG. is a schematic diagram of the process of generating a system hang-up instruction provided by an embodiment of the present application;
[0076] Figure 5A FIG. is the processing process of the voice conversion device receiving the pause instruction provided by an embodiment of the present application;
[0077] Figure 5B FIG. is the processing process of the voice synthesis device receiving the pause instruction provided by an embodiment of the present application;
[0078] Figure 6 FIG. is a control schematic diagram of the call center device corresponding to the interruption instruction provided by an embodiment of the present application;
[0079] Figure 7 Schematic flowchart of a method for controlling session timeout provided by an embodiment of the present application;
[0080] Figure 8 Schematic diagram of a control process for session timeout provided by an embodiment of the present application;
[0081] Figure 9 Schematic diagram of a process of a voice control method provided by an embodiment of the present application;
[0082] Figure 10 Schematic diagram of the structure of a voice control device provided by an embodiment of the present application;
[0083] Figure 11 Schematic diagram of the structure of another voice control device provided by an embodiment of the present application;
[0084] Figure 12 Schematic diagram of the hardware structure of a voice control device provided by the present application.
[0085] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0086] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0087] In the related art, when an intelligent outbound call system receives a user's voice, it can generate a response voice corresponding to the user's voice and play the response voice to the user to achieve voice communication between humans and machines. However, in the actual application process, the scenarios of communication between the intelligent outbound call system and the user are variable. For example, when the intelligent outbound call system plays a voice to the user's client through a call connection, the user may interrupt the played voice, or directly hang up the current call connection, and the intelligent outbound call system can only fixedly process the unprocessed voice and continue to generate the corresponding voice. For example, when the user interrupts the voice being played by the intelligent outbound call system, the intelligent outbound call system will continue to play the voice, which results in a low flexibility of voice control of the intelligent outbound call system.
[0088] To solve the technical problem of low flexibility in voice control in related technologies, an embodiment of the present application provides a voice control method, which is applied to a voice system. The voice system includes a control device, a call center device, a voice conversion device, a response device, and a voice synthesis device. There is a call connection between the call center device and the user device. The control device receives a user hang-up instruction sent by the call center device and determines the user hang-up instruction as the first instruction. Alternatively, the control device receives a user voice sent by the call center device, determines an interruption instruction or a system hang-up instruction based on the user voice, and determines the interruption instruction or the system hang-up instruction as the first instruction. The control device generates a pause instruction according to the first instruction and sends the pause instruction to the voice conversion device, the response device, and the voice synthesis device, so that the voice conversion device, the response device, and the voice synthesis device pause the processing tasks corresponding to the call connection. In this way, during the voice interaction between the voice system and the user, if the user interrupts the voice or hangs up the call connection, the voice system sends a pause instruction to the voice conversion device and the voice synthesis device through a websocket connection, and sends a pause instruction to the response device through the Hyper Text Transfer Protocol (HTTP). This can avoid the situation where the user interrupts the voice while the voice system is still playing the voice, improving the user experience. At the same time, since the control device sends a pause instruction to the voice conversion device and the voice synthesis device through a websocket connection, there is no need to set the HTTP protocol in the voice conversion device and the voice synthesis device, reducing the complexity of the voice conversion device and the voice synthesis device. Moreover, when the user interrupts the voice or hangs up the call connection, the voice system can promptly stop processing the voice corresponding to the call connection, which can not only save the resources of the voice system but also flexibly formulate a pause strategy, thereby improving the flexibility of voice control.
[0089] Next, in combination with Figure 1 , the structure of the voice system involved in the present application will be described.
[0090] Figure 1 FIG. is a schematic structural diagram of a voice system provided by an embodiment of the present application. Please refer to Figure 1 , which includes a voice system and a user device. Among them, the voice system includes a control device, a call center device, a voice conversion device, a response device, and a voice synthesis device. When the voice system and the user device are in a call connection, if the user interrupts the voice broadcast of the voice system through the user device, or the user hangs up the call connection through the user device, the control device can send a pause instruction to the voice conversion device, the response device, and the voice synthesis device, so that the voice conversion device, the response device, and the voice synthesis device pause the processing tasks corresponding to the call connection. This can save the resources of the voice system and improve the flexibility of voice control.
[0091] The technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be described in detail below with specific embodiments. The following several specific embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0092] Figure 2 It is a schematic flowchart of a voice control method provided by an embodiment of the present application. Please refer to Figure 2 and the method may include:
[0093] S201. The control device obtains a first instruction, and the first instruction is a call hang-up instruction or an interruption instruction for the call connection.
[0094] The execution subject of the embodiment of the present application may be a control device in the voice system, or a voice control device provided in the control device. The voice control device may be implemented by software or by a combination of software and hardware.
[0095] The voice system includes a control device, a call center device, a voice conversion device, a response device, and a voice synthesis device. Among them, there is a call connection between the call center device and the user device. The user device may be a device such as a telephone or a computer. For example, the call center device may establish a call connection with the user's telephone. The control device is respectively connected to the call center device, the voice conversion device, and the voice synthesis device through a websocket connection. The control device is connected to the response device through an HTTP connection.
[0096] The websocket connection is used to transmit control instructions and multimedia information. The control instructions include the first instruction and a pause instruction. Optionally, the pause instruction is used to instruct the voice conversion device, the response device, and the voice synthesis device to pause the processing tasks corresponding to the call connection. The multimedia information includes text information and voice information. For example, the multimedia information may be the user voice received by the call center device from the user device, and the multimedia information may also be the corresponding text information generated by the voice conversion device according to the user voice.
[0097] The call center device is configured to send the user voice or a hang-up instruction to the control device through the websocket connection. For example, the call center device may obtain the user voice sent by the user device and send the user voice to the control device through the websocket connection. For example, when the user actively hangs up the call connection, the call center device may generate a hang-up instruction and send the hang-up instruction to the control device through the websocket connection.
[0098] A voice conversion device can convert voice into text. For example, if the voice received by the voice conversion device is "How's the weather today", the voice conversion device can convert this voice into the content of this text. For example, the voice conversion device can be an ASR. When the ASR receives voice, it can convert the voice into text.
[0099] A voice synthesis device can convert text into voice. For example, the voice synthesis device can be a TTS. When the TTS receives the response text, the TTS can convert the response text into a response voice.
[0100] The call hang-up instruction includes a user hang-up instruction and a system hang-up instruction. Among them, the user hang-up instruction is used to indicate that the user hangs up the call connection between the call center device and the user device. For example, when the user communicates with the voice system by phone, if the user hangs up the phone, the voice system generates a hang-up instruction. The system hang-up instruction is that the voice system actively hangs up the call connection between the call center device and the user device. For example, when the user communicates with the voice system by phone, if the voice system generates a system hang-up instruction, the voice system actively hangs up the call connection with the user device.
[0101] The interruption instruction is used to indicate that the user's voice interrupts the voice playback of the call center device. For example, when the call center device plays voice to the user device, if the user emits a voice to interrupt the voice playback of the call center device, the voice system generates an interruption instruction.
[0102] The first instruction can be obtained according to the following feasible implementation method: When the user actively hangs up the connection between the user device and the voice system, the call center device can generate a user hang-up instruction and send the user hang-up instruction to the control device through a websocket connection. After receiving the user hang-up instruction, the control device determines the user hang-up instruction as the first instruction. For example, when the user actively hangs up the call connection between the phone and the call center device, the first instruction obtained by the control device is the user hang-up instruction.
[0103] When the call center device sends the user's voice to the control device, the control device receives the user's voice sent by the call center device and determines the first instruction according to the user's voice. Among them, the first instruction is an interruption instruction or a system hang-up instruction.
[0104] Optionally, there are the following two situations for the control device to determine the first instruction according to the user's voice:
[0105] Situation 1: The first instruction is an interruption instruction.
[0106] When the first instruction is an interruption instruction, the control device can determine the first instruction according to the following feasible implementation methods: The control device sends the user voice to the voice conversion device. For example, the control device obtains the user voice from the call center device and sends the user voice to the voice conversion device through a websocket connection.
[0107] The control device receives the user text corresponding to the user voice sent by the voice conversion device. Optionally, the voice conversion device is used to convert the user voice into user text and send the user text to the control device through a websocket connection. For example, the voice conversion device receives the user voice sent by the control device, converts the user voice into the corresponding user text, and sends the user text to the voice conversion device through a websocket connection.
[0108] The control device obtains the number of characters included in the user text and the time information of the first response voice sent by the control device to the call center device last time. The number of characters is the number of characters in the user text. For example, if the user text includes 10 characters, the number of characters is 10. The first response voice is the response voice sent by the control device to the call center device last time. For example, during the process of the user communicating with the voice system, the control device can send the response voice corresponding to the user voice to the call center device, and the first response voice can be the last response voice sent at the current moment. The time information includes the first sending moment and the first duration of the first response voice. The first sending moment is the moment when the first response voice is sent. The first duration is the playing duration of the first response voice. For example, if the playing duration of the first response voice is 10 seconds, the first duration of the first response voice is 10 seconds. Optionally, the control device can determine the first duration of the first response voice according to the voice segments of the first response voice. For example, if the control device receives 10 voice segments corresponding to the first response voice, and each voice segment is 20 milliseconds, the first duration of the first response voice is 200 milliseconds.
[0109] Determine the interruption instruction according to the number of characters included in the user text and the time information. Optionally, it is possible to determine whether the call center device is playing the first response voice according to the current moment, the first sending moment, and the first duration of the first response voice. For example, determine the time difference according to the current moment and the first sending moment, and then determine whether the call center device is playing the first response voice according to the first duration and the time difference. For example, when the time difference between the current moment and the first sending moment is 10 seconds, if the first duration is 5 seconds, it is determined that the call center device has completed the playback of the first response voice; if the first duration is 15 seconds, it is determined that the call center device has not completed the playback of the first response voice, and the call center device is playing the first response voice.
[0110] If the call center device is playing the first response voice, when the number of characters included in the user text is greater than or equal to a preset threshold, an interruption instruction is generated. For example, when the call center device is playing the first response voice, if the number of characters in the user text obtained by the control device from the voice conversion device is greater than or equal to the preset threshold, it indicates that the user is sending a user voice through the user device. At this time, the control device can determine that the user interrupts the voice played by the call center device, the control device generates an interruption instruction, and determines the interruption instruction as the first instruction; if the number of characters in the user text obtained by the control device from the voice conversion device is less than the preset threshold, it indicates that the user voice sent by the user device is invalid voice (such as ambient noise, etc.). At this time, the control device determines that the user does not interrupt the voice played by the call center device, and the control device does not generate an interruption instruction.
[0111] Next, in combination with Figure 3 , the interruption instruction generated in this case will be described.
[0112] Figure 3 It is a schematic diagram of a process for generating an interruption instruction provided by an embodiment of the present application. Please refer to Figure 3 , including: a control device, a call center device, and a voice conversion device. Among them, the control device sends 10 seconds of voice to the call center device, and the call center device starts playing the voice after receiving the 10 seconds of voice. After 5 seconds, the voice conversion device sends the user text to the control device. The control device receives the user text and determines the number of characters in the user text. When the number of characters in the user text is greater than the preset threshold, the control device generates an interruption instruction and determines the interruption instruction as the first instruction.
[0113] In this case, when the first instruction is an interruption instruction, the control device can accurately determine whether the user interrupts the played voice based on the number of characters in the user text and the time information of the first response voice sent to the call center device last time, and then accurately determine whether the user generates an interruption instruction, improving the accuracy and flexibility of voice control.
[0114] Case 2: The first instruction is a system hang-up instruction.
[0115] When the first instruction is a system hang-up instruction, the control device can determine the first instruction according to the following feasible implementation method: The control device sends the user voice to the voice conversion device and receives the user text corresponding to the user voice sent by the voice conversion device. For example, the control device sends the user voice to the voice conversion device through a websocket connection. After receiving the user voice, the voice conversion device generates the user text corresponding to the user voice and sends the user text corresponding to the user voice to the control device through the websocket connection.
[0116] The control device sends user text to the answering device and receives the system hang-up instruction sent by the answering device. Among them, the answering device is used to determine the answering text corresponding to the user text, and determine whether the current call connection ends according to the answering text. If so, it generates a system hang-up instruction and sends the system hang-up instruction to the control device. Optionally, the answering device may include a multi-round session management system DM and a natural language understanding system NLU. For example, the answering device can obtain multiple answering texts corresponding to the user text through the multi-round session management system, and the natural language understanding system can determine the answering text corresponding to the user text from the multiple answering texts.
[0117] Optionally, the control device and the answering device are connected through HTTP, and the user text is sent to the answering device through the HTTP connection. Optionally, the answering device determines whether the current call connection ends according to the user text. For example, when the answering device receives the user text, the answering device simulates the answering results of various scenarios of the user text through the multi-round session management system DM, and then obtains the answering texts of the user text in various scenarios, and determines a unique answering text corresponding to the user text from the answering texts in various scenarios through the natural language understanding system NLU, and then determines whether the current call connection ends according to the answering text. For example, if the user text received by the answering device is "I'm very busy and have no time", the answering device determines that the answering text is "Okay, goodbye" and determines that the current call connection ends.
[0118] When the answering device determines that the current call connection ends, the answering device can send the system hang-up instruction to the control device through the HTTP connection. For example, when the answering device determines that the current call connection ends, the answering device sends a Hangup signaling to the control device through HTTP. When the control device receives the system hang-up instruction, it determines the system hang-up instruction as the first instruction.
[0119] Next, in combination with Figure 4 , the process of generating the system hang-up instruction will be described.
[0120] Figure 4 FIG. is a schematic diagram of a process for generating a system hang-up instruction provided by an embodiment of the present application. Please refer to Figure 4, including a voice conversion device, a control device, and a response device. The control device sends the voice "I'm not convenient to answer the phone right now" to the voice conversion device. The voice conversion device converts the voice into the corresponding text and sends the text "I'm not convenient to answer the phone right now" to the control device. After receiving the text, the control device sends the text "I'm not convenient to answer the phone right now" to the response device. The response device determines that the response text corresponding to the voice is "Okay, goodbye". When the response device determines that the response text is "Okay, goodbye", the response device determines that the current call connection ends and sends a system interruption instruction to the control device. Optionally, the "Okay, goodbye" text determined by the response device may not be sent to the control device, but directly send a system interruption instruction, or send the "Okay, goodbye" text to the control device, and then generate a response voice of "Okay, goodbye" through the control device and the voice synthesis device. This application embodiment does not make a limitation on this.
[0121] In this case, when the first instruction is a system hang-up instruction, the control device can send the user text corresponding to the user voice to the response device. If the response device determines that the response text corresponding to the user text indicates the end of the current call connection, the response device can send a system hang-up instruction to the control device. After receiving the system hang-up instruction, the control device can actively hang up the call connection between the voice system and the user device. This can not only improve the diversity of voice control customization, but also improve the flexibility of voice control.
[0122] S202. The control device generates a pause instruction according to the first instruction.
[0123] Optionally, when the call center device calls the user device, the call center device can establish a session room for this call, and the voice information of this call is saved in the session room. Optionally, the call center device may include an outbound management system and a call center middleware. The outbound management system can obtain the accounts of the user devices and establish a session room for each account respectively. The call center middleware can call the accounts of the user devices.
[0124] Optionally, connections need to be established between the control device and the call center device, the voice conversion device, the response device, and the voice synthesis device for each session room. For example, if the voice system calls 2 user devices simultaneously, the call center device establishes 2 session rooms respectively. In each session room, the control device is connected to the call center device, the voice conversion device, and the voice synthesis device through websocket respectively, and is connected to the response device through HTTP.
[0125] The pause instruction includes the call identifier of the voice call. Among them, the call identifier can be the identifier of the session room. For example, the voice system can simultaneously call multiple user devices and establish session rooms for each user device respectively. When generating a pause instruction, the control device can add the identifier of the session room to the pause instruction, so as to accurately control each session room and improve the accuracy of voice control.
[0126] Optionally, the specific process for the control device to generate a pause instruction according to the first instruction is as follows: If the control device receives the first instruction, the control device generates a pause instruction. For example, if the control device receives a user hang-up instruction sent by the call center device, the control device generates a pause instruction, and the pause instruction includes the call identifier of this voice call; if the control device determines the first instruction as an interruption instruction or a system hang-up instruction according to the user voice, the control device generates a pause instruction.
[0127] Optionally, the pause instruction can be a custom instruction. For example, if the user pre-customizes the pause instruction as an EOS signaling, when the control device receives the first instruction, the control device generates an EOS signaling.
[0128] S203. Send the pause instruction to the voice conversion device, the answering device, and the voice synthesis device.
[0129] The pause instruction is used to instruct the voice conversion device, the answering device, and the voice synthesis device to pause the processing tasks corresponding to the call connection. Among them, the processing task of the voice conversion device is the user voice that has not been processed in this call connection, the processing task of the answering device is the user text that has not been processed in this call connection, and the processing task of the voice synthesis device is the answering text that has not been processed in this call connection.
[0130] Optionally, the control device is further configured to send the pause instruction to the voice conversion device and the voice synthesis device through a websocket connection, and send the pause instruction to the answering device through an HTTP connection. When the voice conversion device, the voice synthesis device, and the answering device receive the pause instruction, the voice conversion device, the voice synthesis device, and the answering device pause the processing tasks of this session connection.
[0131] Optionally, when the control device sends the pause instruction to the answering device through an HTTP connection, the answering device stops the processing task of this call connection.
[0132] Next, in combination with Figure 5A - Figure 5B , the process of the control device sending the pause instruction to the voice conversion device and the voice synthesis device will be described in detail.
[0133] Figure 5A This is the processing process of the voice conversion device provided by the embodiment of the present application when receiving the pause instruction. Please refer to Figure 5A, including a control device and a voice conversion device. The control device and the voice conversion device are connected through a websocket. The control device sends a pause instruction to the voice conversion device through the websocket connection. Optionally, the pause instruction can be a signaling defined by the user. For example, the pause instruction can be an EOS signaling. When the voice conversion device receives the pause instruction, it stops processing the remaining user voice of the current call connection. After the voice conversion device finishes processing the user voice that is currently being parsed, it sends a processing end instruction to the control device through the websocket connection. Among them, the processing end instruction can be an instruction defined by the user. For example, the processing end instruction can be a signaling with ended=1. When the control device receives the processing end instruction, the control device disconnects the websocket connection with the voice conversion device.
[0134] Figure 5B This is the processing procedure when the voice synthesis device provided in the embodiment of the present application receives a pause instruction. Please refer to Figure 5B , including a control device and a voice synthesis device. The control device and the voice synthesis device are connected through a websocket. The control device sends a pause instruction to the voice synthesis device through the websocket connection. When the voice synthesis device receives the pause instruction, it stops processing the remaining response text of the current call connection. After the voice synthesis device finishes processing the response text that is currently being parsed, it sends a processing end instruction to the control device through the websocket connection. When the control device receives the processing end instruction, the control device disconnects the websocket connection with the voice synthesis device. In this way, when the user hangs up or the system hangs up the call connection, the control device can control the voice synthesis device, the voice conversion device, and the response device to stop the processing tasks of the current call connection, thereby saving the resources of the voice system. At the same time, the control device can send multimedia information and control instructions to the voice conversion device and the voice synthesis device through the websocket connection. Only the websocket protocol needs to be set in the voice conversion device and the voice synthesis device, and there is no need to set the HTTP protocol for signaling transmission, thereby reducing the complexity of the voice system. The control device can accurately identify the scenarios of voice communication between the voice system and the user device, and thus accurately formulate control strategies, improving the accuracy and flexibility of voice control.
[0135] Optionally, if the first instruction obtained by the control device is a system hang-up instruction, the control device can also send a pause instruction to the call center device through the websocket connection. After the call center device receives the pause instruction, it hangs up through the BYE signaling of the SIP protocol. When the call connection between the call center device and the user device is interrupted, the control device disconnects the websocket connection with the call center device.
[0136] Optionally, when the first instruction is an interruption instruction, the control device can also send a pause instruction to the call center device through a websocket connection.
[0137] Next, in combination with Figure 6 , the process of the control device sending a pause instruction to the call center device when the first instruction is an interruption instruction will be described.
[0138] Figure 6 It is a control schematic diagram of the call center device corresponding to the interruption instruction provided by the embodiment of the present application. Please refer to Figure 6 , including a control device and a call center device. Among them, the control device and the call center device are connected through a websocket connection. When the first instruction obtained by the control device is an interruption instruction, the control device sends a pause instruction to the call center device. For example, the pause instruction can be a CLEAR signaling defined by the user. When the call center device receives the pause instruction, the call center device can clear the voice generated by the voice conversion device in the buffer. In this way, when the user interrupts the voice played by the call center device, the call center device can timely clear the voice that has not been played yet, avoiding the overlap of the user's voice and the voice system, and improving the flexibility of voice control.
[0139] An embodiment of the present application provides a voice control method, which is applied to a voice system. The voice system includes a control device, a call center device, a voice conversion device, a response device, and a voice synthesis device. There is a call connection between the call center device and the user device. The control device receives a user hang-up instruction sent by the call center device and determines the user hang-up instruction as the first instruction. Alternatively, the control device receives a user voice sent by the call center device and determines the first instruction according to the user voice. The control device generates a pause instruction according to the first instruction and sends the pause instruction to the voice conversion device, the response device, and the voice synthesis device, so that the voice conversion device, the voice synthesis device, and the response device pause the processing tasks corresponding to the call connection. In this way, during the voice interaction between the voice system and the user, if the user interrupts the voice or hangs up the call connection, the control device can send a pause instruction to the voice conversion device and the voice synthesis device through a websocket connection, and send a pause instruction to the response device through an HTTP connection. This can avoid the situation where the user interrupts the voice while the voice system is still playing the voice, improving the user experience. At the same time, since the control device sends multimedia information and pause instructions to the voice conversion device and the voice synthesis device through a websocket connection, there is no need to set the HTTP protocol in the voice conversion device and the voice synthesis device, reducing the complexity of the voice conversion device and the voice synthesis device. Moreover, when the user interrupts the voice or hangs up the call connection, the voice system can stop processing the voice corresponding to the call connection in time, which can not only save the resources of the voice system, but also flexibly formulate a pause strategy, thereby improving the flexibility of voice control.
[0140] The voice control method provided by the embodiment of the present application further includes a control process for session timeout. Based on the embodiment shown in Figure 2 below, the control process for session timeout will be described in conjunction with Figure 7 .
[0141] Figure 7 It is a schematic flowchart of a control method for session timeout provided by an embodiment of the present application. Please refer to Figure 7 . The method includes:
[0142] S701. After the control device sends a second response voice to the call center device, obtain the second sending time and the second duration of the second response voice.
[0143] Session timeout means that no new user voice is sent within a period of time after the user device obtains the response voice. For example, if the call center device does not receive a new user voice 10 seconds after playing the response voice, it is determined that the session timeout of the current call connection occurs.
[0144] Optionally, the second response voice can be any response voice sent by the control device to the call center device. The second sending time is the time when the control device sends the second response voice to the call center device. The second duration is the duration of the second response voice. Optionally, when the voice synthesis device sends the second response voice to the control device, the voice synthesis device can send the second response voice to the control device in fragments, and the control device determines the second duration of the second response voice according to the received voice fragments. For example, the voice conversion device sends the response voice to the control device at a sampling rate of 8 kHz, with 320 bytes per fragment. The total number of bytes T of the response voice is added to the first 4-byte fragment. The first fragment is the One-Piece signaling. When the control device receives the first fragment, it can obtain the total number of bytes and determine the total number of fragments of the response voice, and then determine the playing duration of the response voice through the duration of each fragment.
[0145] For example, if the total number of fragments of the response voice sent by the voice synthesis device to the control device is 100 and each fragment is 20 milliseconds, the playing duration corresponding to the response voice is 2000 milliseconds.
[0146] S702: If the control device does not receive the user voice within the third duration after the second sending time, the control device performs a preset operation.
[0147] The third duration is greater than or equal to the sum of the second duration and the preset duration. The preset duration can be the duration preset for session timeout determination. For example, if the session times out when the user has not sent a voice after hearing the response voice for 10 seconds, the preset duration is 10 seconds.
[0148] Optionally, the preset operations include: the control device sends a preset voice to the call center device, and the control device sends a system hang-up instruction to the call center device. The preset voice can be a preset script voice. For example, the preset voice can be script voices such as "What do you think?" and "How do you feel?". For example, if the user device does not send a new user voice after receiving the response voice for the preset duration, the control device sends the voice "How do you feel?" to the call center device, and the call center device plays the voice to the user device. For example, if the user device does not send a new user voice after receiving the response voice for the preset duration, the control device sends a system hang-up instruction to the call center device, and the call center device actively hangs up the call connection through the BYE signaling in the SIP protocol.
[0149] Next, in combination with Figure 8 , the control process for session timeout will be described.
[0150] Figure 8 FIG. [Reference number to be filled] is a schematic diagram of a control process for session timeout provided by an embodiment of the present application. Please refer to Figure 8, including: a voice synthesis device, a control device, a call center device, and a user device. Among them, the control device is respectively connected to the voice synthesis device and the call center device through websocket, and the call center device is connected to the user device for a call.
[0151] Please refer to Figure 8 , the voice synthesis device sends a response voice clip to the control device. The control device determines the playback duration of the response voice according to the response voice clip. After the third duration, if the control device does not receive the user voice sent by the call center device, the control device sends a system hang-up instruction to the call center device. When the call center device receives the system hang-up instruction, it sends a BYE signaling in the SIP protocol to the user device to hang up the call connection between the call center device and the user device.
[0152] The embodiment of the present application provides a method for controlling session timeout. After the control device sends the second response voice to the call center device, it obtains the second sending time and the second duration of the second response voice. If the control device does not receive the user voice within the third duration after the second sending time, the control device performs a preset operation, where the third duration is greater than or equal to the sum of the second duration and the preset duration. In this way, when the user does not reply to the response voice sent by the call center device for a long time, the control device can actively determine that the current session times out and send a system hang-up instruction to the call center device, so that the call center device actively hangs up the call connection with the user device, thereby saving the resources of the voice system and improving the flexibility of voice control.
[0153] Based on any of the above embodiments, below, in combination with Figure 9 , the process of the above voice control method is described by way of example.
[0154] Figure 9 is a schematic diagram of the process of a voice control method provided by an embodiment of the present application. In Figure 9 In the shown embodiment, the system hang-up instruction is the first instruction. Please refer to Figure 9 , including: a user device, a call center device, a control device, a voice conversion device, a response device, and a voice synthesis device. Among them, the control device is respectively connected to the call center device, the voice conversion device, and the voice synthesis device through websocket, and the control device is connected to the response device through HTTP, and the call center device is connected to the user device for a call.
[0155] Please refer to Figure 9, the user equipment sends the user voice to the call center equipment. The call center equipment sends the user voice to the control equipment through a websocket connection. The control equipment sends the user voice to the voice conversion equipment through a websocket connection. After receiving the user voice, the voice conversion equipment generates the user text corresponding to the user voice and sends the user text to the control equipment through a websocket connection. The control equipment sends the user text to the answering equipment through an HTTP connection. The answering equipment determines the end of this call according to the user text and determines that the system actively hangs up the call.
[0156] Please refer to Figure 9 , the answering equipment sends a system hang-up instruction to the control equipment through an HTTP connection. The control equipment generates a pause instruction and sends the pause instruction to the voice conversion equipment, the answering equipment, and the voice synthesis equipment, so that the voice conversion equipment, the answering equipment, and the voice synthesis equipment pause the processing tasks of this call connection. The control equipment disconnects the websocket connections with the voice synthesis equipment and the voice conversion equipment.
[0157] Please refer to Figure 9 , the control equipment can also send a pause instruction to the call center equipment. When the call center equipment receives the pause instruction, it sends a BYE signaling in the SIP protocol to the user equipment to disconnect the call connection with the user equipment. When the call connection between the user equipment and the call center equipment is disconnected, the control equipment disconnects the websocket connection with the call center equipment. In this way, since the control equipment sends multimedia information and pause instructions to the voice conversion equipment and the voice synthesis equipment through a websocket connection, there is no need to set up an HTTP protocol in the voice conversion equipment and the voice synthesis equipment, reducing the complexity of the voice conversion equipment and the voice synthesis equipment. When the voice system actively hangs up the call connection, the control equipment can use the pause instruction to make the voice conversion equipment, the answering equipment, and the voice synthesis equipment pause the processing tasks of this call connection, saving the processing resources of the voice system and being able to flexibly formulate pause strategies, thereby improving the flexibility of voice control.
[0158] Figure 10 It is a schematic structural diagram of a voice control device provided by an embodiment of the present application. Please refer to Figure 10 , the voice control device 10 can be set in the control equipment. The voice control device 10 includes a first acquisition module 11, a generation module 12, and a sending module 13, where:
[0159] The first acquisition module 11 is used to acquire a first instruction, where the first instruction is a call hang-up instruction or an interruption instruction for the call connection, and the interruption instruction is used to indicate that the user voice interrupts the voice playback of the call center equipment;
[0160] The generating module 12 is configured to generate a pause instruction according to the first instruction, where the pause instruction includes the call identifier of the voice call;
[0161] The sending module 13 is configured to send the pause instruction to the voice conversion device, the answering device, and the voice synthesis device, where the pause instruction is used to instruct the voice conversion device, the answering device, and the voice synthesis device to pause the processing tasks corresponding to the call connection.
[0162] In a possible implementation manner, the first obtaining module 11 is specifically configured to:
[0163] The control device receives the user hang-up instruction sent by the call center device and determines the user hang-up instruction as the first instruction; or,
[0164] The control device receives the user voice sent by the call center device and determines the first instruction according to the user voice, where the first instruction is the interruption instruction or the system hang-up instruction; the call hang-up instruction includes the user hang-up instruction and the system hang-up instruction.
[0165] In a possible implementation manner, the first obtaining module 11 is specifically configured to:
[0166] The control device sends the user voice to the voice conversion device;
[0167] The control device receives the user text corresponding to the user voice sent by the voice conversion device;
[0168] The control device obtains the number of characters included in the user text, and the time information of the first response voice sent by the control device to the call center device last time, where the time information includes the first sending moment and the first duration of the first response voice;
[0169] Determine the interruption instruction according to the number of characters included in the user text and the time information.
[0170] In a possible implementation manner, the first obtaining module 11 is specifically configured to:
[0171] Determine whether the call center device is playing the first response voice according to the current moment, the first sending moment, and the first duration of the first response voice;
[0172] If so, generate the interruption instruction when the number of characters included in the user text is greater than or equal to a preset threshold.
[0173] In a possible implementation manner, the first obtaining module 11 is specifically configured to:
[0174] The control device sends the user voice to the voice conversion device;
[0175] The control device receives the user text corresponding to the user voice sent by the voice conversion device;
[0176] The control device sends the user text to the answering device;
[0177] The control device receives the system hang-up instruction sent by the answering device;
[0178] The control device determines the system hang-up instruction as the first instruction.
[0179] The voice control device provided by the embodiment of the present application can execute the technical solution shown in the above method embodiment, and its implementation principle and beneficial effects are similar, which will not be elaborated here.
[0180] The voice control device shown in the embodiment of the present application can be a chip, a hardware module, a processor, etc. Of course, the voice control device can be in other forms, and the embodiment of the present application does not make specific limitations on this.
[0181] Figure 11 For another structural schematic diagram of the voice control device provided by the embodiment of the present application. On the basis of the embodiment shown in Figure 10 please refer to Figure 11 , the voice control device 10 further includes a second acquisition module 14, and the second acquisition module 14 is used for:
[0182] After the control device sends the second response voice to the call center device, the control device acquires the second sending time and the second duration of the second response voice;
[0183] If the control device does not receive the user voice within the third duration after the second sending time, the control device performs a preset operation, and the preset operation includes: the control device sends a preset voice to the call center device, and the control device sends a system hang-up instruction to the call center device, and the third duration is greater than or equal to the sum of the second duration and the preset duration.
[0184] In a possible implementation manner, the control device is respectively connected to the call center device, the voice conversion device, and the voice synthesis device through a websocket connection; the websocket connection is used to transmit control instructions and multimedia information, the control instructions include the first instruction and the pause instruction, and the multimedia information includes text information and voice information.
[0185] The voice control device provided in the embodiment of the present application can execute the technical solution shown in the above method embodiment, and its implementation principle and beneficial effects are similar, which will not be repeated here.
[0186] The voice control device shown in the embodiment of the present application can be a chip, a hardware module, a processor, etc. Of course, the voice control device can be in other forms, and the embodiment of the present application does not specifically limit this.
[0187] Figure 12 The hardware structure diagram of the voice control device provided for this application. Figure 12 The voice control device 20 may include: a processor 21 and a memory 22, wherein the processor 21 and the memory 22 can communicate; exemplarily, the processor 21 and the memory 22 communicate via a communication bus 23, the memory 22 is used to store program instructions, and the processor 21 is used to call the program instructions in the memory to execute the voice control method shown in any of the above method embodiments.
[0188] Optionally, the voice control device 20 may further include a communication interface, which may include a transmitter and / or a receiver.
[0189] Optionally, the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. The steps of the method disclosed in the present application may be directly implemented as being executed by a hardware processor, or may be implemented by a combination of hardware and software modules in the processor.
[0190] The embodiment of the present application provides a voice system, including a control device, a call center device, a voice conversion device, an answering device and a voice synthesis device, wherein:
[0191] The call center device is used to send user voice or user hang-up instruction to the control device through a websocket connection;
[0192] The control device is used to send the user voice to the voice conversion device through a websocket connection;
[0193] The voice conversion device is used to convert the user voice into user text, and send the user text to the control device through a websocket connection;
[0194] The control device is further configured to determine an interruption instruction according to the user text;
[0195] The control device is further configured to send the user text to the answering device;
[0196] The answering device is configured to determine a system hang-up instruction according to the user text, and send the system hang-up instruction to the control device;
[0197] The control device is further configured to determine the interruption instruction, the system hang-up instruction, or the user hang-up instruction as a first instruction, and generate a pause instruction according to the first instruction;
[0198] The control device is further configured to send the pause instruction to the voice conversion device, the answering device, and the voice synthesis device.
[0199] The present application provides a readable storage medium, on which a computer program is stored; the computer program is used to implement the voice control method as described in any of the above embodiments.
[0200] The embodiment of the present application provides a computer program product, which includes instructions that, when executed, cause a computer to execute the above voice control method.
[0201] All or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a readable memory. When the program is executed, it executes the steps including the above method embodiments; and the foregoing memory (storage medium) includes: read-only memory (abbreviation: ROM), RAM, flash memory, hard disk, solid state drive, magnetic tape, floppy disk, optical disc, and any combination thereof.
[0202] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, a special-purpose computer, an embedded processing machine, or other programmable terminal devices to generate a machine, so that the instructions executed by the processing unit of the computer or other programmable terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0203] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the function specified in one process Figure 1 or more processes and / or blocks Figure 1 or more blocks.
[0204] These computer program instructions can also be loaded onto a computer or other programmable terminal device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the function specified in one process Figure 1 or more processes and / or blocks Figure 1 or more blocks.
[0205] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
[0206] In the present application, the term "including" and its variations may refer to non-limiting inclusion; the term "or" and its variations may refer to "and / or". In the present application, terms such as "first", "second", etc. are used to distinguish similar objects and do not necessarily have to describe a specific order or sequence. In the present application, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0207] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A voice control method, characterized in that, Applied to a voice system, the voice system includes a control device, a call center device, a voice conversion device, a response device, and a voice synthesis device. There is a call connection between the call center device and the user device. The method includes: The control device receives the user voice or the user hang-up instruction sent by the call center device through a websocket connection, and determines a first instruction. The first instruction is a call hang-up instruction or an interruption instruction for the call connection. The interruption instruction is used to instruct the user voice to interrupt the voice playback of the call center device. The call hang-up instruction includes the user hang-up instruction and the system hang-up instruction. The control device generates a pause instruction according to the first instruction. The pause instruction includes the call identifier of the voice call. Send the pause instruction to the voice conversion device, the response device, and the voice synthesis device through the websocket connection. The pause instruction is used to instruct the voice conversion device, the response device, and the voice synthesis device to pause the processing tasks corresponding to the call connection and clear the voice generated by the voice conversion device in the buffer. After the control device sends the second response voice to the call center device, the control device obtains the second sending time and the second duration of the second response voice. If the control device does not receive the user voice within a third duration after the second sending time, the control device performs a preset operation. The preset operation includes: the control device sends a preset voice to the call center device, and the control device sends a system hang-up instruction to the call center device. The third duration is greater than or equal to the sum of the second duration and the preset duration.
2. The method according to claim 1, wherein The control device obtains the first instruction, including: The control device receives the user hang-up instruction sent by the call center device and determines the user hang-up instruction as the first instruction; or, The control device receives the user voice sent by the call center device and determines the first instruction according to the user voice. The first instruction is the interruption instruction or the system hang-up instruction.
3. The method according to claim 2, characterized in that, The first instruction is an interruption instruction; determining the first instruction according to the user voice information includes: The control device sends the user voice to the voice conversion device. The control device receives the user text corresponding to the user voice sent by the voice conversion device. The control device obtains the number of characters included in the user text, and the time information of the first response voice sent by the control device to the call center device last time. The time information includes the first sending time and the first duration of the first response voice. Determine the interruption instruction according to the number of characters included in the user text and the time information.
4. The method according to claim 3, wherein Determining the interruption instruction according to the number of characters included in the user text and the time information includes: Determine whether the call center device is playing the first response voice according to the current time, the first sending time, and the first duration of the first response voice. If so, when the number of characters included in the user text is greater than or equal to a preset threshold, the interruption instruction is generated.
5. The method according to claim 2, wherein The first instruction is a system hang-up instruction; determining the first instruction according to the user voice information includes: The control device sends the user voice to the voice conversion device. The control device receives the user text corresponding to the user voice sent by the voice conversion device. The control device sends the user text to the answering device. The control device receives the system hang-up instruction sent by the answering device. The control device determines the system hang-up instruction as the first instruction.
6. The method according to any one of claims 1-5, characterized in that, The control device is respectively connected to the call center device, the voice conversion device, and the voice synthesis device through websockets; the websocket connections are used to transmit control instructions and multimedia information, the control instructions include the first instruction and the pause instruction, and the multimedia information includes text information and voice information.
7. A voice system, characterized in that, It includes a control device, a call center device, a voice conversion device, an answering device, and a voice synthesis device. There is a call connection between the call center device and the user device. Among them, The call center device is used to send the user voice or the user hang-up instruction to the control device through a websocket connection. The control device is used to send the user voice to the voice conversion device through a websocket connection. The voice conversion device is used to convert the user voice into a user text and send the user text to the control device through a websocket connection. The control device is further used to determine an interruption instruction according to the user text. The control device is further used to send the user text to the answering device. The answering device is used to determine a system hang-up instruction according to the user text and send the system hang-up instruction to the control device. The control device is further used to determine the interruption instruction, the system hang-up instruction, or the user hang-up instruction as the first instruction, and generate a pause instruction according to the first instruction. The control device is further used to send the pause instruction to the voice conversion device, the answering device, and the voice synthesis device through a websocket connection. The pause instruction is used to instruct the voice conversion device, the answering device, and the voice synthesis device to pause the processing tasks corresponding to the call connection and clear the voice generated by the voice conversion device in the buffer area. The control device is further used to obtain the second sending moment and the second duration of the second answering voice after sending the second answering voice to the call center device. If no user voice is received within a third duration after the second sending moment, a preset operation is performed. The preset operation includes: the control device sends a preset voice to the call center device, and the control device sends a system hang-up instruction to the call center device. The third duration is greater than or equal to the sum of the second duration and a preset duration.
8. The voice system according to claim 7, characterized in that, The control device is further used to execute the method according to any one of claims 2-6.
9. A voice control device, characterized in that, Applied to a voice system, the voice system includes a control device, a call center device, a voice conversion device, a response device, and a voice synthesis device, and there is a call connection between the call center device and a user device; the voice control device includes a first acquisition module, a generation module, a sending module, and a second acquisition module, wherein: The first acquisition module is configured to receive, through a websocket connection, a first instruction sent by the call center device, the first instruction being a call hang-up instruction or an interruption instruction for the call connection, the interruption instruction being used to instruct user voice to interrupt the voice playback of the call center device, and the call hang-up instruction including a user hang-up instruction and a system hang-up instruction; The generation module is configured to generate, according to the first instruction, a pause instruction, the pause instruction including a call identifier of the voice call; The sending module is configured to send, through a websocket connection, the pause instruction to the voice conversion device, the response device, and the voice synthesis device, the pause instruction being used to instruct the voice conversion device, the response device, and the voice synthesis device to pause the processing tasks corresponding to the call connection and empty the voice generated by the voice conversion device in the buffer; The second acquisition module is configured to, after the control device sends a second response voice to the call center device, the control device acquires a second sending time and a second duration of the second response voice; If the control device does not receive user voice within a third duration after the second sending time, the control device performs a preset operation, the preset operation including: the control device sends a preset voice to the call center device, the control device sends a system hang-up instruction to the call center device, and the third duration is greater than or equal to the sum of the second duration and a preset duration.
10. A voice control device, characterized in that, The voice control device includes: a memory, a processor, and a voice control program stored on the memory and executable on the processor, and when the voice control program is executed by the processor, the steps of the voice control method as described in any one of claims 1 to 6 are implemented.
11. A computer-readable storage medium, characterized in that, A voice control program is stored on the computer-readable storage medium, and when the voice control program is executed by a processor, the steps of the voice control method as described in any one of claims 1 to 6 are implemented.
12. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, the voice control method as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Voice interruption method and device
CN111540349A