Voice Processing Method, Device, Equipment, Storage Medium

By directly transmitting data using Websocket connection in an intelligent outbound call system, the problem of low voice processing efficiency in the prior art is solved, and more efficient and real-time voice information processing is achieved.

CN113611312BActive Publication Date: 2025-07-22WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110963645.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-20
Publication Date
2025-07-22
Estimated Expiration
2041-08-20

AI Technical Summary

Technical Problem

The communication connection between devices in the existing intelligent out-call system obtains media information through polling, resulting in poor real-time transmission of media information, which in turn affects the efficiency of voice processing.

Method used

Websocket connection is used to directly transmit data between the control device and the call center device, voice conversion device, answering device and voice synthesis device in the intelligent outbound call system, avoid polling and realize real-time voice information processing.

Benefits of technology

Connecting through Websocket reduces the complexity of the voice system, improves the efficiency and real-timeness of voice processing, and ensures that users can quickly obtain answering voice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113611312B_ABST
    Figure CN113611312B_ABST
Patent Text Reader

Abstract

The present invention discloses a voice processing method, apparatus, device, and storage medium, which are applied to a voice system. The method includes: the control device receives the first voice information of the user from the call center device through a Websocket connection, and the control device is connected to the call center device through a Websocket connection; the control device obtains the first text corresponding to the first voice information from the voice conversion device through a Websocket connection, and the control device is connected to the voice conversion device through a Websocket connection; the control device obtains the response text corresponding to the first text from the response device, and obtains the response voice corresponding to the response text from the voice synthesis device through a Websocket connection; the control device sends the response voice to the call center device through a Websocket connection, so that the call center device sends the response voice to the client of the user. The efficiency of voice processing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent voice technology, and in particular, to a voice processing method, device, equipment, and storage medium. Background Art

[0002] An intelligent outbound calling system can automatically call a user's client and conduct simple voice communication with the user through an intelligent robot. For example, the intelligent outbound calling system generates a response voice corresponding to the voice information according to the user's voice information to communicate with the user.

[0003] Currently, when the intelligent outbound calling system generates a response voice, communication connections need to be established between various devices in the intelligent outbound calling system to transmit media information. For example, a voice gateway sends the user's voice to a voice conversion device through a communication connection. However, the existing communication connections between intelligent outbound calling systems obtain media information through a polling method. For example, the voice conversion device sends a voice acquisition request to the voice gateway at regular intervals, and the voice gateway can send the user voice to the voice conversion device only when it receives the voice acquisition request. This makes the real-time performance of media information transmission between intelligent outbound calling systems poor, and thus leads to low voice processing efficiency. Summary of the Invention

[0004] The main purpose of the present invention is to provide a voice processing method, device, equipment, and storage medium, aiming to solve the technical problem of low voice processing efficiency in the prior art.

[0005] To achieve the above object, in a first aspect, an embodiment of the present invention provides a voice processing method, which is applied to a voice system. The voice system includes a control device, a call center device, a voice conversion device, a response device, and a voice synthesis device. The method includes:

[0006] The control device receives the user's first voice information from the call center device through a Websocket connection. The control device is connected to the call center device through a Websocket connection;

[0007] The control device obtains the first text corresponding to the first voice information from the voice conversion device through a Websocket connection. The control device is connected to the voice conversion device through a Websocket connection;

[0008] The control device obtains the response text corresponding to the first text from the response device and obtains the response voice corresponding to the response text from the voice synthesis device through a Websocket connection;

[0009] The control device sends the response voice to the call center device through a Websocket connection, so that the call center device sends the response voice to the client of the user.

[0010] In a possible implementation manner, after the control device obtains the first text corresponding to the first voice information from the voice conversion device through a Websocket connection, it further includes:

[0011] The control device obtains the number of characters included in the first text, and the time information of the historical response voice that the control device sent to the call center device last time, where the time information includes the first sending moment and the first duration of the first response voice;

[0012] The control device generates an interruption instruction according to the number of characters included in the first text and the time information.

[0013] In a possible implementation manner, generating an interruption instruction according to the number of characters included in the first text and the time information includes:

[0014] Determine whether the call center device is playing the first response voice according to the current moment, the first sending moment, and the first duration of the historical response voice;

[0015] If so, when the number of characters included in the first text is greater than or equal to a preset threshold, generate the interruption instruction.

[0016] In a possible implementation manner, the control device obtains the first text corresponding to the first voice information from the voice conversion device through a Websocket connection, including:

[0017] The control device sends the first voice information to the voice conversion device through a Websocket connection;

[0018] The control device receives the first text sent by the voice conversion device through a Websocket connection.

[0019] In a possible implementation manner, obtaining the response voice corresponding to the response text from the speech synthesis device through a Websocket connection includes:

[0020] The control device sends the response text to the speech synthesis device through a Websocket connection;

[0021] The control device receives the response voice from the speech synthesis device through a Websocket connection.

[0022] In a possible implementation, the control device obtains the response text corresponding to the first text from the response device, including:

[0023] The control device sends the first text to the response device through an HTTP connection;

[0024] The control device receives the response text from the response device through an HTTP connection.

[0025] In a possible implementation, before the control device receives the first voice information of the user from the call center device through a Websocket connection, it includes:

[0026] The control device receives the Websocket connection establishment request sent by the call center device;

[0027] The control device establishes a Websocket connection with the call center device according to the Websocket connection establishment request.

[0028] In a possible implementation, before the control device obtains the first text corresponding to the first voice information from the voice conversion device through a Websocket connection, it further includes:

[0029] After the control device determines that a call connection is established between the client and the call center device, it sends a Websocket connection establishment request to the voice conversion device;

[0030] The control device receives the Websocket connection establishment response corresponding to the Websocket connection establishment request sent by the voice conversion device;

[0031] The control device establishes a Websocket connection with the voice conversion device according to the Websocket connection establishment response.

[0032] In a possible implementation, before the control device obtains the response voice corresponding to the response text from the voice synthesis device through a Websocket connection, it includes:

[0033] After the control device determines that a call connection is established between the client and the call center device, it sends a Websocket connection establishment request to the voice synthesis device;

[0034] The control device receives the Websocket connection establishment response corresponding to the Websocket connection establishment request sent by the voice synthesis device;

[0035] The control device establishes a Websocket connection with the speech synthesis device according to the Websocket connection establishment response.

[0036] In a second aspect, an embodiment of the present application provides a voice system, including a control device, a call center device, a voice conversion device, a response device, and a speech synthesis device, wherein,

[0037] The call center device is configured to send the first voice information of the user to the control device through a Websocket connection;

[0038] The control device is configured to send the first voice information to the voice conversion device through a Websocket connection;

[0039] The voice conversion device is configured to convert the first voice information into a first text, and send the first text to the control device through a Websocket connection;

[0040] The control device is further configured to send the first text to the response device;

[0041] The response device is configured to determine a response text corresponding to the first text, and send the response text to the control device;

[0042] The control device is further configured to send the response text to the speech synthesis device through a Websocket connection;

[0043] The speech synthesis device is configured to convert the response text into a response voice, and send the response voice to the control device through a Websocket connection;

[0044] The control device is further configured to send the response voice to the call center device through a Websocket connection;

[0045] The call center device is further configured to send the response voice to the client of the user.

[0046] In a possible implementation manner, the control device is further configured to execute the method described in the first aspect.

[0047] In a third aspect, an embodiment of the present application provides a voice processing device, which is applied to a voice system. The voice system includes a control device, a call center device, a voice conversion device, a response device, and a speech synthesis device. The voice processing device includes a receiving module, a first obtaining module, a second obtaining module, and a sending module, wherein:

[0048] The receiving module is configured to receive the first voice information of the user from the call center device through a Websocket connection, and the control device is connected to the call center device through a Websocket connection;

[0049] The first obtaining module is configured to obtain the first text corresponding to the first voice information from the voice conversion device through a Websocket connection, and the control device is connected to the voice conversion device through a Websocket connection;

[0050] The second obtaining module is configured to obtain the response text corresponding to the first text from the response device, and obtain the response voice corresponding to the response text from the voice synthesis device through a Websocket connection;

[0051] The sending module is configured to send the response voice to the call center device through a Websocket connection, so that the call center device sends the response voice to the client of the user.

[0052] In a possible implementation manner, the first obtaining module is specifically configured to:

[0053] Send the first voice information to the voice conversion device through a Websocket connection;

[0054] Receive the first text sent by the voice conversion device through a Websocket connection.

[0055] In a possible implementation manner, the first obtaining module is specifically configured to:

[0056] Send the response text to the voice synthesis device through a Websocket connection;

[0057] Receive the response voice from the voice synthesis device through a Websocket connection.

[0058] In a possible implementation manner, the second obtaining module is specifically configured to:

[0059] Send the first text to the response device through an HTTP connection;

[0060] Receive the response text from the response device through an HTTP connection.

[0061] In another possible implementation manner, the receiving module is further configured to:

[0062] Receive the Websocket connection establishment request sent by the call center device;

[0063] Establish a Websocket connection with the call center device according to the Websocket connection establishment request.

[0064] In another possible implementation, the sending module is further configured to:

[0065] After determining that the client has established a call connection with the call center device, the control device sends a Websocket connection establishment request to the voice conversion device;

[0066] The control device receives a Websocket connection establishment response corresponding to the Websocket connection establishment request sent by the voice conversion device;

[0067] The control device establishes a Websocket connection with the voice conversion device according to the Websocket connection establishment response.

[0068] In another possible implementation, the sending module is further configured to:

[0069] After determining that the client has established a call connection with the call center device, the control device sends a Websocket connection establishment request to the voice synthesis device;

[0070] The control device receives a Websocket connection establishment response corresponding to the Websocket connection establishment request sent by the voice synthesis device;

[0071] The control device establishes a Websocket connection with the voice synthesis device according to the Websocket connection establishment response.

[0072] In another possible implementation, the first acquisition module is further configured to:

[0073] The control device acquires the number of characters included in the first text and the time information of the historical response voice that the control device last sent to the call center device, where the time information includes the first sending moment and the first duration of the first response voice;

[0074] The control device generates an interruption instruction according to the number of characters included in the first text and the time information.

[0075] In a possible implementation, the first acquisition module is configured to:

[0076] Determine whether the call center device is playing the first response voice according to the current moment, the first sending moment, and the first duration of the historical response voice;

[0077] If so, when the number of characters included in the first text is greater than or equal to a preset threshold, the interruption instruction is generated.

[0078] In a fourth aspect, an embodiment of the present application provides a voice processing device, including a processor and a memory;

[0079] The memory stores computer execution instructions;

[0080] The processor executes the computer execution instructions stored in the memory, so that the processor executes the voice processing method as described in the first aspect.

[0081] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the voice processing method described in the first aspect.

[0082] In a sixth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the voice processing method described in the first aspect.

[0083] An embodiment of the present invention provides a voice processing method, device, equipment, and storage medium, which are applied to a voice system. The voice system includes a control device, a call center device, a voice conversion device, a response device, and a voice synthesis device. The control device is respectively connected to the call center device, the voice conversion device, and the voice synthesis device through Websocket. The control device receives the user's first voice information from the call center device through the Websocket connection. The control device obtains the first text corresponding to the first voice information from the voice conversion device through the Websocket connection. The control device obtains the response text corresponding to the first text from the response device, and obtains the response voice corresponding to the response text from the voice synthesis device through the Websocket connection. The control device sends the response voice to the call center device through the Websocket connection, so that the call center device sends the response voice to the user's client. In this way, the Websocket connection can directly perform data transmission without polling, which can not only reduce the complexity of the voice system, but also the control device can obtain the response voice corresponding to the user's voice in real time through the voice conversion device, the response device, and the voice conversion device, thereby improving the efficiency of voice processing. Description of the Drawings

[0084] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following briefly introduces the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0085] Figure 1 It is a schematic structural diagram of a voice system provided by an embodiment of the present application;

[0086] Figure 2 It is a schematic flowchart of a voice processing method provided by an embodiment of the present application;

[0087] Figure 3 It is a schematic diagram of the process of a control device receiving the first voice information provided by an embodiment of the present application;

[0088] Figure 4 It is a schematic diagram of the process of a control device obtaining the first text provided by an embodiment of the present application;

[0089] Figure 5 It is a schematic diagram of the process of obtaining a response voice provided by an embodiment of the present application;

[0090] Figure 6 It is a flowchart of a connection method between a control device and a call center device provided by an embodiment of the present application;

[0091] Figure 7 It is a schematic diagram of a connection establishment method between a control device and a voice synthesis device and a voice conversion device provided by an embodiment of the present application;

[0092] Figure 8 It is a schematic structural diagram of a voice processing device provided by an embodiment of the present application;

[0093] Figure 9 It is a schematic hardware structure diagram of a voice processing device provided by the present application.

[0094] The realization of the object of the present invention, functional features and advantages will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0095] The following will describe the exemplary embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0096] In the related art, communication connections need to be established between various devices in an intelligent outbound call system to transmit media information. For example, a voice gateway sends user voice to a voice conversion device and sends response text to a voice answering device through a communication connection. However, the existing communication connections between intelligent outbound call systems obtain media information through a polling method. For example, the voice answering device sends a text acquisition request to the voice gateway at regular intervals, and only when the voice gateway receives the text acquisition request can it send the text corresponding to the user voice to the voice answering device. This makes the real-time performance of media information transmission between intelligent outbound call systems poor, and thus leads to low efficiency of voice processing.

[0097] To solve the technical problem of low efficiency of voice processing in the related art, an embodiment of the present application provides a voice processing method, which is applied to a voice system. The voice system includes a control device, a call center device, a voice conversion device, a response device, and a voice synthesis device. The control device receives the first voice information of the user from the call center device through a Websocket connection, and sends the first voice information to the voice conversion device through the Websocket connection. The voice conversion device converts the first voice information into corresponding text information, and sends the text information to the control device through the Websocket connection. When the control device receives the text information, it sends the text information to the response device through the Hyper Text Transfer Protocol (HTTP). The response device determines the response text according to the text information, and sends the response text to the control device through the HTTP. When the control device receives the response text, it sends the response text to the voice synthesis device through the Websocket connection. After the voice synthesis device synthesizes the response text into voice, it sends the response voice to the control device through the Websocket connection. When the control device receives the response voice, it sends the response voice to the call center device through the Websocket connection, so that the call center device sends the response voice to the client of the user. In this way, the complexity of the voice system is reduced through the Websocket connection. Since the Websocket connection can directly perform data transmission without polling, through the control device, the first voice information can be processed flexibly and in real time. When obtaining the first voice information of the user, the voice system can synthesize the response voice corresponding to the first voice information in real time, improving the efficiency of voice processing.

[0098] Next, in combination with Figure 1 , the structure of the voice system involved in the present application will be described.

[0099] Figure 1 The following is a schematic structural diagram of a voice system provided by an embodiment of the present application. Please refer to Figure 1, including a voice system and a client. Among them, the voice system includes a control device, a call center device, a voice conversion device, a response device, and a voice synthesis device. The voice system can receive the voice information input by the user at the client, and determine the response voice according to the received voice information. When determining the response voice, the voice system can send the response voice to the client so that the user can have a conversation through the client and the voice system, which can improve the efficiency of voice processing.

[0100] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the accompanying drawings.

[0101] Figure 2 It is a schematic flowchart of a voice processing method provided by an embodiment of this application. Please refer to Figure 2 , the method may include:

[0102] S201. The call center device sends the user's first voice information to the control device through a Websocket connection, and the control device is connected to the call center device through a Websocket connection.

[0103] The execution subject of the embodiment of this application can be a voice system or a voice processing device set in the voice system. The voice processing device can be implemented by software or by a combination of software and hardware.

[0104] Websocket is a protocol for full-duplex communication on a single Transmission Control Protocol (TCP) connection. For example, after device A and device B are connected through Websocket, two-way data transmission can be carried out simultaneously. The control device and the call center device are connected through Websocket. Optionally, the call center device can call the user's client. For example, when the call center device receives the user's phone number, the call center device can make a call to the user's phone number. Optionally, when the call center device calls the user's client, the call center device can create a session room for this call, and the voice information of this call is saved in the session room. Optionally, the call center device may include an outbound management system and a call center middleware. The outbound management system can obtain the accounts of the user's clients and create a session room for each account respectively, and the call center middleware can call the accounts of the clients.

[0105] The call center device is used to send the first voice message of the user to the control device through a Websocket connection. For example, after the control device and the call center device are connected through a Websocket connection, the call center device can send the first voice message of the user to the control device through the Websocket connection. The first voice message is the voice message output by the user's client. For example, the call center device can make a call to the user's mobile phone. After the user answers the call, the mobile phone can send the voice message after the user answers to the call center device. For example, after the call center device establishes a call connection with the user's mobile phone, if the voice uttered by the user is "Hello", the first voice message received by the call center device is "Hello". When the call center device receives the first voice message, it can directly send the first voice message "Hello" to the control device through the Websocket connection. For example, when the call center device obtains the first voice message, the call center device can send the first voice message to the control device, and the control device receives the first voice message according to the Websocket connection.

[0106] Next, in conjunction with Figure 3 , the process of the control device receiving the first voice message of the user from the call center device through the Websocket connection will be described.

[0107] Figure 3 It is a schematic diagram of the process for a control device to receive the first voice message provided by an embodiment of the present application. Please refer to Figure 3 , which includes a control device, a call center device, and a client. Among them, after the call center device calls the client, the client accepts the call from the call center device, and the call center device establishes a Websocket connection with the control device.

[0108] Please refer to Figure 3 , the client sends the voice message "What's the weather like today" to the call center device. The call center device receives the voice message sent by the client and sends the voice message to the control device through the Websocket connection.

[0109] S202. The control device obtains the first text corresponding to the first voice message from the voice conversion device through the Websocket connection. The control device is connected to the voice conversion device through the Websocket connection.

[0110] The voice conversion device is used to convert the first voice message into the first text and send the first text to the control device through the Websocket connection. The first text is the text content corresponding to the first voice message. For example, if the first voice message is the voice "What did you have for breakfast", the first text is the text "What did you have for breakfast".

[0111] A voice conversion device can convert voice information into text information. For example, the voice conversion device can perform speech recognition on the first voice information sent by the control device and convert the recognized voice into the corresponding text. For example, if the voice received by the voice conversion device is "How's the weather today", the voice conversion device can convert this voice into the text content "How's the weather today". For example, the voice conversion device can be a device supporting ASR technology. When the ASR receives voice information, it can convert the voice information into text information.

[0112] The control device and the voice conversion device are connected via Websocket. Optionally, the control device can obtain the first text corresponding to the first voice information through the following feasible implementation methods: The control device sends the first voice information to the voice conversion device via the Websocket connection. For example, when the control device receives the first voice information of the user sent by the call center device, the control device can send this first voice information to the voice conversion device via the Websocket connection.

[0113] The control device receives the first text sent by the voice conversion device via the Websocket connection. For example, when the voice conversion device receives the first voice information sent by the control device, the voice conversion device can convert the first voice information into the first text in real time and send the first text corresponding to the first voice information to the control device in real time via the Websocket connection. For example, if the user voice obtained by the control device from the call center device via the Websocket connection is "Hello", the control device sends the voice "Hello" to the voice conversion device via the Websocket connection. After the voice conversion device receives the voice "Hello" via the Websocket connection, the voice conversion device recognizes the voice and converts the recognized voice into the text "Hello". The voice conversion device can send the text "Hello" to the control device via the Websocket connection.

[0114] Next, in combination with Figure 4 , the process of the control device obtaining the first text corresponding to the first voice information from the voice conversion device via the Websocket connection will be described.

[0115] Figure 4 This is a schematic diagram of the process for a control device to obtain the first text provided by an embodiment of this application. Please refer to Figure 4, including a control device, a call center device, and a voice conversion device. Among them, the call center device is connected to the control device via a Websocket, and the control device is connected to the voice conversion device via a Websocket. The first voice message sent by the call center to the control device is "Is it raining today?". When the control device receives the first voice message, it sends the first voice message to the voice conversion device via the Websocket connection. When the voice conversion device receives the voice message "Is it raining today?", it can process the voice message, convert the voice message into the text content "Is it raining today?", and send the text content "Is it raining today?" to the control device via the Websocket connection.

[0116] Optionally, after the control device obtains the first text, the control device can also generate an interruption instruction based on the first text. The interruption instruction can be generated according to the following feasible implementation methods: The control device obtains the number of characters included in the first text and the time information of the historical response voice sent by the control device to the call center device last time. The number of characters is the number of characters in the first text. For example, if the first text includes 10 characters, the number of characters is 10. The historical response voice is the response voice sent by the control device to the call center device last time. For example, during the process of a user communicating with the voice system, the control device can send the response voice corresponding to the user's voice to the call center device, and the historical response voice can be the last response voice sent at the current moment. The time information includes the first sending moment and the first duration of the historical response voice. The first sending moment is the moment when the historical response voice is sent. The first duration is the playing duration of the historical response voice. For example, if the playing duration of the historical response voice is 10 seconds, the first duration of the historical response voice is 10 seconds. Optionally, the control device can determine the first duration of the historical response voice according to the voice slices of the historical response voice. For example, if the control device receives 10 voice slices corresponding to the historical response voice, and each voice slice is 20 milliseconds, the first duration of the historical response voice is 200 milliseconds.

[0117] Determine the interruption instruction according to the number of characters included in the first text and the time information. Optionally, it is possible to determine whether the call center device is playing the historical response voice according to the current moment, the first sending moment, and the first duration of the historical response voice. For example, determine the time difference according to the current moment and the first sending moment, and then determine whether the call center device is playing the historical response voice according to the first duration and the time difference. For example, when the time difference between the current moment and the first sending moment is 10 seconds, if the first duration is 5 seconds, it is determined that the call center device has completed the playing of the historical response voice; if the first duration is 15 seconds, it is determined that the call center device has not completed the playing of the historical response voice and the call center device is playing the historical response voice.

[0118] If the call center device is playing historical response voice, when the number of characters included in the first text is greater than or equal to a preset threshold, an interruption instruction is generated. For example, when the call center device is playing historical response voice, if the number of characters in the first text obtained by the control device from the voice conversion device is greater than or equal to the preset threshold, it indicates that the user is sending user voice through the user device. At this time, the control device can determine that the user interrupts the voice played by the call center device, and the control device generates an interruption instruction; if the number of characters in the first text obtained by the control device from the voice conversion device is less than the preset threshold, it indicates that the user voice sent by the user device is invalid voice (such as environmental noise, etc.). At this time, the control device determines that the user does not interrupt the voice played by the call center device, and the control device does not generate an interruption instruction. The control device can flexibly determine whether to generate an interruption instruction based on the number of characters in the first text and the time information of the historical response voice, thereby improving the flexibility of the voice processing of the control device. S203. The control device obtains the response text corresponding to the first text from the response device, and obtains the response voice corresponding to the response text from the speech synthesis device through a Websocket connection.

[0119] Optionally, the control device is further configured to send the first text to the response device. The response device is configured to determine the response text corresponding to the first text and send the response text to the control device. Optionally, the response device may include a multi-round session management system DM and a natural language understanding system NLU. For example, the response device can obtain multiple response texts corresponding to the first text through the multi-round session management system, and the natural speech understanding system can determine the response text corresponding to the first text from the multiple response texts.

[0120] Optionally, the control device and the response device are connected through an HTTP connection. The control device can obtain the response text corresponding to the first text according to the following feasible implementation methods: The control device sends the first text to the response device through an HTTP connection. For example, when the control device receives the first text sent by the voice conversion device, the control can send the first text to the response device through an HTTP connection. For example, after the control device receives the text "Hello" sent by the voice conversion device, it can send the text "Hello" to the response device through an HTTP connection. Optionally, the control device and the response device can also be connected through a Websocket connection, and then send the first text through the Websocket connection. In this way, the first text can be quickly sent to the response device through HTTP, improving the real-time performance of voice processing.

[0121] The control device receives a response text from the response device via an HTTP connection. Optionally, the response device may determine the response text corresponding to the first text based on the first text. For example, when the response device receives the first text, the response device simulates the response results of multiple scenarios of the first text through a multi-round session management system (DM), and then obtains the response texts of the first text in multiple scenarios. The response device then determines a unique response text corresponding to the first text from the response texts in multiple scenarios through a natural language understanding system (NLU). After the response device determines the response text, the response device may send the response text corresponding to the first text to the control device via HTTP.

[0122] The control device is connected to the speech synthesis device via a Websocket connection. The control device is also used to send the response text to the speech synthesis device via the Websocket connection. For example, when the control device receives the response text corresponding to the first text sent by the response device, the control device may send the response text to the speech synthesis device via the Websocket connection.

[0123] The speech synthesis device is used to convert the response text into a response voice and send the response voice to the control device via the Websocket connection. For example, the speech synthesis device may be a TTS. When the TTS receives the response text, the TTS may convert the response text into a response voice. For example, if the response text corresponding to the first text is the text content "It is sunny today", then the TTS may convert this text content into the voice message "It is sunny today".

[0124] Optionally, the control device sends the response text to the speech synthesis device via the Websocket connection and receives the response voice from the speech synthesis device via the Websocket connection. For example, when the control device obtains the response text corresponding to the first text through the response device, the control device may send the response text to the speech synthesis device via the Websocket. When the speech synthesis device receives the response text, it may convert the response text into a response voice and send the response voice to the control device. For example, the first text sent by the control device to the response device is "What's the weather today", and the response device determines the response text as "It is sunny today" through the first text. The control device sends the text content "It is sunny today" to the speech synthesis device, and the speech synthesis device may generate the voice message "It is sunny today" and send the voice message to the control device.

[0125] Next, in conjunction with Figure 5 , the process of the control device obtaining the response voice corresponding to the response text will be described.

[0126] Figure 5 FIG. [FIGURE NUMBER] is a schematic diagram of a process for obtaining a response voice provided by an embodiment of the present application. Please refer toFigure 5 , including a voice conversion device, a control device, a response device, and a voice synthesis device. Among them, the voice conversion device is connected to the control device through Websocket, the response device is connected to the control device through HTTP, and the control device and the voice synthesis device are connected through Websocket.

[0127] Please refer to Figure 5 , the voice conversion device sends the first text "How's the weather today" to the control device through the Websocket connection. When the control device receives the first text through the Websocket, it sends the first text to the response device through HTTP. When the response device receives the first text through HTTP, it can determine the response text "It's raining heavily today" corresponding to the first text according to the first text, and send the response text to the control device through HTTP. When the control device receives the response text, it sends the response text to the voice synthesis device through the Websocket. The voice synthesis device generates a response voice "It's raining heavily today" through the response text, and sends the response voice "It's raining heavily today" to the control device through the Websocket.

[0128] S204. The control device sends the response voice to the call center device through the Websocket connection, so that the call center device sends the response voice to the user's client.

[0129] Optionally, the control device is further configured to send the response voice to the call center device through the Websocket connection. For example, when the control device receives the response voice sent by the voice synthesis device, the control device can send the response voice to the call center device through the Websocket connection.

[0130] Optionally, the call center device is further configured to send a response voice to the user's client. For example, when the call center device receives the response voice sent by the control device, it can send the response voice to the user's client. For example, if the response voice corresponding to the first voice message received by the control device is "It's raining heavily today", the control device can send this response voice to the call center device. When the call center device receives the response voice, it can send this response voice to the user's client so that the user can hear the voice "It's raining heavily today" through the client. For example, the user sends a user voice "What's the weather like today" to the call center device through the client. After the call center device receives the user voice, it sends the user voice to the control device. The control device can send the user voice "What's the weather like today" to the voice conversion device. The voice conversion device recognizes the user voice "What's the weather like today" to obtain the user text "What's the weather like today". The voice conversion device sends the user text to the control device. After the control device receives the user text, it can send the user text to the response device. The response device generates a response text "It's sunny today" according to the user text and sends the response text to the control device. When the control device receives the response text, it sends the response text to the voice synthesis device. The voice synthesis device converts the response text into a response voice "It's sunny today" and sends the response voice to the control device. When the control device receives the response voice, it sends the response voice to the call center device. After the call center device receives the response voice, it can play the response voice "It's sunny today" to the user's client.

[0131] An embodiment of the present application provides a voice processing method. The call center device sends the user's first voice message to the control device through a Websocket connection. The control device is connected to the call center device through a Websocket connection. The control device obtains the first text corresponding to the first voice message from the voice conversion device through a Websocket connection. The control device is connected to the voice conversion device through a Websocket connection. The control device obtains the response text corresponding to the first text from the response device and obtains the response voice corresponding to the response text from the voice synthesis device through a Websocket connection. The control device sends the response voice to the call center device through a Websocket connection so that the call center device sends the response voice to the user's client. According to the above method, the control device in the voice system is respectively connected to the call center device, the voice conversion device and the voice synthesis device through a Websocket connection, reducing the complexity of the voice system and the cost of the voice system. Moreover, the control device can generate the response voice corresponding to the first voice message in real time, enabling the user to quickly obtain the response voice through the client and improving the efficiency of voice generation.

[0132] In Figure 2Based on the illustrated embodiments, before the control device receives the first voice message of the user from the call center device through a Websocket connection, the above voice processing method further includes the connection process between the control device and the call center device. Next, in combination with Figure 6 , the connection process between the control device and the call center device will be described.

[0133] Figure 6 FIG. is a flowchart of a method for connecting a control device and a call center device provided by an embodiment of the present application. Please refer to Figure 6 , the method includes:

[0134] S601. The control device receives a Websocket connection establishment request sent by the call center device.

[0135] Optionally, when the call center device calls the client of the user, if a call connection is established between the client of the user and the call center device, the call center device sends a Websocket connection establishment request to the control device. For example, when the call center device calls the user's mobile phone, if the user answers the call, the call center device sends a Websocket connection establishment request to the control device; if the user does not answer the call, the call center device does not send a Websocket connection establishment request to the control device. In this way, after the call center device determines that a call connection has been established with the client of the user, the call center device will send a Websocket connection establishment request to the control device, thereby avoiding the situation of resource waste caused by the control device establishing a Websocket connection with the call center device while the client of the user has not established a call connection with the call center device, and improving the resource utilization rate of the voice system.

[0136] Optionally, the call center device may also send a Websocket connection establishment request to the control device before calling the client of the user.

[0137] S602. The control device establishes a Websocket connection with the call center device according to the Websocket connection establishment request.

[0138] Optionally, when the control device receives the Websocket connection request sent by the call center device, it can establish a Websocket connection with the call center device.

[0139] An embodiment of the present application provides a method for establishing a Websocket connection between a control device and a call center device. After the call center device determines that a call connection has been established with the client of the user, the call center device sends a Websocket connection establishment request to the control device, which can avoid resource waste and improve the resource utilization rate of the voice system.

[0140] Based on any of the above embodiments, the voice processing method of the present application further includes the process of establishing a Websocket connection between the control device and the voice conversion device and the voice synthesis device. Next, in combination with Figure 7 , the process of establishing a Websocket connection between the control device and the voice synthesis device and the voice conversion device will be described.

[0141] Figure 7 It is a schematic diagram of a method for establishing a connection between a control device and a voice synthesis device and a voice conversion device provided by an embodiment of the present application. Please refer to Figure 7 , this method includes:

[0142] S701. After the control device determines that a call connection is established between the client and the call center device, it sends a Websocket connection establishment request to the voice conversion device and the voice synthesis device.

[0143] Optionally, before the control device obtains the first text corresponding to the first voice information from the voice conversion device through the Websocket connection, the control device can also determine whether a call connection is established between the client and the call center device. If a call connection is established between the client and the call center device, the control device sends a Websocket connection establishment request to the voice conversion device. For example, after the control device determines that a call connection is established between the client and the call center device, it sends a Websocket connection establishment request to the voice conversion device.

[0144] Optionally, before the control device obtains the response voice corresponding to the response text from the voice synthesis device through the Websocket connection, the control device can also determine whether a call connection is established between the client and the call center device. If a call connection is established between the client and the call center device, the control device sends a Websocket connection establishment request to the voice synthesis device. For example, after the control device determines that a call connection is established between the client and the call center device, it sends a Websocket connection establishment request to the voice synthesis device.

[0145] Optionally, the control device can send a Websocket connection establishment request to the voice synthesis device and the voice conversion device simultaneously, or send a Websocket connection establishment request to the voice synthesis device and the voice conversion device in a preset order. The embodiments of the present application do not limit this.

[0146] S702. The control device receives the Websocket connection establishment response corresponding to the Websocket connection establishment request sent by the voice conversion device and the voice synthesis device.

[0147] Optionally, when the voice conversion device receives a Websocket connection establishment request, it can send a Websocket connection establishment response to the control device. When the voice synthesis device receives a Websocket connection establishment request, it can send a Websocket connection establishment response to the control device.

[0148] S703. The control device establishes a Websocket connection with the voice conversion device and the voice synthesis device according to the Websocket connection establishment response.

[0149] The control device establishes a Websocket connection with the voice synthesis device according to the Websocket connection establishment response of the voice synthesis device, and the control device establishes a Websocket connection with the voice conversion device according to the Websocket connection establishment response of the voice conversion device.

[0150] The embodiment of the present application provides a method for establishing a connection between a control device and a voice synthesis device and a voice conversion device. After the control device determines that a call connection is established between the client and the call center device, it sends a Websocket connection establishment request to the voice conversion device and the voice synthesis device. The control device receives the Websocket connection establishment response corresponding to the Websocket connection establishment request sent by the voice conversion device and the voice synthesis device, and the control device establishes a Websocket connection with the voice conversion device and the voice synthesis device according to the Websocket connection establishment response. According to the above method, the resource utilization rate of the voice system can be improved, the complexity of the voice system can be reduced, and the efficiency of voice processing can be improved.

[0151] Figure 8 It is a schematic structural diagram of a voice processing device provided by an embodiment of the present application. Please refer to Figure 8 , the voice processing device 10 is applied to a voice system. The voice system includes a control device, a call center device, a voice conversion device, a response device, and a voice synthesis device. The voice processing device 10 includes a receiving module 11, a first obtaining module 12, a second obtaining module 13, and a sending module 14, where:

[0152] The receiving module 11 is configured to receive the first voice information of the user from the call center device through a Websocket connection. The control device is connected to the call center device through a Websocket connection;

[0153] The first obtaining module 12 is configured to obtain the first text corresponding to the first voice information from the voice conversion device through a Websocket connection. The control device is connected to the voice conversion device through a Websocket connection;

[0154] The second acquisition module 13 is configured to acquire the response text corresponding to the first text from the response device, and acquire the response voice corresponding to the response text from the speech synthesis device through a Websocket connection;

[0155] The sending module 14 is configured to send the response voice to the call center device through a Websocket connection, so that the call center device sends the response voice to the client of the user.

[0156] In a possible implementation manner, the first acquisition module 12 is specifically configured to:

[0157] Send the first voice message to the speech conversion device through a Websocket connection;

[0158] Receive the first text sent by the speech conversion device through a Websocket connection.

[0159] In a possible implementation manner, the first acquisition module 12 is specifically configured to:

[0160] Send the response text to the speech synthesis device through a Websocket connection;

[0161] Receive the response voice from the speech synthesis device through a Websocket connection.

[0162] In a possible implementation manner, the second acquisition module 13 is specifically configured to:

[0163] Send the first text to the response device through an HTTP connection;

[0164] Receive the response text from the response device through an HTTP connection.

[0165] In another possible implementation manner, the receiving module 11 is further configured to:

[0166] Receive a Websocket connection establishment request sent by the call center device;

[0167] Establish a Websocket connection with the call center device according to the Websocket connection establishment request.

[0168] In another possible implementation manner, the sending module 14 is further configured to:

[0169] After determining that the client and the call center device establish a call connection, the control device sends a Websocket connection establishment request to the speech conversion device;

[0170] The control device receives a Websocket connection establishment response corresponding to the Websocket connection establishment request sent by the voice conversion device;

[0171] The control device establishes a Websocket connection with the voice conversion device according to the Websocket connection establishment response.

[0172] In another possible implementation manner, the sending module 14 is further configured to:

[0173] After the control device determines that a call connection is established between the client and the call center device, the control device sends a Websocket connection establishment request to the voice synthesis device;

[0174] The control device receives a Websocket connection establishment response corresponding to the Websocket connection establishment request sent by the voice synthesis device;

[0175] The control device establishes a Websocket connection with the voice synthesis device according to the Websocket connection establishment response.

[0176] In another possible implementation manner, the first acquisition module 12 is further configured to:

[0177] The control device acquires the number of characters included in the first text, and the time information of the historical response voice that the control device sent to the call center device last time, where the time information includes a first sending moment and a first duration of the first response voice;

[0178] The control device generates an interruption instruction according to the number of characters included in the first text and the time information.

[0179] In a possible implementation manner, the first acquisition module 12 is configured to:

[0180] Determine whether the call center device is playing the first response voice according to the current moment, the first sending moment, and the first duration of the historical response voice;

[0181] If so, when the number of characters included in the first text is greater than or equal to a preset threshold, generate the interruption instruction.

[0182] The voice processing device provided in the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and its implementation principles and beneficial effects are similar, and will not be described in detail here.

[0183] The speech processing device shown in the embodiment of the present application may be a chip, a hardware module, a processor, etc. Of course, the speech processing device may be in other forms, which are not specifically limited in the embodiment of the present application.

[0184] Figure 9 The hardware structure diagram of the voice processing device provided for this application. Figure 9 The speech processing device 20 may include: a processor 21 and a memory 22, wherein the processor 21 and the memory 22 can communicate; exemplarily, the processor 21 and the memory 22 communicate via a communication bus 23, the memory 22 is used to store program instructions, and the processor 21 is used to call the program instructions in the memory to execute the speech processing method shown in any of the above method embodiments.

[0185] Optionally, the speech processing device 20 may further include a communication interface, which may include a transmitter and / or a receiver.

[0186] Optionally, the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the present application may be directly implemented as being executed by a hardware processor, or may be implemented by a combination of hardware and software modules in the processor.

[0187] The embodiment of the present application provides a voice system, including a control device, a call center device, a voice conversion device, an answering device, and a voice synthesis device, wherein the call center device is used to send a user's first voice information to the control device through a Websocket connection;

[0188] The control device is used to send the first voice information to the voice conversion device through a Websocket connection;

[0189] The voice conversion device is used to convert the first voice information into a first text, and send the first text to the control device through a Websocket connection;

[0190] The control device is further used to send the first text to the answering device;

[0191] The response device is used to determine a response text corresponding to the first text, and send the response text to the control device;

[0192] The control device is further configured to send the response text to the speech synthesis device through a Websocket connection;

[0193] The speech synthesis device is configured to convert the response text into a response voice and send the response voice to the control device through a Websocket connection;

[0194] The control device is further configured to send the response voice to the call center device through a Websocket connection;

[0195] The call center device is further configured to send the response voice to the user's client.

[0196] The control device is further configured to execute the speech processing method of any one of the above embodiments.

[0197] The present application provides a readable storage medium, on which a computer program is stored; the computer program is used to implement the speech processing method as described in any of the above embodiments.

[0198] The embodiment of the present application provides a computer program product, the computer program product includes instructions, when the instructions are executed, the computer is made to execute the above speech processing method.

[0199] All or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a readable memory. When the program is executed, it executes the steps including the above method embodiments; and the foregoing memory (storage medium) includes: read-only memory (abbreviation: ROM), RAM, flash memory, hard disk, solid state drive, magnetic tape, floppy disk, optical disc and any combination thereof.

[0200] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable terminal devices to generate a machine, so that the instructions executed by the processing unit of the computer or other programmable terminal devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1a device for the functions specified in one or more boxes

[0201] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions specified in the process Figure 1 one process or more processes and / or boxes Figure 1 a box or more boxes

[0202] These computer program instructions can also be loaded onto a computer or other programmable terminal device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one process or more processes and / or boxes Figure 1 a box or more boxes

[0203] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these changes and modifications therein

[0204] In the present application, the term "including" and its variations may refer to non-limiting inclusion; the term "or" and its variations may refer to "and / or". In the present application, terms such as "first" and "second" are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. In the present application, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after

[0205] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token

Claims

1. A voice processing method, characterized in that, Applied to a voice system, the voice system includes a control device, a call center device, a voice conversion device, a response device, and a voice synthesis device. The method includes: The control device receives the user's first voice message from the call center device through a Websocket connection. The control device is connected to the call center device through a Websocket connection. The control device obtains the first text corresponding to the first voice message from the voice conversion device through a Websocket connection. The control device is connected to the voice conversion device through a Websocket connection. The control device obtains the response text corresponding to the first text from the response device and obtains the response voice corresponding to the response text from the voice synthesis device through a Websocket connection. The control device sends the response voice to the call center device through a Websocket connection, so that the call center device sends the response voice to the user's client. After the control device obtains the first text corresponding to the first voice message from the voice conversion device through a Websocket connection, it further includes: The control device obtains the number of characters included in the first text and the time information of the historical response voice that the control device sent to the call center device last time. The time information includes the first sending moment and the first duration of the historical response voice. The first duration is the playback duration of the historical response voice. The control device generates an interruption instruction according to the number of characters included in the first text and the time information.

2. The method according to claim 1, wherein Generating an interruption instruction according to the number of characters included in the first text and the time information includes: Determining whether the call center device is playing the historical response voice according to the current moment, the first sending moment, and the first duration of the historical response voice. If so, when the number of characters included in the first text is greater than or equal to a preset threshold, the interruption instruction is generated.

3. The method according to any one of claims 1-2, characterized in that The control device obtaining the response text corresponding to the first text from the response device includes: The control device sends the first text to the response device through an HTTP connection. The control device receives the response text from the response device through an HTTP connection.

4. The method according to any one of claims 1 to 3, characterized in that, Before the control device receives the user's first voice message from the call center device through a Websocket connection, it includes: The control device receives the Websocket connection establishment request sent by the call center device. The control device establishes a Websocket connection with the call center device according to the Websocket connection establishment request.

5. The method according to any one of claims 1 to 3, characterized in that, Before the control device obtains the first text corresponding to the first voice message from the voice conversion device through a Websocket connection, it further includes: After the control device determines that a call connection is established between the client and the call center device, it sends a Websocket connection establishment request to the voice conversion device. The control device receives a Websocket connection establishment response corresponding to the Websocket connection establishment request sent by the voice conversion device; The control device establishes a Websocket connection with the voice conversion device according to the Websocket connection establishment response.

6. The method according to any one of claims 1 to 3, characterized in that Before the control device obtains the response voice corresponding to the response text from the speech synthesis device through the Websocket connection, it includes: After the control device determines that a call connection is established between the client and the call center device, it sends a Websocket connection establishment request to the speech synthesis device; The control device receives a Websocket connection establishment response corresponding to the Websocket connection establishment request sent by the speech synthesis device; The control device establishes a Websocket connection with the speech synthesis device according to the Websocket connection establishment response.

7. A voice system, characterized in that, It includes a control device, a call center device, a voice conversion device, a response device, and a speech synthesis device, where: The call center device is used to send the first voice information of the user to the control device through a Websocket connection; The control device is used to send the first voice information to the voice conversion device through a Websocket connection; The voice conversion device is used to convert the first voice information into a first text and send the first text to the control device through a Websocket connection; The control device is further used to send the first text to the response device; The response device is used to determine a response text corresponding to the first text and send the response text to the control device; The control device is further used to send the response text to the speech synthesis device through a Websocket connection; The speech synthesis device is used to convert the response text into a response voice and send the response voice to the control device through a Websocket connection; The control device is further used to send the response voice to the call center device through a Websocket connection; The call center device is further used to send the response voice to the client of the user; After the control device obtains the first text, it is further used to obtain the number of characters included in the first text, and the time information of the historical response voice that the control device sent to the call center device last time. The time information includes a first sending moment and a first duration of the historical response voice, and the first duration is the playing duration of the historical response voice; according to the number of characters included in the first text and the time information, a interruption instruction is generated.

8. The voice system according to claim 7, characterized in that, The control device is further used to execute the method according to any one of claims 2-6.

9. A voice processing device, characterized in that, Applied to a voice system, the voice system includes a control device, a call center device, a voice conversion device, a response device, and a speech synthesis device. The voice processing device includes a receiving module, a first obtaining module, a second obtaining module, and a sending module, where: The receiving module is configured to receive the first voice information of the user from the call center device through a Websocket connection, and the control device is connected to the call center device through a Websocket connection; The first obtaining module is configured to obtain the first text corresponding to the first voice information from the voice conversion device through a Websocket connection, and the control device is connected to the voice conversion device through a Websocket connection; The second obtaining module is configured to obtain the response text corresponding to the first text from the response device, and obtain the response voice corresponding to the response text from the voice synthesis device through a Websocket connection; The sending module is configured to send the response voice to the call center device through a Websocket connection, so that the call center device sends the response voice to the client of the user; The first obtaining module is further configured to obtain the number of characters included in the first text, and the time information of the historical response voice last sent by the control device to the call center device. The time information includes the first sending moment and the first duration of the historical response voice, and the first duration is the playing duration of the historical response voice; and generate an interruption instruction according to the number of characters included in the first text and the time information.

10. A voice processing device, characterized in that, The voice processing device includes: a memory, a processor, and a voice processing program stored on the memory and executable on the processor. When the voice processing program is executed by the processor, the steps of the voice processing method according to any one of claims 1 to 6 are implemented.

11. A computer-readable storage medium, characterized in that, A voice processing program is stored on the computer-readable storage medium. When the voice processing program is executed by a processor, the steps of the voice processing method according to any one of claims 1 to 6 are implemented.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Intelligent quality inspection method and system for voice communication

    CN111128241A

  • Intelligent calling method, device, equipment and medium

    CN111666380A