A method and system for realizing real-time voice human-computer dialogue

By establishing a long connection between the user and the agent server and the voice server, and using the set byte threshold and heartbeat packet mechanism, the problem of voice stuttering and disconnection in human-computer interaction is solved, and the call quality and user experience are improved.

CN114242073BActive Publication Date: 2025-07-04LINGXI TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111506364.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-07-04
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

In large-scale human-computer interaction scenarios, voice stream playback stuttering and network signals disconnection during the voice interaction between users and agents, resulting in poor call quality and low work efficiency.

Method used

By establishing a long connection between the user server side and the agent server side and the voice server, and using the set byte threshold to read voice information, combined with the heartbeat packet mechanism to detect the connection status in real time, ensuring stable transmission and reconstruction of information, avoiding lags and disconnection.

Benefits of technology

It achieves smooth communication between users and agents, improves call quality and user experience, and improves communication efficiency and work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114242073B_ABST
    Figure CN114242073B_ABST
Patent Text Reader

Abstract

Some embodiments of the present application provide a method and system for realizing real-time voice human-machine dialogue, including: a user server, a voice server, and an agent server. Among them, the voice server establishes a first long connection with the user server and a second long connection with the agent server; the user server can read the user's voice according to a set byte threshold to obtain voice information; after sending the voice information to the voice server through the first long connection, it is then forwarded to the agent server; the agent server can read the customer service voice according to a set byte threshold to obtain customer service voice information; after sending the customer service voice information to the voice server through the second long connection, it is then forwarded to the user server. This embodiment realizes the real-time forwarding and interaction of voices, avoids the situation of jamming and communication disconnection during the interaction between the user and the agent, and improves the call quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent interaction technologies, and more specifically, to a method and system for realizing real-time voice human-machine conversation. Background Art

[0002] With the development of the field of intelligent interaction, human-machine conversation technology has gradually been applied to various interaction scenarios.

[0003] Currently, in a relatively large-scale interaction field, the number of users is large. When users directly conduct voice interaction with seat personnel, during the forwarding process of the voice stream, situations such as voice stream playback jamming and network signal disconnection are likely to occur, resulting in poor call quality and low work efficiency.

[0004] Therefore, how to improve the call quality of human-machine conversation has become a technical problem that urgently needs to be solved. Summary of the Invention

[0005] Some embodiments of the present application aim to provide a method and system for realizing real-time voice human-machine conversation. Through the technical solutions of the embodiments of the present application, it is possible to avoid situations of jamming and network disconnection during the interaction between users and seat personnel, thereby improving the call quality of both parties.

[0006] In a first aspect, some embodiments of the present application provide a method for realizing real-time voice human-machine conversation, which is applied to a user server side and includes: answering a user's voice; reading the voice according to a set byte threshold to obtain voice information; sending the voice information to a voice server, where the voice server forwards the voice information to a seat server side so that the seat server side can obtain the voice; receiving the customer service voice of the seat server side sent by the voice server.

[0007] Some embodiments of the present application forward the user's voice read by the user server side according to a set byte threshold to the seat server side through a voice server, realizing real-time forwarding of the user voice stream. At the same time, reading the user's voice according to the set byte threshold can ensure smooth communication between users and seat personnel and avoid the situation of voice signal jamming. In addition, some embodiments of the present application also use the voice server to communicate with the seat side, further improving the call quality and user experience.

[0008] In some embodiments, the byte threshold is set by the following method: based on the server configuration parameters of the user server side, set the buffer parameters of the user server side, and use the buffer parameters as the byte threshold.

[0009] In some embodiments of the present application, by setting the size of the buffer parameter of the server as a byte threshold for different servers, the fluency and clarity of the voice information received at the agent server side can be ensured.

[0010] In some embodiments, before reading the voice according to the set byte threshold to obtain voice information, the method further includes: sending a request to establish a long connection to the voice server; receiving a successful establishment flag sent by the voice server, so that the user server side establishes a long connection with the voice server.

[0011] In some embodiments of the present application, establishing a long connection between the user server side and the voice server realizes the stability of the connection between the servers and can guarantee the call quality.

[0012] In some embodiments, during the period when the user server side establishes a long connection with the voice server, the method further includes: sending a heartbeat packet to the voice server according to a set time period; if a connection normal flag sent by the voice server is received within the set time period, it is confirmed that the connection between the user server side and the voice server is normal; if a connection normal flag sent by the voice server is not received within the set time period, a request to re - establish a long connection is sent to the voice server and a successful establishment flag sent by the voice server is received, so that the user server side re - establishes a long connection with the voice server.

[0013] In some embodiments of the present application, the user server side sends a heartbeat packet to the voice server at regular time intervals to detect the connection status between the two in real time, avoiding connection interruption during the interaction process, and at the same time, the connection can be re - established in real time for the interrupted situation.

[0014] In a second aspect, some embodiments of the present application provide a method for realizing real - time voice human - machine dialogue, which is applied to the agent server side and includes: receiving the voice information of the user sent by the voice server; obtaining the customer service voice corresponding to the voice information; reading the customer service voice according to the set byte threshold to obtain customer service voice information; sending the customer service voice information to the voice server, where the voice server forwards the customer service voice information to the user server side so that the user server side obtains the customer service voice.

[0015] In some embodiments of the present application, the customer service voice of the seat personnel read by the seat server end through a set byte threshold is forwarded to the user server end through the voice server, realizing the real-time forwarding of the customer service voice stream. At the same time, reading the voice of the seat personnel according to the set byte threshold can ensure smooth communication between the seat personnel and the user, avoid the situation of voice signal jamming, improve the call quality and user experience, and thus improve the communication efficiency.

[0016] In some embodiments, the byte threshold is set by the following method: based on the server configuration parameters of the seat server end, the buffer parameters of the seat server end are set, and the buffer parameters are used as the byte threshold.

[0017] In some embodiments of the present application, by setting different buffer parameters for different seat servers as the size of the byte threshold, the smoothness and clarity of the customer service voice information received by the user server end can be ensured, thus guaranteeing the communication effect.

[0018] In some embodiments, before receiving the voice information of the user sent by the receiving voice server, the method further includes: sending a request to establish a long connection to the voice server; receiving the connection success flag sent by the voice server, so that the seat server end establishes a long connection with the voice server.

[0019] In some embodiments of the present application, by establishing a long connection between the seat server end and the voice server, the stability of the connection between the seat server end and the voice server is realized, which can guarantee the call quality.

[0020] In some embodiments, during the period when the seat server end establishes a long connection with the voice server, the method further includes: sending a heartbeat packet to the voice server according to a set time period; if the connection normal flag sent by the voice server is received within the set time period, it is confirmed that the connection between the seat server end and the voice server is normal; if the connection normal flag sent by the voice server is not received within the set time period, a request to establish a long connection is resent to the voice server and the connection success flag sent by the voice server is received, so that the seat server end re - establishes a long connection with the voice server.

[0021] In some embodiments of the present application, the seat server end sends a heartbeat packet to the voice server every certain time period to detect the connection situation between the two in real time, avoid the situation of connection interruption during the voice information interaction process, and at the same time, for the interrupted situation, it can also realize real - time re - establishment of the connection.

[0022] In some embodiments, before obtaining the customer service voice corresponding to the voice information, the method further includes: allocating corresponding seat personnel at the agent server end according to the voice transfer rate; and allocating the voice information to the corresponding seat personnel.

[0023] In some embodiments of the present application, sufficient seat personnel are allocated according to the voice transfer rate, and users do not need to queue up when interacting with seat personnel, improving the user experience and communication efficiency.

[0024] In a third aspect, some embodiments of the present application provide a method for implementing real-time voice human-machine dialogue, which is applied to a voice server and includes: receiving a long connection establishment request sent by a first server end, where the first server end is at least used to answer the user's voice or at least used to obtain customer service voice information according to the user's voice information; sending a connection establishment success flag to the first server end so that the first server end establishes a first long connection with the voice server; and sending the information from the first server end to a second server end at least through the first long connection.

[0025] The voice server in some embodiments of the present application realizes real-time forwarding of voice information by establishing a long connection with the first server end. Through the long connection, real-time interaction between the first server end and the second server end can be achieved, and voice information can be transmitted stably, clearly, and smoothly, improving the call efficiency and quality.

[0026] In some embodiments, the first server end is a user server end, and the second server end is an agent server end; the method further includes: establishing a second long connection between the agent server end and the voice server; where the sending the information from the first server end to the second server end at least through the first long connection includes: sending the voice of the user from the user server end to the agent server end through the first long connection and the second long connection so that the agent server end can obtain the voice.

[0027] In some embodiments of the present application, the first server end is set as the user server end, and the second server end is set as the agent server end. The two respectively establish a first long connection and a second long connection with the voice server to ensure that the voice received by the user server end is transmitted to the agent server end in real time and stably, guaranteeing the call communication effect.

[0028] In some embodiments, the first server end is an agent server end, and the second server end is a user server end; the method further includes: establishing a second long connection between the user server end and the voice server; wherein, sending the information from the first server end to the second server end through at least the first long connection includes: sending the customer service voice from the agent server end to the user server end through the first long connection and the second long connection, so that the user server end can obtain the customer service voice.

[0029] In some embodiments of the present application, the first server end is set as the agent server end, and the second server end is set as the user server end. The two establish a first long connection and a second long connection with the voice server respectively, so that the customer service voice received by the agent server end can be transmitted to the user server end in real time and stably, ensuring the call communication effect.

[0030] In some embodiments, during the period when the first server end establishes a first long connection with the voice server, the method further includes: if a heartbeat packet sent by the first server end is received within a set time period, sending a connection normal flag to the first server end, where the connection normal flag is used to indicate that the network connection between the voice server and the first server end is normal; if the heartbeat packet sent by the first server end is not received within the set time period, receiving a long connection establishment request re-sent by the first server end and sending a connection establishment success flag to the first server end, so that the voice server and the first server end re-establish a long connection.

[0031] In some embodiments of the present application, the connection quality of the long connection between the first server end and the voice server is detected by whether a heartbeat packet is received within a specified time period, avoiding abnormal situations of disconnection. When a disconnection occurs, it can be detected in the first time, and a long connection can be re-established in time, so that the voice call can be forwarded in real time, providing an effective guarantee for the call quality.

[0032] In a fourth aspect, some embodiments of the present application provide a user server end, including: a listening module configured to answer the user's voice; a reading module configured to read the voice according to a set byte threshold to obtain voice information; a sending module configured to send the voice information to a voice server, where the voice server forwards the voice information to an agent server end, so that the agent server end can obtain the voice; an information receiving module configured to receive the customer service voice of the agent server end sent by the voice server.

[0033] Fifth aspect, some embodiments of the present application provide a seat server side, including: a receiving module configured to receive voice information of a user sent by a voice server; an obtaining module configured to obtain a customer service voice corresponding to the voice information; a voice reading module configured to read the customer service voice according to a set byte threshold to obtain customer service voice information; an information sending module configured to send the customer service voice information to the voice server, where the voice server forwards the customer service voice information to a user server side so that the user server side obtains the customer service voice.

[0034] Sixth aspect, some embodiments of the present application provide a voice server, including: a request receiving module configured to receive a long connection establishment request sent by a first server side, where the first server side is at least used to answer a user's voice or at least used to obtain customer service voice information according to the user's voice information; a request confirmation module configured to send a successful establishment flag to the first server side so that the first server side establishes a first long connection with the voice server; an information forwarding module configured to send information from the first server side to a second server side at least through the first long connection.

[0035] Seventh aspect, some embodiments of the present application provide a system for realizing real-time voice human-machine dialogue, including: a user server side, a voice server, and a seat server side, where the voice server establishes a first long connection with the user server side and a second long connection with the seat server side; the user server side is configured to: answer a user's voice; read the voice according to a set byte threshold to obtain voice information; send the voice information to the voice server through the first long connection, where the voice server forwards the voice information to the seat server side so that the seat server side obtains the voice; receive the customer service voice of the seat server side sent by the voice server; the seat server side is configured to: receive voice information of a user sent by the voice server; obtain a customer service voice corresponding to the voice information; read the customer service voice according to a set byte threshold to obtain customer service voice information; send the customer service voice information to the voice server through the second long connection, where the voice server forwards the customer service voice information to the user server side so that the user server side obtains the customer service voice.

[0036] Eighth aspect, some embodiments of the present application provide an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor, where when the processor executes the program, the method described in any embodiment of the first aspect, the second aspect, or the third aspect can be implemented.

[0037] In a ninth aspect, some embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the methods described in any of the embodiments of the first aspect, the second aspect, or the third aspect can be implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] To more clearly illustrate the technical solutions of some embodiments of the present application, the drawings required for use in some embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 It is a structural diagram of a system for implementing real-time voice human-machine dialogue provided by some embodiments of the present application;

[0040] Figure 2 It is one of the flowcharts of a method for implementing real-time voice human-machine dialogue provided by some embodiments of the present application;

[0041] Figure 3 It is the second flowchart of a method for implementing real-time voice human-machine dialogue provided by some embodiments of the present application;

[0042] Figure 4 It is the third flowchart of a method for implementing real-time voice human-machine dialogue provided by some embodiments of the present application;

[0043] Figure 5 It is an interaction flowchart of a user server 100, an agent server 200, and a voice server 300 provided by some embodiments of the present application;

[0044] Figure 6 It is a block diagram of the composition of a user server provided by some embodiments of the present application;

[0045] Figure 7 It is a block diagram of the composition of an agent server provided by some embodiments of the present application;

[0046] Figure 8 It is a block diagram of the composition of a voice server provided by some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] Next, the technical solutions in some embodiments of the present application will be described in conjunction with the drawings in some embodiments of the present application.

[0048] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. At the same time, in the description of the present application, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0049] In the related art examples, due to the convenience of human-computer interaction, human-computer interaction is involved in the service field (such as, bank systems, insurance systems, after-sales systems, etc.). When the existing technology directly communicates between the user and the agent, when the connection is unstable, problems such as lag or disconnection are likely to occur.

[0050] In view of this, some embodiments of the present application provide a method and system for realizing real-time voice human-machine dialogue, which avoid the occurrence of lag and network disconnection during the interaction between the user and the agent, improve the call quality while enhancing the user experience, and ensure work efficiency.

[0051] The following exemplarily introduces a method for realizing real-time voice human-machine dialogue provided by some embodiments of the present application.

[0052] As Figure 1 shown, some embodiments of the present application provide a structural diagram of a system for realizing real-time voice human-machine dialogue. Figure 1 The structural diagram of the forwarding system includes a user server 100, an agent server 200, and a voice server 300.

[0053] Compared with the technical solutions with many technical defects brought about by the direct communication between the user server 100 and the agent server 200 in the related art, some embodiments of the present application introduce a voice server. Based on meeting the information forwarding requirements, the voice server can also detect the long connection status in real time, effectively avoiding the lag problem existing in the prior art during the call.

[0054] Figure 1 The voice server 300 establishes a first long connection with the user server 100 and a second long connection with the agent server 200. That is to say, in some embodiments of the present application, both the user server 100 and the agent server 200 can bidirectionally transmit data information to the voice server 300 through long connections.

[0055] Figure 1 The user server 100 can send a request to establish a long connection, user voice, or a heartbeat packet for detecting whether the network is disconnected to the voice server 300. Figure 1The seat server 200 can send a request to establish a long connection, customer service voice, or a heartbeat packet for detecting whether the network is disconnected to the voice server 300. Correspondingly, the voice server 300 can send an identifier indicating whether the connection is successfully established to the user server 100 and the seat server 200 respectively, and can also forward the user voice to the seat server 200, or forward the customer service voice to the user server 100.

[0056] It should be noted that the user server 100 can be deployed on a terminal device, enabling the terminal device to have the function of acquiring user voice. The seat server 200 can be deployed on a seat-side device, enabling the seat-side device to have the function of acquiring the voice of the seat personnel. Both the terminal device and the seat-side device can establish a long connection with the voice server 300 through wireless network devices and wired network devices to achieve data transmission.

[0057] It can be understood that the user server 100 and the seat server 200 can be deployed either on mobile terminal devices or on non-portable computer terminals. The voice server 300 can also be deployed on non-portable computers or portable terminals, etc. The present application does not limit the specific device types.

[0058] The following is an exemplary elaboration Figure 1 of the related functions of each unit.

[0059] In some embodiments of the present application, the voice server 300 establishes a first long connection with the user server 100 and a second long connection with the seat server 200.

[0060] In some embodiments of the present application, the user server 100 is at least configured to: answer the user's voice; read the voice according to a set byte threshold to obtain voice information; send the voice information to the voice server through the first long connection, where the voice server forwards the voice information to the seat server so that the seat server can obtain the voice; receive the customer service voice of the seat server sent by the voice server.

[0061] In some embodiments of the present application, the seat server 200 is at least configured to: receive the voice information of the user sent by the voice server; obtain the corresponding customer service voice for the voice information; read the customer service voice according to a set byte threshold to obtain customer service voice information; send the customer service voice information to the voice server through the second long connection, where the voice server forwards the customer service voice information to the user server so that the user server can obtain the customer service voice.

[0062] The following is a specific elaboration in combination with the attached Figure 2 Specifically elaborateFigure 1 The implementation process of the method for implementing real-time voice human-machine dialogue executed by the user server 100.

[0063] Please refer to the appendix Figure 2 For some embodiments of the present application, the method for implementing real-time voice human-machine dialogue executed by the user server 100 may include: S210, answering the user's voice; S220,

[0064] Reading the voice according to the set byte threshold to obtain voice information; S230, sending the voice information to the voice server, where the voice server forwards the voice information to the agent server end so that the agent server end can obtain the voice; S240, receiving the customer service voice of the agent server end sent by the voice server.

[0065] The above process will be described exemplarily below.

[0066] To ensure the transmission quality of the voice information of the user server end, in some embodiments of the present application, when executing Figure 2 step S220, it is necessary to set the buffer parameter of the user server end based on the server configuration parameter of the user server end, and use the buffer parameter as the byte threshold.

[0067] For example, in some embodiments of the present application, according to the basic configuration performance of the server, by pre-manually testing the clarity, fluency, and latency performance of voice transmission, continuously adjust the size of the cache block (that is, the buffer parameter) of the user server end until the voice can be clearly, smoothly, and meet the latency requirements for real-time transmission to the agent end. As an example of the present application, the audio sampling frequency of the user server end can be set to 8000 Hz, and the cache block size can be set to 1280 frames / time, that is, 1280 frames of audio data in the user's voice are read each time for transmission at the set audio sampling frequency.

[0068] In some embodiments of the present application, to ensure the stable transmission of voice information, before executing S210, the method further includes sending a request to establish a long connection to the voice server; receiving the success flag sent by the voice server to establish a long connection between the user server end and the voice server.

[0069] For example, to ensure the quality of the long connection, after the user server end sends a request to the voice server, the voice server will feedback the identification information of whether the connection is successfully established to the user server end to ensure that the two successfully establish a connection.

[0070] In some embodiments of the present application, to ensure that there is no network disconnection during the interaction between the user and the seat personnel, during the period when the long connection is established between the user server and the voice server, the method further includes: sending a heartbeat packet to the voice server at a set time period; if the connection normal flag sent by the voice server is received within the set time period, it is confirmed that the connection between the user server and the voice server is normal; if the connection normal flag sent by the voice server is not received within the set time period, a long connection establishment request is resent to the voice server and the establishment success flag sent by the voice server is received, so that the user server re - establishes a long connection with the voice server.

[0071] For example, since the interaction duration between the user and the seat personnel is uncertain, the connection between the user server and the voice server may be disconnected during a long call. Therefore, it is necessary to send a heartbeat packet to the voice server at regular intervals to detect the connection quality and avoid disconnection. When a disconnection occurs, the connection can be re - established in time to ensure the continuation of the call.

[0072] The following combines the attached Figure 3 Specifically elaborate Figure 1 the implementation process of the method for realizing real - time voice human - machine dialogue executed by the seat server 200 in the middle.

[0073] Please refer to the attached Figure 3 In some embodiments of the present application, the method for realizing real - time voice human - machine dialogue executed by the seat server 200 may include: S310, receiving the voice information of the user sent by the voice server; S320, obtaining the customer service voice corresponding to the voice information; S330, reading the customer service voice according to the set byte threshold to obtain the customer service voice information; S340, sending the customer service voice information to the voice server, where the voice server forwards the customer service voice information to the user server so that the user server can obtain the customer service voice.

[0074] The above process is described below by way of example.

[0075] To ensure the voice call quality, in some embodiments of the present application, when the seat server 200 executes S330, the buffer parameter of the seat server can be set based on the server configuration parameters of the seat server, and the buffer parameter is used as the byte threshold.

[0076] For example, in some embodiments of the present application, according to the parameters of the basic configuration performance of the agent server, by manually testing the clarity, fluency, and latency performance of voice transmission in advance, the size of the cache block (i.e., the buffer parameter) of the agent server is continuously adjusted until the voice can be clearly, smoothly, and in real-time transmitted to the user side to meet the latency requirements. As an example of the present application, the audio sampling frequency of the agent server can be set to 8000 Hz, and the cache block size can be set to 1280 frames / time, that is, 1280 frames of audio data in the voice of the agent are read and transmitted each time at the set audio sampling frequency.

[0077] In some embodiments of the present application, in order to ensure the stable transmission of voice information, before executing S310, the method further includes sending a request to establish a long connection to the voice server; receiving the establishment success flag sent by the voice server, so that the agent server establishes a long connection with the voice server.

[0078] In some embodiments of the present application, in order to avoid abnormal situations during the call or be able to initiate remedial measures in a timely manner when an abnormality occurs, during the period when the agent server establishes a long connection with the voice server, the method further includes: sending a heartbeat packet to the voice server according to a set time period; if the connection normal flag sent by the voice server is received within the set time period, it is confirmed that the connection between the agent server and the voice server is normal; if the connection normal flag sent by the voice server is not received within the set time period, a request to re-establish a long connection is sent to the voice server and the establishment success flag sent by the voice server is received, so that the agent server re-establishes a long connection with the voice server.

[0079] In some embodiments of the present application, in order to improve the user experience, before executing S320, the method further includes: allocating corresponding agent personnel at the agent server according to the voice transfer rate; and allocating the voice information to the corresponding agent personnel.

[0080] For example, according to the number of robot agents that can be undertaken in different interaction scenarios, combined with the incoming call connection rate and the rate of transferring to agent personnel, an appropriate number of agent personnel is allocated, so that users do not need to queue up when they need agent services, improving the call efficiency. Among them, the number of agent personnel is obtained by the following method: the number of agent personnel = (the number of robot seats * the average number of incoming calls per hour * the incoming call connection rate * the average call duration per call) / the set call duration per hour for the seat. Among them, the incoming call connection rate is the ratio of the number of connected calls to the total number of user incoming calls. The rate of transferring to agent personnel is the ratio of the number of calls transferred to agent personnel to the number of connected calls.

[0081] The following is combined with the attachedFigure 4 Specific elaboration Figure 1 The implementation process of the method for implementing real-time voice human-machine dialogue executed by the voice server 300 in

[0082] Please refer to the attached Figure 4 For some embodiments of the present application, the method for implementing real-time voice human-machine dialogue executed by the voice server 300 may include: S410, receiving a long connection establishment request sent by the first server end, where the first server end is at least used to answer the user's voice or at least used to obtain customer service voice information according to the user's voice information; S420, sending a successful establishment flag to the first server end so that the first server end establishes a first long connection with the voice server; S430, sending the information from the first server end to the second server end at least through the first long connection.

[0083] It should be noted that the first server end may be a user server end or a seat server end. When the first server end is a user server end, the second server end is a seat server end. When the first server end is a seat server end, the second server end is a user server end.

[0084] The above process is elaborated below by way of example.

[0085] In some embodiments of the present application, the first server end is set as the user server end and the second server end is set as the seat server end; the above method further includes: establishing a second long connection between the seat server end and the voice server; where S430 executed by the voice server 300 includes: sending the voice of the user from the user server end to the seat server end through the first long connection and the second long connection so that the seat server end obtains the voice.

[0086] For example, the user server end establishes a first long connection with the voice server, and the seat server end establishes a second long connection with the voice server. Through the long connection, the user's voice can be forwarded to the seat server end in real time for the seat personnel to obtain the user's needs.

[0087] In some other embodiments of the present application, the first server end is set as the seat server end and the second server end is set as the user server end; the method further includes: establishing a second long connection between the user server end and the voice server; where S430 executed by the voice server 300 may further include: sending the customer service voice from the seat server end to the user server end through the first long connection and the second long connection so that the user server end obtains the customer service voice.

[0088] For example, the agent server establishes a first long connection with the voice server, and the user server establishes a second long connection with the voice server. Through the long connection, the customer service voice of the agent can be forwarded to the user server in real time for the user to answer the voice message of the customer service reply.

[0089] In some embodiments of the present application, in order to enable the voice server to effectively realize the real-time forwarding of the voice of the first server and avoid interruption, during the period of establishing the first long connection between the first server and the voice server, the method further includes: if a heartbeat packet sent by the first server is received within a set time period, a connection normal flag is sent to the first server, where the connection normal flag is used to indicate that the network connection between the voice server and the first server is normal; if the heartbeat packet sent by the first server is not received within the set time period, the voice server receives the re-sent long connection establishment request sent by the first server and sends a connection success flag to the first server, so that the voice server and the first server re-establish a long connection.

[0090] It should be noted that the above first server can be a user server or an agent server, and both can realize the problem of real-time detection of the long connection quality.

[0091] The following Figure 5 exemplarily elaborates Figure 1 the interaction process among the user server 100, the agent server 200, and the voice server 300, and improves the call quality by realizing the real-time forwarding of voice.

[0092] S1. Both the user server 100 and the agent server 200 send requests to establish long connections to the voice server 300.

[0093] For example, while answering a user call, Figure 5 the voice client (as a specific example of the user server 100) and the voice customer service client (as a specific example of the agent server 200) in

[0094] It should be noted that both the user server 100 and the agent server 200 are provided with virtual sound cards to store the voices of users or agents.

[0095] S2. The voice server 300 sends connection success flags to both the user server 100 and the agent server 200 to confirm the successful establishment of the communication connection.

[0096] For example, in some embodiments of the present application, during the establishment of a long connection between the user server 100 and the voice server 300, the user server 100 sends heartbeat packets to the voice server 300 at a set time period (for example, the time period can be 3 ms or 5 cm, etc., and a suitable period can be set according to the actual situation, which is not limited here). If the user server 100 receives a connection normal flag sent by the voice server 300 within the set time period, it is confirmed that the connection between the user server 100 and the voice server 300 is normal. If the user server 100 does not receive the connection normal flag sent by the voice server 300 within the set time period, the user server 100 resends a long connection establishment request to the voice server 300 and receives a connection establishment success flag sent by the voice server 300 to re - establish a long connection between the user server 100 and the voice server 300.

[0097] It should be understood that the agent server 200 also sends heartbeat packets to the voice server 300 at a set time period to detect the long - connection quality between the agent server 200 and the voice server 300. The specific detection process is similar to the detection process of the above - mentioned user server 100 and the voice server 300. To avoid repetition, the detailed description is omitted here.

[0098] S3. The user server 100 answers the user's voice.

[0099] For example, in some embodiments of the present application, the user server 100 stores the user's voice through an internal virtual sound card.

[0100] S4. The user server 100 reads the voice according to a set byte threshold to obtain voice information.

[0101] For example, in order to ensure clear text and smooth speech rate of the obtained voice information, in some embodiments of the present application, a single - read byte threshold can be set to read the user's voice on the virtual sound card to obtain the user's voice information. As an example of the present application, in an environment where the audio sampling frequency is set to 8000 Hz and the buffer block size of PyAudio is 1280 frames per time, the user's voice is read.

[0102] S5. The user server 100 sends the voice information in S4 to the voice server 300 through the long - connection channel.

[0103] S6. The voice server 300 forwards the voice information of the user server 100 to the agent server 200 through the long - connection channel.

[0104] After the seat server 200 receives the voice message, the seat personnel will reply to the voice message, and the seat server 200 obtains the customer service voice.

[0105] For example, in some embodiments of the present application, the virtual sound card provided in the seat server 200 stores the voice of the seat personnel, that is, the customer service voice.

[0106] S8, the seat server 200 reads the customer service voice according to the set byte threshold to obtain the customer service voice information.

[0107] For example, in order to ensure the clarity of the text and the smoothness of the speech rate of the obtained customer service voice information, in some embodiments of the present application, the byte threshold for single reading can be set to read the customer service voice on the virtual sound card to obtain the customer service voice information. As an example of the present application, the audio sampling frequency can be set to 8000Hz, and the customer service voice can be read in an environment where the buffer block size of PyAudio is 1280 frames / time.

[0108] S9, the seat server 200 sends the customer service voice information in S8 to the voice server 300 through the long connection channel.

[0109] S10, the voice server 300 forwards the customer service voice information of the seat server 200 to the user server 100 through the long connection channel.

[0110] It can be understood that if the user needs to have multiple rounds of interaction with the seat personnel, the implementation process of each round of interaction is the same as that of S1~S10. To avoid repetition, it will not be elaborated here.

[0111] S11, the user server 100 monitors that the user hangs up the phone.

[0112] S12, the user server 100 or the seat server 200 sends a request to disconnect the long connection to the voice server 300.

[0113] S13, the voice server 300 disconnects the long connection with both the user server 100 and the seat server 200.

[0114] It should be noted that in some embodiments of the present application, as long as any one of the user server 100 and the seat server 200 sends a request to disconnect the long connection to the voice server 300, the user server 100 and the seat server 200 can respectively disconnect the long connection with the voice server 300.

[0115] It can be understood that in some embodiments of the present application, a long connection is established with the voice server 300 when the user server 100 and the agent server 200 interact, and the long connection is disconnected when the interaction ends. This can effectively avoid the problem of voice jamming and poor connection quality caused by the user server 100 and the agent server 200 being connected to the voice server 300 all the time.

[0116] In addition, in some other embodiments of the present application, since the memory of the voice server 300 that can carry forwarded data is limited, multiple voice servers 300 can be set according to the actual application scenario requirements. When the memory of a single voice server 300 reaches the maximum upper limit, the user server 100 and the seat server 200 can automatically match other voice servers 300 that have not reached the memory upper limit and establish a long connection, avoiding abnormal situations of call jams and ensuring call quality and communication efficiency.

[0117] Please refer to Figure 6 , Figure 6 The following is a block diagram showing a composition of a user server provided by some embodiments of the present application. It should be understood that the user server is similar to the above-mentioned Figure 2 Corresponding to the method embodiment, each step involved in the above method embodiment can be executed. The specific functions of the user server can be found in the description above. To avoid repetition, the detailed description is appropriately omitted here.

[0118] Figure 6 The user server end includes at least one software function module that can be stored in the memory or solidified in the user server end in the form of software or firmware, and the user server end includes: a monitoring module 610, configured to answer the user's voice; a reading module 620, configured to read the voice according to a set byte threshold and obtain voice information; a sending module 630, configured to send the voice information to a voice server, wherein the voice server forwards the voice information to the agent server end so that the agent server end obtains the voice; an information receiving module 640, configured to receive the customer service voice of the agent server end sent by the voice server.

[0119] Please refer to Figure 7 , Figure 7 The following is a block diagram showing the composition of a seat server provided by some embodiments of the present application. It should be understood that the seat server is similar to the above-mentioned Figure 3 Corresponding to the method embodiment, each step involved in the above method embodiment can be executed. The specific functions of the agent server can refer to the description above. To avoid repetition, the detailed description is appropriately omitted here.

[0120] Figure 7The agent server side includes at least one software function module that can be stored in the memory in the form of software or firmware or solidified in the agent server side, and the agent server side includes: a receiving module 710, configured to receive the user's voice information sent by the voice server; an acquisition module 720, configured to acquire the customer service voice corresponding to the voice information; a voice reading module 730, configured to read the customer service voice according to a set byte threshold and acquire the customer service voice information; an information sending module 740, configured to send the customer service voice information to the voice server, wherein the voice server forwards the customer service voice information to the user server side, so that the user server side acquires the customer service voice.

[0121] Please refer to Figure 8 , Figure 8 FIG. 1 shows a block diagram of a voice server provided in some embodiments of the present application. It should be understood that the voice server is similar to the above-mentioned Figure 4 Corresponding to the method embodiment, each step involved in the above method embodiment can be executed. The specific functions of the voice server can be found in the description above. To avoid repetition, the detailed description is appropriately omitted here.

[0122] Figure 8 The voice server includes at least one software function module that can be stored in a memory in the form of software or firmware or solidified in the voice server, and the voice server includes: a request receiving module 810, configured to receive a request for establishing a long connection sent by a first server end, wherein the first server end is at least used to answer the user's voice, or at least used to obtain customer service voice information based on the user's voice information; a request confirmation module 820, configured to send an establishment success mark to the first server end, so that the first server end establishes a first long connection with the voice server; an information forwarding module 830, configured to send information from the first server end to the second server end at least through the first long connection.

[0123] Some embodiments of the present application further provide an electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, Figure 2 , Figure 3 or Figure 4 The method described in any embodiment of the present invention.

[0124] Some embodiments of the present application also provide a computer-readable storage medium having a computer program stored thereon, which can implement Figure 2 , Figure 3 or Figure 4 The method described in any embodiment of the present invention.

[0125] The above are only examples of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application. It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0126] As described above, this is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0127] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

Claims

1. A method for realizing real-time voice human-machine dialogue, characterized in that, Applied to the user server side, including: Answer the user's voice; Read the voice according to the set byte threshold to obtain voice information; Send the voice information to the voice server, where the voice server forwards the voice information to the agent server side so that the agent server side can obtain the voice; the number of voice servers is multiple, and when the memory of a single voice server reaches the upper limit value, it will automatically match to other voice servers that have not reached the memory upper limit; Receive the customer service voice of the agent server side sent by the voice server; where the agent server side configures the corresponding number of agent personnel according to the number of robot agents undertaken in different interaction scenarios, combined with the incoming call connection rate and the transfer to agent manual rate; Before reading the voice according to the set byte threshold to obtain voice information, the method further includes: sending a request to establish a long connection to the voice server; receiving the connection success flag sent by the voice server to establish a long connection between the user server side and the voice server; When it is detected that the user hangs up the phone, either the user server side or the agent server side sends a request to disconnect the long connection to the voice server to facilitate disconnecting the connection with the voice server; Set the byte threshold: Based on the server configuration parameters of the user server side and the clarity, fluency, and latency performance of manual test voice transmission, set the buffer parameters of the user server side and use the buffer parameters as the byte threshold.

2. The method according to claim 1, wherein During the period when the user server side establishes a long connection with the voice server, the method further includes: Send a heartbeat packet to the voice server according to the set time period; If the connection normal flag sent by the voice server is received within the set time period, confirm that the connection between the user server side and the voice server is normal; If the connection normal flag sent by the voice server is not received within the set time period, resend a request to establish a long connection to the voice server and receive the connection success flag sent by the voice server to re - establish a long connection between the user server side and the voice server.

3. A method for realizing real-time voice human-machine dialogue, characterized in that, Applied to the agent server side, where the agent server side configures the corresponding number of agent personnel according to the number of robot agents undertaken in different interaction scenarios, combined with the incoming call connection rate and the transfer to agent manual rate; including: Receive the user's voice information sent by the voice server; the number of voice servers is multiple, and when the memory of a single voice server reaches the upper limit value, it will automatically match to other voice servers that have not reached the memory upper limit; Obtain the customer service voice corresponding to the voice information; Read the customer service voice according to the set byte threshold to obtain customer service voice information; Send the customer service voice information to the voice server, where the voice server forwards the customer service voice information to the user server side so that the user server side can obtain the customer service voice; Before receiving the voice information of the user sent by the receiving voice server, the method further includes: sending a request for establishing a long connection to the voice server; receiving an establishment success identifier sent by the voice server, so that the agent server end establishes a long connection with the voice server; After the user hangs up the phone, either the user server end or the agent server end sends a request for disconnecting the long connection to the voice server, so as to disconnect the connection with the voice server; Set the byte threshold: Based on the server configuration parameters of the agent server end and the clarity, fluency, and latency performance of manual test voice transmission, set the buffer parameter of the agent server end, and use the buffer parameter as the byte threshold.

4. The method according to claim 3, wherein During the period when the agent server end establishes a long connection with the voice server, the method further includes: Sending a heartbeat packet to the voice server according to a set time period; If a connection normal identifier sent by the voice server is received within the set time period, confirm that the connection between the agent server end and the voice server is normal; If a connection normal identifier sent by the voice server is not received within the set time period, resend a request for establishing a long connection to the voice server and receive an establishment success identifier sent by the voice server, so that the agent server end re - establishes a long connection with the voice server.

5. The method according to claim 3, wherein Before obtaining the customer service voice corresponding to the voice information, the method further includes: Allocating corresponding agents at the agent server end according to the voice transfer rate; Allocating the voice information to the corresponding agent.

6. A method for realizing real-time voice human-machine dialogue, characterized in that, Applied to a voice server, the number of voice servers is multiple. When the memory of a single voice server reaches the upper limit value, it is automatically matched to other voice servers that have not reached the memory upper limit; including: Receiving a request for establishing a long connection sent by the first server end, where the first server end is at least used for answering the user's voice or at least used for obtaining customer service voice information according to the user's voice information; Sending an establishment success identifier to the first server end, so that the first server end establishes a first long connection with the voice server; Sending the information from the first server end to the second server end at least through the first long connection; Receiving a request for disconnecting the long connection sent by either the first server end or the second server end, and disconnecting the first long connection and the second long connection with the first server end and the second server end; When the first server end or the second server end is the agent server end, the agent server end allocates the corresponding number of agents according to the number of robot agents undertaken in different interaction scenarios, combined with the incoming call connection rate and the transfer - to - agent manual rate; Among them, the first server and the second server read the voice according to a set byte threshold; the byte threshold is set as follows: based on the server configuration parameters of the first server and the second server and the clarity, fluency and latency performance of the artificial test voice transmission, the buffer parameters of the first server and the second server are set respectively, and the buffer parameters are used as the byte threshold.

7. The method according to claim 6, characterized in that, The first server is the user server, and the second server is the agent server; The method further includes: establishing a second long connection between the agent server and the voice server; Among them, The at least sending the information from the first server to the second server through the first long connection includes: Sending the voice of the user from the user server to the agent server through the first long connection and the second long connection, so that the agent server can obtain the voice.

8. The method according to claim 6, wherein The first server is the agent server, and the second server is the user server; The method further includes: establishing a second long connection between the user server and the voice server; Among them, The at least sending the information from the first server to the second server through the first long connection includes: Sending the customer service voice from the agent server to the user server through the first long connection and the second long connection, so that the user server can obtain the customer service voice.

9. The method according to claim 6, wherein During the period of establishing the first long connection between the first server and the voice server, the method further includes: If a heartbeat packet sent by the first server is received within a set time period, a connection normal flag is sent to the first server, where the connection normal flag is used to indicate that the network connection between the voice server and the first server is normal; If the heartbeat packet sent by the first server is not received within a set time period, the re-sent long connection establishment request from the first server is received and a connection establishment success flag is sent to the first server, so that the voice server and the first server re-establish a long connection.

10. A system for realizing real-time voice human-machine dialogue, characterized in that, The system is used to execute the method according to any one of claims 1-9, including: a user server, a voice server and an agent server, where The voice server establishes a first long connection with the user server and a second long connection with the agent server; The user server is configured to: Answer the voice of the user; read the voice according to a set byte threshold to obtain voice information; send the voice information to the voice server through the first long connection, where the voice server forwards the voice information to the agent server so that the agent server can obtain the voice; receive the customer service voice of the agent server sent by the voice server; The agent server is configured to: Receive the voice information of the user sent by the voice server; obtain the customer service voice corresponding to the voice information; read the customer service voice according to the set byte threshold to obtain customer service voice information; send the customer service voice information to the voice server through the second long connection, where the voice server forwards the customer service voice information to the user server side so that the user server side can obtain the customer service voice; The voice server is configured to receive the requests for disconnecting the long connection sent by the user server side and the agent server side, and disconnect the first long connection and the second long connection.

Citation Information

Patent Citations

  • Service method, client, system, electronic device and readable storage medium

    CN112200654A

  • Voice conversation method and device, storage medium and electronic equipment

    CN112767936A

  • Data exchange method, device, and system for group communication

    WO2014194647A1