Audio communication method, electronic device, storage medium and computer program product

By sending the pre-determined data of the target audio in advance in the intelligent voice dialogue system and sending the remaining data in periodic periods, the problem of large delay in the intelligent voice dialogue system is solved and the user experience is improved.

CN120223753APending Publication Date: 2025-06-27BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510405489.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the intelligent voice dialogue system, the delay in conversation between users and intelligent robots is large, which affects the user experience, which is mainly caused by network delay, voice recognition algorithm delay and jitter buffer zone delay.

Method used

After generating the target audio, the audio data of the previous predetermined time is sent to the target terminal at one time. After completion, the remaining audio data is sent in multiple predetermined periods, and the scaling adjustment function of the jitter buffer area is turned off at the target terminal.

Benefits of technology

By sending audio data of a predetermined duration in advance, the delay caused by the jitter cache area is eliminated, the overall delay of smart voice conversations is reduced, and the user experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223753A_ABST
    Figure CN120223753A_ABST
Patent Text Reader

Abstract

The invention relates to an audio communication method, electronic equipment, a storage medium and a computer program product. The audio communication method comprises the following steps: in response to receiving an input audio of a target terminal, generating a target audio corresponding to the input audio; first audio data in the target audio are sent to the target terminal at a time, and the first audio data are audio data of the previous preset duration in the target audio; and in response to completion of sending of the first audio data, sending second audio data in the target audio to the target terminal in a plurality of predetermined periods, the second audio data being audio data except the first audio data in the target audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of speech processing, and in particular, to an audio communication method, an electronic device, a storage medium, and a computer program product. Background Art

[0002] An intelligent voice dialogue system is a voice interaction system based on artificial intelligence technology that simulates the communication ability of humans to have conversations with users. The user uses the software on the terminal to connect to an intelligent robot equipped with an intelligent voice dialogue system through the network. The user's voice is collected by the microphone of the terminal and sent to the server. The intelligent robot on the server generates synthesized voice segments through technologies such as speech recognition, natural language understanding, and speech synthesis. The server then sends these voice segments to the software on the user's terminal for playback.

[0003] When a user has a conversation with an intelligent robot, latency is a very important factor affecting the user experience. The smaller the latency, the more natural the intelligent robot will feel to the user. However, there are many nodes on the link between the user and the intelligent robot that can introduce latency, such as network latency, latency of algorithms such as speech recognition, and latency of the jitter buffer on the user terminal, which results in a relatively large latency, making the intelligent robot feel unnatural to the user and the user experience being poor. Summary of the Invention

[0004] The present disclosure provides an audio communication method, an electronic device, a storage medium, and a computer program product to solve at least one of the above-related technical problems.

[0005] According to a first aspect of an embodiment of the present disclosure, an audio communication method is provided, which is applied to a server. The audio communication method includes: in response to receiving input audio from a target terminal, generating target audio corresponding to the input audio; sending first audio data in the target audio to the target terminal at one time, where the first audio data is audio data of a pre-determined duration at the beginning of the target audio; in response to the completion of sending the first audio data, sending second audio data in the target audio to the target terminal in multiple predetermined periods, where the second audio data is audio data in the target audio other than the first audio data.

[0006] Optionally, the completion of sending the first audio data is determined by the following method: calculating the time difference between the current time and the start time, and calculating the sum of the time difference and the pre-determined duration, where the start time is the time when sending the first audio data to the target terminal begins; in response to the sum being less than or equal to the total duration of the audio data already sent in the target audio, stopping continuously sending the audio data of the target audio to the target terminal and determining that the sending of the first audio data is completed.

[0007] Optionally, in response to the sum of times being greater than the total duration of the audio data already sent in the target audio, continuously send the audio data of the target audio to the target terminal.

[0008] Optionally, the first audio data and the second audio data are sent in the form of audio packets, and the audio communication method further includes: receiving an acknowledgment message feedback by the target terminal for each audio packet; in response to not receiving the acknowledgment message corresponding to any audio packet, repeatedly send any audio packet to the target terminal within a predetermined duration until the acknowledgment message for any audio packet is received.

[0009] Optionally, the audio data of the target audio sent to the target terminal is stored in the jitter buffer of the target terminal, wherein the stretching and adjusting function of the jitter buffer is turned off.

[0010] According to a second aspect of the embodiments of the present disclosure, there is provided an audio communication method applied to a target terminal. The audio communication method includes: sending input audio to a server; receiving, at one time by the server, the first audio data in the target audio corresponding to the input audio, where the first audio data is the audio data of the first predetermined duration in the target audio; receiving the second audio data in the target audio sent by the server in multiple predetermined periods and playing the received audio data of the target audio, where the second audio data is the audio data in the target audio other than the first audio data.

[0011] Optionally, the first audio data and the second audio data are received in the form of audio packets, and the audio communication method further includes: for each audio packet, feedback a corresponding acknowledgment message to the server; in response to not receiving any audio packet, receive, within a predetermined duration, any audio packet repeatedly sent by the server until any audio packet is received, and feedback an acknowledgment message for any audio packet to the server.

[0012] Optionally, the audio communication method further includes: storing the received audio data of the target audio in the jitter buffer of the target terminal, wherein the stretching and adjusting function of the jitter buffer is turned off; obtaining and playing the received audio data of the target audio from the jitter buffer.

[0013] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: at least one processor; at least one memory storing computer-executable instructions, wherein when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the above audio communication method.

[0014] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are run by at least one processor, causing at least one processor to execute the above audio communication method.

[0015] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product including computer instructions which, when executed by a processor, implement the above audio communication method.

[0016] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: According to the audio communication method, electronic device, storage medium and computer program product of the present disclosure, after generating the target audio corresponding to the input audio, the audio data of the first predetermined duration in the target audio is sent to the target terminal at one time, and then after the audio data of the first predetermined duration is sent, the remaining audio data in the target audio is sent to the target terminal in multiple predetermined cycles, so that when the target terminal plays the target audio later, it has already received the audio data of the predetermined duration in advance, thereby eliminating the delay caused by the jitter buffer on the target terminal without sacrificing network resistance, and further reducing the overall delay of the intelligent voice conversation and improving the experience of the intelligent voice conversation.

[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0019] Figure 1 is a schematic flowchart of a traditional RTC scheme according to an exemplary embodiment of the present disclosure; Figure 2 is a schematic diagram of IAT probability distribution according to an exemplary embodiment of the present disclosure; Figure 3 is a schematic diagram of an implementation scenario of an audio communication method according to an exemplary embodiment of the present disclosure; Figure 4 is a flowchart of an audio communication method according to an exemplary embodiment of the present disclosure; Figure 5 is a flowchart of another audio communication method according to an exemplary embodiment of the present disclosure; Figure 6 is a schematic system flowchart of an audio communication method according to an exemplary embodiment of the present disclosure; Figure 7A is a schematic diagram of the delay of a traditional RTC scheme according to an exemplary embodiment of the present disclosure; Figure 7B is a schematic diagram of the delay of the solution of the present disclosure according to an exemplary embodiment of the present disclosure; Figure 8is a block diagram of an audio communication system shown according to an exemplary embodiment of the present disclosure; Figure 9 is a view of a computing environment coupled to a user interface shown according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0020] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0022] It should be noted here that "at least one of several items" in the present disclosure all represents the three types of parallel situations including "any one of the several items", "any combination of multiple items of the several items", and "the whole of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "performing at least one of step one and step two" means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing step one and step two.

[0023] Currently, there are two processing methods for intelligent voice dialogue systems: 1) The voice generated by the intelligent robot will be sent to the user in the form of a file, and the user will play it after receiving the complete file. This method requires waiting until the file reception is completed before playing, resulting in a certain delay; 2) A real-time communication (RTC) channel is established between the intelligent robot and the user. The voice generated by the intelligent robot is divided into multiple audio frames, and the audio frames are then transmitted through the RTC channel. In this method, the user receiving end needs to have a jitter buffer (JB) to counteract network jitter and network packet loss, and the jitter buffer itself will introduce delay.

[0024] Here is a brief introduction to the traditional RTC transmission scenario. When using the RTC channel for voice transmission, the user terminal will have a jitter buffer to cache the audio packets received from the intelligent robot on the server, so as to resist the jitter that appears on the network, such as Figure 1 shown. Specifically, the intelligent robot sends audio frames in units of audio packets at fixed intervals (for example, 20 ms). The user terminal calculates the probability distribution of the Inter Arrival Time (IAT) based on the time interval between the arrival of audio packets, that is, the Inter Arrival Time, and calculates the current network jitter according to the probability distribution. Suppose as Figure 2 shown, 98% of the IAT is within 200 ms. At this time, it can be considered that the network jitter is 200 ms, and the target water level (that is, the storage size) of the jitter buffer can be adjusted to 200 ms. In this way, when the network jitters within 200 ms, the smoothness of the sound on the user terminal can still be guaranteed. The adjustment strategy for maintaining the target water level of the jitter buffer is simply as follows: when the audio duration stored in the jitter buffer is higher than the target water level, the stored audio is played at an accelerated speed to reduce the audio stored in the jitter buffer; when the audio duration stored in the jitter buffer is lower than the target water level, the stored audio is played at a slow speed to increase the audio stored in the jitter buffer.

[0025] In view of the above problems, the present disclosure proposes an audio communication method. After generating the target audio corresponding to the input audio, the audio data of the first predetermined duration in the target audio is sent to the target terminal at one time. Then, after the audio data of the first predetermined duration is sent, the remaining audio data in the target audio is sent to the target terminal in multiple predetermined cycles and the target terminal is notified to play the received audio data of the target audio, so that when the target terminal plays the received audio data of the target audio, it has already received the audio data of the predetermined duration in advance, thereby eliminating the delay of the jitter buffer in the RTC solution without losing network resistance, and thus reducing the delay from the intelligent robot to the user receiving end.

[0026] Figure 3 is a schematic diagram of an implementation scenario of an audio communication method shown according to an exemplary embodiment of the present disclosure, such as Figure 3 described. This implementation scenario includes a server 300, a user terminal 310, and a user terminal 320. Among them, the number of user terminals is not limited to 2, including but not limited to devices such as mobile phones and personal computers. The user terminal can install an application program for intelligent voice dialogue. The server can be a single server, or a server cluster composed of several servers, or a cloud computing platform or a virtualization center.

[0027] The server 300 receives the input audio from the user terminal 310 or the user terminal 320, such as a question, a request, etc., and then generates a target audio corresponding to the input audio in response to receiving the input audio of the target terminal, and sends the audio data of the first predetermined duration in the target audio to the user terminal 310 or the user terminal 320 at one time; in response to the completion of the transmission of the first audio data, the remaining audio data in the target audio is sent to the user terminal 310 or the user terminal 320 in multiple predetermined periods, and then the user terminal 310 or the user terminal 320 will play the received audio data of the target audio.

[0028] Next, an audio communication method, an electronic device, a storage medium, and a computer program product according to an exemplary embodiment of the present disclosure will be described in detail with reference to the accompanying drawings.

[0029] Figure 4 is a flowchart of an audio communication method shown according to an exemplary embodiment of the present disclosure, as Figure 4 shown, the audio communication method is applied to a server and includes the following steps: In step S401, in response to receiving the input audio of the target terminal, a target audio corresponding to the input audio is generated. As an example, after receiving the input audio of the target terminal, the intelligent robot on the server can generate one or more segments of speech at one time based on the input audio, that is, the target audio includes one or more segments, so it is equivalent to knowing the "future" speech information.

[0030] As an example, after generating the target audio, the target audio can be divided into a predetermined number of audio frames, and when the target terminal is sent subsequently, the target audio is transmitted to the target terminal in the order of the audio frames.

[0031] In step S402, the first audio data in the target audio is sent to the target terminal at one time, where the first audio data is the audio data of the first predetermined duration in the target audio.

[0032] As an example, the above-mentioned predetermined duration can be set as needed, and the present disclosure does not limit this.

[0033] As an example, since the target audio contains "future" voice information, when starting to send the target audio to the target terminal, the audio data of a previous period of the target audio can be sent to the target terminal in advance, so as to eliminate the delay caused by the jitter buffer. For example, assuming that the target audio is a 1-minute audio and the predetermined duration is 1 second, when starting to send the target audio to the target terminal, the audio data of the first 1 second of the target audio can be sent to the target terminal in advance, and the remaining 59 seconds of audio data of the target audio can be sent to the target terminal in the original manner, which is not limited in this disclosure.

[0034] It should be noted that the audio data of the target audio can be sent to the target terminal in units of audio packets, and each audio packet can contain a predetermined number of audio frames, which is not limited in this disclosure.

[0035] In step S403, in response to the completion of the transmission of the first audio data, the second audio data in the target audio is sent to the target terminal in multiple predetermined periods, where the second audio data is the audio data in the target audio other than the first audio data. As an example, after the transmission of the first audio data is completed, the remaining audio data in the target audio other than the first audio data can be sent to the target terminal in the original transmission manner, such as divided into multiple predetermined periods after the start time, which is not limited in this disclosure; and after the transmission of the first audio data is completed, the target terminal is notified to play the received audio data of the target audio. At this time, when the target terminal plays the received audio data of the target audio, since the audio data of the predetermined duration has been received in advance, the delay caused by the jitter buffer can be eliminated.

[0036] As an example, the server can notify the target terminal to play the received target audio after the transmission of the first audio data is completed. After receiving the notification, the target terminal starts to play the target audio; the target terminal can also play the target audio according to the timing of playing the target audio in the original periodic transmission manner, which is not limited in this disclosure.

[0037] According to the exemplary embodiment of the present disclosure, the completion of the transmission of the first audio data is determined in the following manner: calculate the time difference between the current time and the start time, and calculate the sum of the time difference and the predetermined duration; in response to the sum being less than or equal to the total duration of the audio data already sent in the target audio, stop continuously sending the audio data of the target audio to the target terminal, and determine that the first audio data transmission is completed. Through this embodiment, it is convenient and accurate to know that the first audio data transmission is completed and start the subsequent periodic transmission, so that the delay caused by the jitter buffer on the target terminal can be continuously eliminated, and the jitter buffer can also be prevented from storing too much data.

[0038] According to an exemplary embodiment of the present disclosure, in response to the sum of time being greater than the total duration of the audio data that has been sent in the target audio, the audio data of the target audio is continuously sent to the target terminal. Through this embodiment, it can be ensured that the first audio data is completely sent.

[0039] As an example, when starting to send the target audio, record the current time as start_time. Assume that the total duration of the audio data that has been sent is n, which will increase as each audio frame is sent, and the current time is current_time. If (current_time – start_time) + t > n, then continuously send the audio data of the target audio to the target terminal; if (current_time – start_time) + t <= n, then stop continuously sending the audio data of the target audio to the target terminal, determine that the first audio data is sent completely, and then wait for the timer to trigger and perform the subsequent periodic sending process.

[0040] According to an exemplary embodiment of the present disclosure, the first audio data and the second audio data are sent in the form of audio packets. After sending the audio packets of the target audio to the target terminal, it is also possible to receive the confirmation messages feedback by the target terminal for each audio packet; in response to the confirmation message corresponding to any audio packet not being received, the any audio packet is repeatedly sent to the target terminal within a predetermined duration until the confirmation message of the any audio packet is received. Through this embodiment, the retransmission mechanism is used to resist the loss of audio quality caused by network packet loss. Because the audio data of a predetermined duration is sent in advance, as long as the retransmission is completed within the predetermined duration, the packet loss can be recovered, thereby avoiding the loss of audio quality.

[0041] As an example, the audio data of the target audio can be sent to the target terminal in units of audio packets. After receiving the audio packet, the target terminal will feedback a confirmation message to the server, informing the server that the audio packet has been successfully received. However, if there is a packet loss situation due to network or other reasons, the target terminal will not feedback a confirmation message to the server. At this time, the server can retransmit the lost audio packet to the target terminal until the target terminal successfully receives the audio packet. It should be noted that in the present disclosure, because the audio data of a predetermined duration is sent in advance, as long as the lost audio packet is retransmitted and sent within the predetermined duration, the packet loss can be recovered.

[0042] According to an exemplary embodiment of the present disclosure, the audio data sent to the target terminal in the target audio is stored in the jitter buffer of the target terminal, wherein the function of stretching and adjusting the jitter buffer is turned off. Through this embodiment, by turning off the function of stretching and adjusting the jitter buffer, it is possible to maintain the target audio playing at a normal speed without the need to accelerate or slow down the target audio, thereby improving the user comfort.

[0043] As an example, the audio data of the target audio successfully sent by the server to the target terminal is stored in the jitter buffer of the target terminal, which facilitates the target terminal to obtain and play the corresponding audio data from the jitter buffer. Moreover, the jitter buffer can turn off the stretching adjustment function so that the received audio data can be played at the normal speed.

[0044] Figure 5 is a flowchart of another audio communication method shown according to an exemplary embodiment of the present disclosure. As Figure 5 shown, the audio communication method is applied to the target terminal and includes the following steps: In step S501, an input audio is sent to the server.

[0045] As an example, the user can initiate a question or chat content (i.e., the input audio) to the intelligent robot through the target terminal, that is, send the input audio to the server through the target terminal. After receiving the input audio from the target terminal, the intelligent robot on the server can generate one or more segments of speech at one time based on the input audio, that is, the target audio includes one or more segments. Therefore, it is equivalent to knowing the "future" speech information.

[0046] In step S502, the first audio data in the target audio corresponding to the input audio sent by the server at one time is received, where the first audio data is the audio data of the target audio in the previous predetermined duration.

[0047] As an example, the above-mentioned predetermined duration can be set as needed, and the present disclosure does not limit this.

[0048] As an example, since the target audio contains "future" speech information, when the server starts to send the target audio to the target terminal, it can send the audio data of the previous duration in the target audio to the target terminal in advance, so as to eliminate the delay caused by the jitter buffer. For example, assuming that the target audio is a 1-minute audio and the predetermined duration is 1 second, when the server starts to send the target audio to the target terminal, the first 1-second audio data of the target audio can be sent to the target terminal in advance, and the remaining 59-second audio data of the target audio can be sent to the target terminal in the original way. The present disclosure does not limit this.

[0049] In step S503, the second audio data in the target audio sent by the server in multiple predetermined periods is received and the audio data of the received target audio is played, where the second audio data is the audio data of the target audio except the first audio data.

[0050] As an example, after the first audio data is sent, the remaining audio data in the target audio except the first audio data can be sent to the target terminal in the original sending manner, such as divided into multiple predetermined periods after the start time, and the present disclosure does not limit this; after the first audio data is sent, the target terminal plays the received audio data.

[0051] According to an exemplary embodiment of the present disclosure, the first audio data and the second audio data are received in the form of audio packets. After the target terminal receives the audio packets of the target audio, corresponding confirmation messages can also be fed back to the server for each audio packet; in response to not receiving any audio packet, any audio packet repeatedly sent by the server is received within a predetermined duration until any audio packet is received, and a confirmation message for any audio packet is fed back to the server. Through this embodiment, the retransmission mechanism is used to resist the loss of sound quality caused by network packet loss. Because the audio data of a predetermined duration is sent in advance, as long as the retransmission is completed within the predetermined duration, the packet loss can be recovered, thereby avoiding the loss of sound quality.

[0052] As an example, the audio data of the target audio can be sent to the target terminal in units of audio packets. After the target terminal receives an audio packet, it will feed back a confirmation message to the server to inform the server that the audio packet has been successfully received. However, if packet loss occurs due to reasons such as the network, the target terminal does not receive the lost audio packet, and thus will not feed back the confirmation message corresponding to the lost audio packet to the server. At this time, the server can retransmit the lost audio packet to the target terminal until the target terminal successfully receives the audio packet. It should be noted that since the audio data of a predetermined duration is sent in advance in the present disclosure, as long as the lost audio packet is retransmitted and sent within the predetermined duration, the packet loss can be recovered.

[0053] According to an exemplary embodiment of the present disclosure, the target terminal can also store the received audio data of the target audio in the jitter buffer of the target terminal, where the stretching and adjusting function of the jitter buffer is turned off; the received audio data of the target audio is obtained from the jitter buffer and played. Through this embodiment, by turning off the stretching and adjusting function of the jitter buffer, the target audio can be played at a normal speed, and there is no need to accelerate or slow down the playback of the target audio, improving user comfort.

[0054] As an example, the audio data of the target audio successfully sent by the server to the target terminal is stored in the jitter buffer of the target terminal, which is convenient for obtaining and playing the corresponding audio data from the jitter buffer, and the jitter buffer can turn off the stretching and adjusting function so that the received audio data can be played at a normal speed.

[0055] To facilitate the understanding of the present disclosure, the following is combined with Figure 6 the description of the system.

[0056] Figure 6 The system flow diagram of the audio communication method is shown. As Figure 6 shown, after the intelligent robot receives the input audio of the user at the target terminal, it will generate the target audio corresponding to the input audio, such as Figure 6 the audio files 1, 2, 3...n described therein. It sends the audio data of the previous predetermined duration in the audio file to the target terminal in advance. For example, it sends the audio data of duration t in advance. After sending the audio data of duration t3, it periodically sends the remaining audio data. The target terminal stores the received audio data in the jitter buffer. After sending the audio data of duration t3, that is, when periodically sending the remaining audio data, the target terminal starts to play the received audio data. In this way, when the target terminal plays the audio data of the received target audio, it has already received the audio data of the predetermined duration t3 in advance, so as to eliminate the delay of the jitter buffer in the RTC scheme without losing network resistance, thereby reducing the delay from the intelligent robot to the user receiving end. Furthermore, the stretch adjustment function of the jitter buffer of the present disclosure is turned off, so that the target terminal can play the received audio data at a normal speed.

[0057] Figure 7A The delay diagram of the traditional RTC scheme is shown. As Figure 7A shown, the audio data at time t1 will be played at time t1 + t2 + t3, with a delay of t2 + t3. Therefore, the traditional RTC scheme not only has the delay t2 brought by the transmission process between the server and the target terminal, but also has the delay t3 brought by the jitter buffer. Figure 7B The delay diagram of the scheme of the present disclosure is shown. As Figure 7B shown, the audio data at time t1 will be played at time t1 + t2, with a delay of t2. Therefore, the scheme of the present disclosure only has the delay t2 brought by the transmission process between the server and the target terminal, and no longer has the delay t3 brought by the jitter buffer.

[0058] Figure 8 is a block diagram of an audio communication system shown according to an exemplary embodiment of the present disclosure. Referring to Figure 8 , the system 800 includes a server 801 and a target terminal 802.

[0059] The target terminal 802 is configured to send input audio to the server 801; the server 801 is configured to generate target audio corresponding to the input audio in response to receiving the input audio of the target terminal 802; send the first audio data in the target audio to the target terminal 802 at one time, where the first audio data is the audio data of the first predetermined duration in the target audio; in response to the completion of the sending of the first audio data, send the second audio data in the target audio to the target terminal 802 in multiple predetermined periods, where the second audio data is the audio data in the target audio other than the first audio data. Optionally, the completion of the sending of the first audio data is determined in the following manner: calculate the time difference between the current time and the start time, and calculate the sum of the time difference and the predetermined duration, where the start time is the time when the sending of the first audio data to the target terminal 802 starts; in response to the sum of the time being greater than the total duration of the audio data already sent in the target audio, continuously send the audio data of the target audio to the target terminal 802; in response to the sum of the time being less than or equal to the total duration of the audio data already sent in the target audio, stop continuously sending the audio data of the target audio to the target terminal 802, and determine that the sending of the first audio data is completed.

[0060] Optionally, the first audio data and the second audio data are sent in the form of audio packets, and the server 801 receives the confirmation messages feedback by the target terminal 802 for each audio packet; in response to the confirmation message corresponding to any audio packet not being received, repeatedly send any audio packet to the target terminal 802 within a predetermined duration until the confirmation message of any audio packet is received.

[0061] Optionally, the audio data of the target audio received by the target terminal 802 is stored in the jitter buffer of the target terminal 802, where the stretching and adjustment function of the jitter buffer is turned off; obtain and play the audio data of the received target audio from the jitter buffer.

[0062] Figure 9 A computing environment 910 coupled to a user interface 950 is shown. The computing environment 910 can be part of a data processing server. The computing environment 910 includes a processor 920, a memory 930, and an input / output (I / O) interface 940.

[0063] The processor 920 generally controls the overall operation of the computing environment 910, such as operations associated with display, data acquisition, data communication, and image processing. The processor 920 may include one or more processors for executing instructions to perform all or some of the steps in the above methods. In addition, the processor 920 may include one or more modules that facilitate the interaction between the processor 920 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip microcomputer, a graphics processing unit (GPU), etc.

[0064] Memory 930 is configured to store various types of data to support the operation of computing environment 910. Memory 930 may include predetermined software 932. Examples of such data include instructions for any application or method operating on computing environment 910, video data sets, image data, and the like. Memory 930 may be implemented by using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0065] I / O interface 940 provides an interface between processor 920 and peripheral interface modules (such as a keyboard, click wheel, buttons, etc.). The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 940 may be coupled to an encoder and a decoder.

[0066] In an embodiment, computing environment 910 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.

[0067] According to an embodiment of the present disclosure, an electronic device may be provided, the electronic device including at least one memory and at least one processor, a set of computer-executable instructions being stored in the at least one memory, and when the set of computer-executable instructions is executed by the at least one processor, an audio communication method according to an embodiment of the present disclosure is performed.

[0068] As an example, the electronic device may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above set of instructions. Here, electronic device 1000 does not have to be a single electronic device, and may also be any assembly of devices or circuits capable of executing the above instructions (or instruction sets) alone or jointly. The electronic device may also be a part of an integrated control system or a system manager, or may be configured to be interconnected with a local or remote (e.g., via wireless transmission) interface as a portable electronic device.

[0069] In addition, the electronic device may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device may be connected to each other via a bus and / or a network.

[0070] According to an embodiment of the present disclosure, a computer-readable storage medium may also be provided, wherein when instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the audio communication method of the embodiment of the present disclosure. Examples of the computer-readable storage medium here include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0071] According to an embodiment of the present disclosure, a computer program product including a plurality of programs, such as those in memory 930, is provided, and the plurality of programs may be executed by a processor 920 in a computing environment 910 to perform the above method. For example, the computer program product may include a non-transitory computer-readable storage medium.

[0072] Unless otherwise specifically stated, the order of steps of the method according to the present disclosure is only illustrative, and the steps of the method according to the present disclosure are not limited to the specific order described above, but may be changed according to the actual situation. In addition, at least one of the steps of the method according to the present disclosure may be adjusted, combined, or deleted according to actual needs.

[0073] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are indicated by the appended claims.

[0074] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. An audio communication method, characterized in that: Applied to a server, the audio communication method comprises: In response to receiving input audio from a target terminal, generating target audio corresponding to the input audio; Sending the first audio data in the target audio to the target terminal at one time, wherein the first audio data is the audio data of the first predetermined time length in the target audio; In response to completion of sending the first audio data, second audio data in the target audio is sent to the target terminal in a plurality of predetermined cycles, wherein the second audio data is audio data in the target audio excluding the first audio data.

2. The audio communication method according to claim 1, characterized in that: The sending of the first audio data is determined to be completed by: Calculating a time difference between a current time and a start time, and calculating a time difference between the time difference and the predetermined duration, wherein the start time is a time when the first audio data starts to be sent to the target terminal; In response to the time being less than or equal to the total duration of the audio data that has been sent in the target audio, continuously sending the audio data of the target audio to the target terminal is stopped, and it is determined that the sending of the first audio data is completed.

3. The audio communication method according to claim 2, characterized in that: Also includes: In response to the time being greater than the total duration of the audio data that has been sent in the target audio, the audio data of the target audio is continuously sent to the target terminal.

4. The audio communication method according to claim 1, wherein: The first audio data and the second audio data are sent in the form of audio packets, and the audio communication method further includes: Receiving a confirmation message fed back by the target terminal for each audio package; In response to the confirmation message corresponding to any audio package not being received, repeatedly sending the any audio package to the target terminal within the predetermined time period until a confirmation message of the any audio package is received.

5. The audio communication method according to claim 1, wherein: The audio data in the target audio that is sent to the target terminal is stored in a jitter buffer area of ​​the target terminal, wherein a scaling adjustment function of the jitter buffer area is disabled.

6. An audio communication method, characterized in that: Applied to a target terminal, the audio communication method comprises: Send input audio to the server; Receiving first audio data in the target audio corresponding to the input audio sent by the server at one time, wherein the first audio data is audio data of a first predetermined time length in the target audio; Receive second audio data in the target audio sent by the server in multiple predetermined periods, wherein the second audio data is audio data in the target audio except the first audio data.

7. The audio communication method according to claim 6, characterized in that: The first audio data and the second audio data are received in the form of audio packets, and the audio communication method further includes: For each audio packet, feeding back a corresponding confirmation message to the server; In response to not receiving any audio package, receiving any audio package repeatedly sent by the server within the predetermined time period until any audio package is received, and feeding back a confirmation message for any audio package to the server.

8. The audio communication method according to claim 6, characterized in that: The audio communication method further comprises: storing the received audio data of the target audio in a jitter buffer area of ​​the target terminal, wherein a scaling adjustment function of the jitter buffer area is disabled; The received audio data of the target audio is acquired from the jitter buffer and played.

9. An electronic device, characterized in that: include: at least one processor; at least one memory storing computer executable instructions, Wherein, when the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to perform the audio communication method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor is prompted to perform the audio communication method according to any one of claims 1 to 8.

11. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the audio communication method according to any one of claims 1 to 8 is implemented.