A method, apparatus, electronic device, and storage medium for adjusting buffer length
The method dynamically adjusts buffer zone length based on jitter and speech characteristics to enhance audio playback smoothness and quality in voice communications.
Patent Information
- Application Number
- CN202111312574.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-08
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-08
AI Technical Summary
In the prior art, the adjustment accuracy of the buffer length is not high, resulting in problems of lag and sound quality damage when playing audio signals under network jitter.
By detecting the jitter value and speech detection results of the audio data packet, combined with the speech speed detection results, the buffer length is dynamically adjusted to reduce the buffer length when there is no speech data and increase the buffer redundancy space when there is voice data, avoiding signal delay and loss.
It improves the smoothness of audio signal playback, reduces delay and lag, and improves the playback quality of audio signal.
Smart Images

Figure CN116095395B_ABST
Abstract
Description
Background Art
[0002] Currently, with the development of network technology, during the process of transmitting each audio data packet of an audio signal to a receiving end, various network problems, such as congestion, network errors, etc., may cause problems such as packet loss and delayed arrival of each audio data packet received by the receiving end, thereby reducing the playback quality of the audio signal obtained after decoding.
[0003] To solve the above problems, after the receiving end receives each audio data packet, each audio data packet can be stored in a buffer for caching and sent to a decoding end for decoding at the same time interval, so that the audio signal obtained after decoding can be played smoothly.
[0004] In practical applications, when the buffer length is set too long, unnecessary signal delay will be increased; when the buffer length is set too short, buffer overload will occur, resulting in packet loss and causing the problem of jerky call sounds. Therefore, how to dynamically adjust the buffer length has become an urgent problem to be solved. Summary of the Invention
[0005] Embodiments of the present application provide a method, apparatus, electronic device, and storage medium for adjusting the buffer length, which are used to improve the accuracy of dynamically adjusting the buffer length, thereby improving the smoothness of playing the audio signal.
[0006] On the one hand, embodiments of the present application provide a method for adjusting the buffer length, including:
[0007] Receiving each audio data packet sent by a sending end, where each audio data packet is obtained by encoding at least one audio frame of an audio signal;
[0008] Respectively obtaining the jitter values detected when receiving the respective audio data packets;
[0009] Obtaining the voice detection result corresponding to each audio frame and the speech rate detection result of the audio signal; wherein, the voice detection result indicates whether the corresponding audio frame contains voice data; the speech rate detection result is a result obtained by detecting the voice content contained in the audio signal per unit time;
[0010] Based on the speech rate detection result, each voice detection result, and each jitter value, and in combination with a preset buffer length adjustment strategy, correspondingly adjust the current buffer length of the buffer.
[0011] On the one hand, embodiments of the present application provide an apparatus for adjusting the buffer length, including:
[0012] A receiving module, configured to receive each audio data packet sent by a sending end, where each audio data packet is obtained by encoding at least one audio frame of an audio signal;
[0013] A jitter detection module, configured to respectively obtain jitter values detected when receiving each of the audio data packets;
[0014] A processing module, configured to obtain a voice detection result corresponding to each audio frame and a speech rate detection result of the audio signal; where the voice detection result indicates whether the corresponding audio frame contains voice data; the speech rate detection result is a result obtained by detecting the voice content contained in the audio signal within a unit time;
[0015] An adjustment module, configured to correspondingly adjust the current buffer length of a buffer based on the speech rate detection result, each voice detection result, and each jitter value, in combination with a preset buffer length adjustment strategy.
[0016] In a possible embodiment, the adjustment module is further configured to:
[0017] Based on the reception time corresponding to each of the audio data packets, determine, from the audio data packets, audio data packets that meet the reception time condition;
[0018] If it is determined that each voice detection result does not contain voice data, correspondingly adjust the current buffer length of the buffer based on a non-voice control strategy and a target jitter value of the determined audio data packets;
[0019] If it is determined that at least one voice detection result contains voice data, correspondingly adjust the buffer length based on the speech rate detection result and the target jitter value, in combination with a voice control strategy.
[0020] In a possible embodiment, when correspondingly adjusting the current buffer length of the buffer based on the non-voice control strategy and the target jitter value of the determined audio data packets, the adjustment module is further configured to:
[0021] If the target jitter value of the determined audio data packet is less than a jitter value threshold, use a preset first length as the current buffer length of the buffer;
[0022] If it is determined that the target jitter value is not less than the jitter value threshold, select any target length from a preset length range as the buffer length, where the length range is generated based on the first length and a second length, and the first length is less than the second length.
[0023] In a possible embodiment, when adjusting the buffer length accordingly based on the speech rate detection result and the target jitter value in combination with a voice control strategy, the adjustment module is further configured to:
[0024] If it is determined that the target jitter value is less than the jitter value threshold, select any target length from a preset length range as the buffer length, where the length range is generated based on the first length and the second length, and the first length is less than the second length;
[0025] If it is determined that the target jitter value is not less than the jitter value threshold, determine the buffer length based on the speech rate detection result and the target jitter value.
[0026] In a possible embodiment, when determining the buffer length based on the speech rate detection result and the target jitter value, the adjustment module is further configured to:
[0027] Determine a speech rate adjustment parameter of the audio signal based on the speech rate detection result;
[0028] Determine a third length based on the speech rate adjustment parameter, the target jitter value, and a preset jitter value adjustment function;
[0029] Select a target length that meets the length condition from the third length and the second length as the buffer length.
[0030] In a possible embodiment, when determining the third length based on the speech rate adjustment parameter, the target jitter value, and a preset jitter value adjustment function, the processing module is further configured to:
[0031] Use the target jitter value as a variable of the preset jitter value adjustment function, and obtain a fourth length based on the target jitter value and the jitter value adjustment function;
[0032] Obtain the third length based on the speech rate adjustment parameter and the fourth length, where the speech rate adjustment parameter is positively correlated with the fourth length.
[0033] In a possible embodiment, when selecting a target length that meets the length condition from the third length and the second length, the processing module is further configured to:
[0034] If it is determined that the third length is greater than the second length, use the second length as the target length;
[0035] If it is determined that the third length is not greater than the second length, use the third length as the target length.
[0036] In a possible embodiment, when obtaining the speech rate detection result of the audio signal, the processing module is further configured to:
[0037] Determine, from each audio frame, each target audio frame whose speech detection result contains speech data;
[0038] Obtain the pitch period state corresponding to each of the target audio frames;
[0039] Based on each pitch period state, determine the number of state switches of the audio signal;
[0040] According to the number of state switches, the number of audio frames corresponding to each of the target audio frames, and a preset speech rate threshold, determine the speech rate detection result of the audio signal.
[0041] In a possible embodiment, when obtaining the pitch period state corresponding to each of the target audio frames, the processing module is further configured to:
[0042] For each of the target audio frames, respectively perform the following operations:
[0043] Perform pitch period detection on a target audio frame to obtain the pitch period value corresponding to the target audio frame;
[0044] Based on the pitch period value of the target audio frame and the pitch period value of the previous target audio frame corresponding to the target audio frame, determine the pitch period value difference;
[0045] According to the pitch period value difference and a difference threshold, determine the pitch period state corresponding to the target audio frame.
[0046] In a possible embodiment, when respectively determining the jitter value corresponding to each audio data packet, the jitter detection module is further configured to:
[0047] For each of the audio data packets, respectively perform the following operations:
[0048] Based on the reception time and transmission time of an audio data packet, and the reception time and transmission time of the previous audio data packet of the audio data packet, determine the time difference between the audio data packet and the previous audio data packet;
[0049] Based on the time difference, a smoothing coefficient, and the jitter value corresponding to the previous audio data packet, obtain the jitter value corresponding to the audio data packet.
[0050] On the one hand, an embodiment of the present application provides an electronic device, which includes a processor and a memory. Among them, the memory stores program codes. When the program codes are executed by the processor, the processor executes the steps of any of the above methods for adjusting the buffer length.
[0051] On the one hand, an embodiment of the present application provides a computer storage medium. The computer storage medium stores computer instructions. When the computer instructions run on a computer, the computer executes the steps of any of the above methods for adjusting the buffer length.
[0052] On the one hand, an embodiment of the present application provides a computer program product, which includes computer instructions. The computer instructions are stored in a computer-readable storage medium. When a processor of an electronic device reads the computer instructions from the computer-readable storage medium, the processor executes the computer instructions, so that the electronic device executes the steps of any of the above methods for adjusting the buffer length.
[0053] Since the embodiments of the present application adopt the above technical solutions, they have at least the following technical effects:
[0054] In the solution of the embodiment of the present application, after obtaining the jitter values detected when receiving each audio data packet respectively, based on each jitter value, the voice detection results and the speech rate detection results of the corresponding audio signals of each audio frame are obtained. Combining the corresponding buffer length adjustment strategy, the current buffer length of the buffer is adjusted accordingly.
[0055] Adopting the above solution, the audio signal is analyzed, and the buffer length is dynamically adjusted by combining the jitter value, the voice detection result and the speech rate detection result. Since losing audio frames that do not contain voice data will not reduce the call quality during a call, when the audio signal does not contain voice data, minimizing the buffer length as much as possible can avoid signal delay problems caused by too long a buffer length. Since losing audio frames that contain voice data will reduce the smoothness of the audio signal playback and cause call stuttering, when there is voice data and the speech rate is fast, increasing the buffer length as much as possible can avoid audio signal loss caused by buffer overload. Therefore, adopting the above solution can improve the accuracy of audio signal playback, thereby improving the smoothness of audio signal playback.
[0056] Other features and advantages of the present application will be described in the subsequent specification, and part of them will become obvious from the specification or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0058] Figure 1 Schematic diagram of the application scenario in the embodiments of the present application;
[0059] Figure 2 Flowchart of the method for adjusting the buffer length in the embodiments of the present application;
[0060] Figure 3 Schematic diagram of the process for obtaining the jitter value in the embodiments of the present application;
[0061] Figure 4 Schematic diagram of the process for obtaining a voice detection result in the embodiments of the present application;
[0062] Figure 5 Another schematic diagram of the process for obtaining a voice detection result in the embodiments of the present application;
[0063] Figure 6 Schematic diagram of the process for obtaining a speech rate detection result in the embodiments of the present application;
[0064] Figure 7 Another schematic diagram of the process for obtaining a speech rate detection result in the embodiments of the present application;
[0065] Figure 8 Schematic diagram of the process for performing speech rate detection on an audio signal in the embodiments of the present application;
[0066] Figure 9 Schematic diagram of the process for obtaining the pitch period state method in the embodiments of the present application;
[0067] Figure 10 Example diagram for determining the number of state transitions of an audio signal in the embodiments of the present application;
[0068] Figure 11 Schematic diagram of the process for adjusting the buffer length in the embodiments of the present application;
[0069] Figure 12 Schematic diagram of the process for adjusting the buffer length based on a non-speech control strategy in the embodiments of the present application;
[0070] Figure 13 Schematic diagram of the process for adjusting the buffer length based on a speech control strategy in the embodiments of the present application;
[0071] Figure 14Flowchart of the method for determining the buffer length in the embodiments of the present application;
[0072] Figure 15 Schematic flowchart of the method for determining the third length in the embodiments of the present application;
[0073] Figure 16 Flowchart of the method for selecting the buffer length in the embodiments of the present application;
[0074] Figure 17 Logical schematic diagram of a method for determining the buffer length in the embodiments of the present application;
[0075] Figure 18 Example diagram of the method for adjusting the buffer length in the embodiments of the present application;
[0076] Figure 19 Schematic diagram of the detection process deployed at the sending end in the embodiments of the present application;
[0077] Figure 20 Schematic diagram of the detection process deployed at the receiving end in the embodiments of the present application;
[0078] Figure 21 Structural schematic diagram of a device for adjusting the buffer length in the embodiments of the present application;
[0079] Figure 22 Structural schematic diagram of an electronic device provided by the embodiments of the present application;
[0080] Figure 23 Another structural schematic diagram of an electronic device in the embodiments of the present application. Detailed implementation manners
[0081] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0082] For the convenience of those skilled in the art to better understand the technical solutions of the present application, some concepts involved in the present application are introduced below.
[0083] Jitter value: During the transmission of an audio signal, the transmission time interval for each audio data packet sent by the sender is the same. That is to say, the sender sends each audio data packet evenly. However, due to various network problems, such as congestion, packet loss, network errors, etc., the time interval between each audio data packet received by the receiver may be different, and may suddenly become larger or smaller, thus causing the transmission delay to change. The jitter value is the measure of the degree of change in the transmission delay.
[0084] Buffer: Also known as the jitter buffer, it is an important module in real-time audio and video applications. The buffer is used to handle situations such as the loss, out-of-order arrival, and delayed arrival of received audio data packets, and smoothly send audio data packets to the receiver.
[0085] The core idea of the buffer is to increase the delay from the sender to the receiver to improve the smoothness of the audio and video call. When the transmission network is unstable and jittery, for example, when an abnormally large number of audio data packets are received in a short period of time, or when the received audio data packets are out of order, the buffer increases the buffer length so that the buffer has enough cache space to receive more audio data packets, avoiding the problem that the incoming audio data packets are forced to be discarded due to insufficient buffer length. Then, the audio data packets in the buffer are re-ordered and other processing is performed so that the received audio data packets can be smoothly output to the decoding end, and thus the decoded audio signal can be played smoothly. When the number of received audio data packets returns to normal, the buffer will return to the normal buffer length to avoid introducing additional end-to-end delay.
[0086] Audio data packet: An audio data packet is obtained by encoding an audio frame of an audio signal.
[0087] Speech detection result: The speech detection result indicates whether the corresponding audio data packet contains speech data, that is, the audio information volume of the corresponding audio data packet. When the speech detection result is that it contains speech data, it is determined that there is valid information during the call, and the audio frame corresponding to the audio data packet is a speech frame. When the speech detection result is that it does not contain speech data, it is determined that there is no valid information during the call, and the audio frame corresponding to the audio data packet is a non-speech frame. The non-speech frame plays a certain role in the transition between speech frames. Therefore, the audio information volume of the speech frame is higher than that of the non-speech frame.
[0088] Speech rate detection result: The speech rate detection result is obtained by detecting the speech content contained in the audio signal per unit time, and can be divided into low speech rate, medium speech rate, and high speech rate. The higher the speech rate, the more speech content is contained per unit time, and the higher the information density.
[0089] Buffer length: It represents the number of audio data packets that the buffer can store. The longer the buffer length, the more audio data packets can be stored; the shorter the buffer length, the fewer audio data packets can be stored.
[0090] Pitch period value: When a person is speaking, according to the different vibration modes of the vocal cords, the sound signal is divided into voiceless and voiced sounds. Among them, voiceless sounds do not require periodic vibration of the vocal cords, while voiced sounds require periodic vibration of the vocal cords. Therefore, for voiced sounds, they have obvious periodicity, and the period of vocal cord vibration is the pitch period value.
[0091] Pitch period state: The pitch period state represents the change state of the pitch period value of the current audio data packet relative to the pitch period value of the previous audio data packet. The pitch period state can be divided into three states: "rising", "flat", and "falling".
[0092] The terms "first" and "second" in the text are only used for descriptive purposes and cannot be understood as explicitly or implicitly indicating relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.
[0093] The following briefly introduces the design concept of the embodiments of the present application:
[0094] Currently, with the development of network technology, usually the time intervals for the sending end to send each audio data packet are the same. However, during the process of transmitting each audio data packet of the audio signal to the receiving end, due to various network problems such as congestion and network errors, the time intervals of each audio data packet received by the receiving end may be different, resulting in problems such as packet loss and delayed arrival, thereby reducing the playback quality of the decoded audio signal.
[0095] To solve the above problems, after the receiving end receives each audio data packet, each audio data packet can be stored in a buffer for caching, re-sorted and other processing on each audio data packet in the buffer, and sent to the decoding end for decoding at the same time interval, so that the decoded audio signal can be played smoothly.
[0096] In actual application processes, buffers are usually divided into two types, one is a static buffer, and the other is a dynamic buffer:
[0097] Static buffer: The static buffer adopts a fixed buffer length and can counter jitter below the buffer length. For example, since the line is stable in some landline applications, a fixed buffer length is used. There is a fixed delay in the static buffer length. However, when the buffer length is set too long, unnecessary signal delay will be increased; when the buffer length is set too short, buffer overload will occur, resulting in packet loss and causing the problem of stuttering in the call sound.
[0098] Dynamic buffer: The buffer length is dynamically adjusted according to the detected jitter value of the audio packet. For example, when the detected jitter value changes, when the jitter value increases, the buffer length is increased, and when the jitter value decreases, the buffer length is decreased.
[0099] However, the method of statically adjusting the buffer length in the related art is not applicable to non-stable network transmission scenarios, such as Voice over Internet Protocol (VoIP), Internet live broadcast, radio, etc. When there is a large network jitter, it is easy to cause problems such as stuttering in the sound and damage to the sound quality; for the method of dynamically adjusting the buffer length, if the buffer length is immediately reduced after detecting that the jitter value decreases, buffer overload may occur because the buffer cannot store the received audio packets, resulting in packet loss. If the buffer length is immediately increased after detecting that the transmission delay change value increases, the signal delay during audio playback may be large because the storage space of the buffer is too large, thereby reducing the smoothness of the played audio signal.
[0100] Therefore, the accuracy of this method of adjusting the buffer length in the related art is not high.
[0101] In view of this, the embodiments of the present application propose a method, device, electronic device and storage medium for adjusting the buffer length. After determining the jitter value corresponding to each audio packet, based on each jitter value, the voice detection result and the speech rate detection result of each audio frame obtained are combined with the corresponding buffer length adjustment strategy to correspondingly adjust the current buffer length of the buffer. In this way, based on each jitter value, the voice detection result and the speech rate detection result of each audio frame, and the corresponding buffer length adjustment strategy, the current buffer length is dynamically adjusted, which can minimize the buffer length as much as possible in the absence of voice data, thereby reducing the end-to-end delay problem. In the case of a fast speech rate of the audio signal, a larger buffer redundancy space is reserved to avoid the overflow of audio packets in the buffer and the loss of audio signals under the transient impact of audio packets during network jitter, thereby improving the smoothness of the audio signal playback.
[0102] The preferred embodiments of the present application will be described below in conjunction with the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0103] As Figure 1 shown, it is a schematic diagram of an application scenario in an embodiment of the present application. The application scenario diagram includes a sending end 110 and a receiving end 120, and the sending end 110 and the receiving end 120 can communicate through a communication network.
[0104] In an alternative embodiment, the communication network can be a wired network or a wireless network.
[0105] In the embodiment of the present application, the sending end 110 and the receiving end 120 are electronic devices used by users, and the electronic devices include but are not limited to personal computers, mobile phones, tablet computers, notebooks, e-book readers, intelligent voice interaction devices, intelligent household appliances, vehicle-mounted terminals and other devices.
[0106] It should be noted that the method for adjusting the buffer length in the embodiment of the present application can be executed separately by the sending end or the receiving end, or can be jointly executed by the sending end and the receiving end. When jointly executed by the sending end and the receiving end, for example, the sending end can send audio data packets to the receiving end, and then the receiving end performs subsequent processing. In the following, mainly taking the receiving end executing alone as an example for illustration, and no specific limitation is made here.
[0107] In specific implementation, the receiving end can receive each audio data packet of the audio signal, and then process the audio data packet by using the method for adjusting the buffer length in the embodiment of the present application to adjust the current buffer length of the buffer.
[0108] Next, in combination with the above-described application scenario, the method for adjusting the buffer length provided by the exemplary embodiment of the present application will be described with reference to the drawings. It should be noted that the above application scenario is only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard. And the embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, signal transmission, voice call, etc.
[0109] Refer to Figure 2 shown, it is a flowchart of the implementation of a method for adjusting the buffer length in an embodiment of the present application. Here, taking the receiving end as the execution entity as an example for introduction, the specific implementation process of the method is as follows:
[0110] S20: Receive each audio data packet sent by the sending end.
[0111] Wherein, each audio data packet is obtained by encoding at least one audio frame of the audio signal.
[0112] In the embodiment of the present application, the sending end frames the audio signal to obtain each audio frame of the audio signal, encodes at least one audio frame of the audio signal respectively to obtain corresponding audio data packets, and sends each encoded audio data packet to the receiving end through a communication network at a preset time interval, so that the receiving end receives each audio data packet sent by the sending end.
[0113] It should be noted that the time interval when the sending end sends each audio data packet is fixed. For example, the time interval between the sending of the first audio data packet and the second audio data packet by the sending end is the same as the time interval between the second audio data packet and the third audio data packet.
[0114] In addition, it should be noted that in the embodiment of the present application, each audio data packet may be obtained by encoding one audio frame of the audio signal or may be obtained by encoding multiple audio frames of the audio signal. The embodiment of the present application does not limit this.
[0115] S21: Obtain the jitter values detected when receiving each audio data packet respectively.
[0116] Wherein, the jitter value characterizes the degree of change in the time interval for the receiving end to receive each audio data packet.
[0117] In the embodiment of the present application, jitter detection is performed on each audio data packet respectively, so as to obtain the jitter values detected when receiving each audio data packet.
[0118] Optionally, in the embodiment of the present application, when executing S21, it is necessary to obtain the jitter values detected when receiving each audio data packet respectively. Specifically, taking any one audio data packet (hereinafter referred to as audio data packet i) as an example, the process of obtaining the jitter value is introduced as follows. Refer to Figure 3 As shown, it is a schematic flowchart of obtaining the jitter value in the embodiment of the present application. The following combines the attached Figure 3 The process of obtaining each jitter value respectively in the embodiment of the present application is described in detail:
[0119] S211: Based on the reception time and transmission time of audio data packet i, and the reception time and transmission time of the previous audio data packet i - 1 of audio data packet i, determine the time difference between audio data packet i and the previous audio data packet i - 1.
[0120] In the embodiments of the present application, since each audio data packet corresponds to a transmission time and a reception time respectively, therefore, determine the reception time and the transmission time of the audio data packet i, and determine the reception time and the transmission time of the previous audio data packet i - 1 of the audio data packet i, and based on the reception time and the transmission time of the audio data packet i, determine the transmission time of the audio data packet i, based on the reception time and the transmission time of the audio data packet i - 1, determine the transmission time of the audio data packet i - 1, then, calculate the difference between the transmission time of the audio data packet i and the transmission time of the audio data packet i - 1, and obtain the time difference between the audio data packet i and the audio data packet i - 1.
[0121] For example, the time difference can be expressed as:
[0122] d(i,i - 1) = (r(i) - r(i - 1)) - (s(i) - s(i - 1)) = (r(i) - s(i)) - (r(i - 1) - s(i - 1))
[0123] Wherein, d(i,i - 1) represents the time difference between the audio data packet i and the audio data packet i - 1, r(i) represents the time when the audio data packet i arrives at the receiving end, s(i) represents the time when the sending end sends the audio data packet i, r(i - 1) represents the time when the audio data i - 1 arrives at the receiving end, and s(i - 1) represents the time when the sending end sends the audio data packet i - 1.
[0124] It should be noted that the units of the reception time and the transmission time are both the sampling rate.
[0125] S212: Determine the detected jitter value when receiving the audio data packet i based on the time difference, the smoothing coefficient, and the jitter value corresponding to the audio data packet i - 1.
[0126] In the embodiments of the present application, after determining the time difference, use the jitter value corresponding to the audio data packet i - 1, the time difference between the audio data packet i and the audio data packet i - 1, and the smoothing coefficient to determine the jitter value when receiving the audio data packet i.
[0127] For example, the jitter value can be expressed as:
[0128]
[0129] Wherein, jitter_value(i) represents the jitter value of the audio data packet i, jitter_value(i - 1) represents the jitter value of the audio data packet i - 1, |d(i - 1,i)| represents the absolute value of the time difference between the audio data packet i and the audio data packet i - 1, and x is the smoothing coefficient.
[0130] Among them, based on the above formula, it can be known that if the interval of the audio data packets received by the receiving end is the same as the sending interval of the sending end, the jitter value is 0, and the smoothing coefficient can be determined based on empirical values. For example, it can be This application embodiment does not limit this.
[0131] It should be noted that the audio data packets are sent based on the Request For Comments (RFC) 3550 communication protocol, and the sent audio data packets are Real-time Transport Protocol (RTP) packets. Therefore, the audio data packet i-1 refers to the previous received audio data packet, rather than being counted according to the RTP sequence number.
[0132] In this way, in the embodiment of this application, based on the smoothing coefficient, the jitter value corresponding to the previous audio data packet, and the time difference, the jitter value when receiving the current audio data packet is calculated, which can eliminate the influence of noise, make the jitter converge within a more reasonable range, and avoid the influence of burst data.
[0133] S22: Obtain the voice detection result corresponding to each audio frame and the speech rate detection result of the audio signal.
[0134] Among them, the voice detection result indicates whether the corresponding audio frame contains voice data.
[0135] In the embodiment of this application, after determining the jitter value corresponding to each audio data packet, obtain the voice detection result corresponding to each audio frame, and obtain the speech rate detection result of the audio signal.
[0136] First, the method for obtaining the speech rate detection result corresponding to each audio frame in the embodiment of this application will be described in detail.
[0137] In the embodiment of this application, two possible implementation manners are provided for obtaining the voice detection result corresponding to each audio frame, specifically including:
[0138] The first method: Obtain the voice detection result from the audio data packet.
[0139] In the embodiment of this application, refer to Figure 4As shown in the figure, it is a schematic flowchart of a method for obtaining a voice detection result in an embodiment of the present application. After the sending end obtains an audio signal, it divides the audio signal into frames at a preset fixed time interval to obtain each audio frame, and respectively performs voice activity detection on each audio frame to determine the voice detection result corresponding to each audio frame. Then, it respectively performs audio encoding on each audio frame to obtain an audio data packet corresponding to each audio frame. Then, it packages the voice detection result corresponding to each audio frame with the corresponding audio data packet, and sends each audio data packet to the receiving end at a fixed time interval through a communication network. After the receiving end receives each audio data packet, it respectively parses each audio data packet to obtain the corresponding voice detection result from each audio data packet.
[0140] The second method: The receiving end recognizes the voice detection result of the audio frame.
[0141] In the embodiment of the present application, refer to Figure 5 As shown in the figure, it is another schematic flowchart of a method for obtaining a voice detection result in an embodiment of the present application. After the sending end obtains an audio signal, it divides the audio signal into frames at a preset fixed time interval to obtain each audio frame, and respectively performs audio encoding on each audio frame to obtain an audio data packet corresponding to each audio frame. Then, it sends each audio data packet to the receiving end at a fixed time interval through a communication network. After the receiving end receives each audio data packet, it respectively decodes each audio data packet to obtain the audio frame corresponding to each audio data packet. Finally, it respectively performs voice activity detection on each audio frame to determine the voice detection result corresponding to each audio frame.
[0142] It should be noted that the voice activity detection method in the embodiment of the present application can be, for example, voice activity detection (VAD). Whether each audio frame contains voice data is identified through VAD. That is, if the result of VAD is 1, it is determined that the audio frame contains voice data; if the VAD result is 0, it is determined that the audio frame does not contain voice data. An audio frame not containing voice data indicates that the audio frame is a silent or noise signal. The embodiment of the present application does not limit the method for obtaining the voice detection result of the audio frame.
[0143] In addition, it should be noted that the preset time interval can be, for example, 20 ms, that is, every 20 ms is divided into 1 frame.
[0144] Secondly, two possible implementation methods are provided for obtaining the speech rate detection result of the audio signal in the embodiment of the present application, specifically including:
[0145] The first method: Obtain the speech rate detection result from the audio data packet.
[0146] In the embodiments of the present application, refer to Figure 6 As shown, it is a schematic flowchart of a process for obtaining a speech rate detection result in the embodiments of the present application. After the sending end obtains an audio signal, it performs a speech rate detection on the audio signal to obtain the speech rate detection result of the audio signal. At the same time, according to a preset fixed time interval, the audio signal is framed to obtain each audio frame, and each audio frame is respectively subjected to audio encoding to obtain an audio data packet corresponding to each audio frame. Then, the speech rate detection result is packed into any one of the audio data packets, and through a communication network, each audio data packet is sent to the receiving end at a fixed time interval. After the receiving end receives each audio data packet, it respectively parses each audio data packet to obtain the audio frame corresponding to each audio data packet and the speech rate detection result of the audio signal.
[0147] It should be noted that in the embodiments of the present application, the speech rate detection result can also be respectively packed into each audio data packet, and the embodiments of the present application do not limit this.
[0148] The second method: The receiving end performs a speech rate detection on the audio signal.
[0149] In the embodiments of the present application, refer to Figure 7 As shown, it is another schematic flowchart of a process for obtaining a speech rate detection result in the embodiments of the present application. After the sending end obtains an audio signal, according to a preset fixed time interval, the audio signal is framed to obtain each audio frame, and each audio frame is respectively subjected to audio encoding to obtain an audio data packet corresponding to each audio frame. Then, through a communication network, each audio data packet is sent to the receiving end at a fixed time interval. After the receiving end receives each audio data packet, it respectively decodes each audio data packet to obtain the audio frame corresponding to each audio data packet. Finally, based on each audio frame, the speech rate detection result of the audio signal is determined.
[0150] Optionally, in the embodiments of the present application, a possible implementation manner is provided for the receiving end to perform a speech rate detection on the audio signal. Refer to Figure 8 As shown, it is a schematic flowchart of a process for performing a speech rate detection on the audio signal in the embodiments of the present application, specifically including:
[0151] S221: Determine, from each audio frame, each target audio frame whose speech detection result is included in the speech data.
[0152] In the embodiments of the present application, since the speech detection result corresponding to each audio frame is included in the speech data or not included in the speech data, therefore, based on the speech detection result corresponding to each audio frame, from each audio frame, the audio frames whose speech detection result is included in the speech data are screened out as the target audio frames.
[0153] For example, assume that there are a total of 10 audio frames, and the speech detection results of these 10 audio frames are 0011101001 respectively, where 0 indicates that the speech detection result of the audio frame does not contain speech data, and 1 indicates that the speech detection result of the audio frame contains speech data. Then, from the above 10 audio frames, the audio frames with the speech detection result of containing speech data are selected, which are the 3rd, 4th, 5th, 7th, and 10th audio frames respectively, and the 3rd, 4th, 5th, 7th, and 10th audio frames are used as target audio frames.
[0154] S222: Obtain the pitch period state corresponding to each target audio frame.
[0155] In the embodiments of the present application, each target audio frame is recognized respectively to obtain the pitch period state corresponding to each target audio frame.
[0156] Optionally, in the embodiments of the present application, when executing S22-1, it is necessary to obtain the pitch period state corresponding to each target audio frame. Specifically, taking any one target audio frame (hereinafter referred to as target audio frame b) as an example, the process of obtaining the pitch period state is introduced as follows. Refer to Figure 9 shown, which is a schematic flow chart of the method for obtaining the pitch period state in the embodiments of the present application. The following combines the attached Figure 9 , and details the process of obtaining the pitch period state respectively in the embodiments of the present application:
[0157] S2221: Perform pitch period detection on target audio frame b to obtain the pitch period value corresponding to target audio frame b.
[0158] In the embodiments of the present application, a preset pitch period detection method is used to perform pitch period detection on target audio frame b, so as to obtain the pitch period value corresponding to target audio frame b.
[0159] Among them, the preset pitch period detection method can be, for example, pitch period detection based on autocorrelation, or, for example, pitch period detection based on linear predictive coding. The embodiments of the present application do not limit this.
[0160] S2222: Determine the pitch period value difference based on the pitch period value of target audio frame b and the pitch period value of the previous target audio frame b-1 corresponding to target audio frame b.
[0161] In the embodiments of the present application, obtain the pitch period value of the previous target audio frame b-1 corresponding to target audio frame b, and subtract the pitch period value of the previous target audio frame b-1 from the audio period value of target audio frame b to obtain the pitch period value difference between target audio frame b and the previous target audio frame b-1.
[0162] S2223: Determine the pitch period state corresponding to the target audio frame b according to the pitch period value difference and the difference threshold value.
[0163] In the embodiments of the present application, each target audio data packet can be divided into three pitch period states: "rising", "flat", and "falling". Specifically, a pitch difference threshold value is preset in advance. After determining the pitch period value difference corresponding to the target audio frame b, it is determined whether the pitch period value difference is greater than the preset difference threshold value. Specifically, it can be divided into the following three cases:
[0164] The first case: The pitch period value difference is less than the preset difference threshold value.
[0165] In the embodiments of the present application, if it is determined that the pitch period value difference is less than the preset difference threshold value, that is, the pitch period value of the target audio frame b is equal to or slightly different from the pitch period value of the previous target audio frame b-1, then it is determined that the pitch period state corresponding to the audio frame b is "flat".
[0166] The second case: The pitch period value difference is not less than the preset difference threshold value, and the pitch period value of the target audio frame b is greater than the pitch period value of the previous target audio frame b-1.
[0167] In the embodiments of the present application, if it is determined that the pitch period value difference is not less than the preset difference threshold value, then it is determined whether the pitch period value of the target audio frame b is greater than the pitch period value of the previous target audio frame b-1. If it is determined that the pitch period value of the target audio frame b is greater than the pitch period value of the previous target audio frame b-1, then it is determined that the pitch period state corresponding to the target audio frame b is "rising", that is, if it is determined that the pitch period value of the target audio frame b is greater than the pitch period value of the previous target audio frame b-1, and the pitch period value difference is greater than the preset difference threshold value, then it is determined that the pitch period state corresponding to the target audio frame b is "rising".
[0168] The third case: The pitch period value difference is not less than the preset difference threshold value, and the pitch period value of the target audio frame b is less than the pitch period value of the previous target audio frame b-1.
[0169] In the embodiments of the present application, if it is determined that the pitch period value difference is not less than the preset difference threshold value, then it is determined whether the pitch period value of the target audio frame b is greater than the pitch period value of the previous target audio frame b-1. If it is determined that the pitch period value of the target audio frame b is less than the pitch period value of the previous target audio frame b-1, then it is determined that the pitch period state corresponding to the target audio frame b is "falling", that is, if it is determined that the pitch period value of the target audio frame b is less than the pitch period value of the previous target audio frame b-1, and the pitch period value difference is greater than the preset difference threshold value, then it is determined that the pitch period state corresponding to the target audio frame b is "falling".
[0170] In this way, by using the difference value of the pitch period and a preset difference threshold value to determine the pitch period state of the target audio frame, the accuracy of determining the pitch period state can be improved.
[0171] S223: Determine the number of state transitions of the audio signal based on each pitch period state.
[0172] In the embodiments of the present application, since each target audio frame is divided into three pitch period states: "ascending", "flat", and "descending", therefore, adjacent target audio frames in the same pitch period state are counted to obtain a pitch period state statistical result, and based on the pitch period state statistical result, the number of state transitions of the audio signal is determined.
[0173] For example, referring to Figure 10 shown in the figure, which is an example diagram for determining the number of state transitions of the audio signal in the embodiments of the present application. Assume that the pitch period states corresponding to ten adjacent target audio frames are 0000111122 respectively, where "0" represents that the pitch period state of the target audio frame is "ascending", "1" represents that the pitch period state of the target audio frame is "flat", and "2" represents that the pitch period state of the target audio frame is "descending". Therefore, the statistical result of the ten adjacent target audio frames is that the cumulative value of the pitch period state of continuous "ascending" is 4, that is, four consecutive 0s, the cumulative value of the pitch period state of continuous "flat" is 4, that is, four consecutive 1s, and the cumulative value of the pitch period state of continuous "descending" is 2, that is, two consecutive 2s. Therefore, the ten adjacent target audio frames have switched 3 times, which are "ascending", "flat", and "descending".
[0174] S224: Determine the speech rate detection result of the audio signal according to the number of state transitions, the number of audio frames corresponding to each target audio frame, and a preset speech rate threshold value.
[0175] In the embodiments of the present application, first, calculate the number of state transitions and the number of audio frames corresponding to each target audio frame to obtain a speech rate value.
[0176] For example, the speech rate value can be expressed as: rate_V = Cnt_P / Cnt_V.
[0177] Wherein, Cnt_V represents the number of audio frames corresponding to the target audio frame, that is, the number of audio frames containing speech data in the speech detection result, Cnt_P represents the number of state transitions of each pitch period state, and rate_V represents the speech rate value, which is used to approximately represent the speech rate situation of the audio signal.
[0178] Then, based on the speech rate value and a preset speech rate threshold value, determine the speech rate detection result of the audio signal.
[0179] Specifically, in the embodiments of the present application, when determining the speech rate detection result of an audio signal based on the speech rate value and a preset speech rate threshold value, it can be specifically divided into the following three cases:
[0180] The first case: The speech rate value is less than or equal to the first speech rate threshold value.
[0181] In the embodiments of the present application, if it is determined that the speech rate value is less than or equal to the first speech rate threshold value, then it is determined that the speech rate detection result is a low speech rate.
[0182] For example, assuming that the first speech rate threshold value is 0.08, if it is determined that the speech rate value rate_V of the audio signal is less than or equal to 0.08, then it is determined that the speech rate detection result of the audio signal is a low speech rate.
[0183] The second case: The speech rate value is greater than the first speech rate threshold value and less than or equal to the second speech rate threshold value.
[0184] In the embodiments of the present application, if it is determined that the speech rate value of the audio signal is greater than the first speech rate threshold value and less than or equal to the second speech rate threshold value, then it is determined that the speech rate detection result is a medium speech rate.
[0185] For example, assuming that the first speech rate threshold value is 0.08 and the second speech rate threshold value is 0.15, if it is determined that the speech rate value rate_V of the audio signal is between 0.08 and 0.15, then it is determined that the speech rate detection result of the audio signal is a medium speech rate.
[0186] It should be noted that in the embodiments of the present application, the first speech rate threshold value and the second speech rate threshold value are determined based on empirical values, and the first speech rate threshold value is less than the second speech rate threshold value. For example, the first speech rate threshold value is 0.08 and the second speech rate threshold value is 0.1. The embodiments of the present application do not limit this.
[0187] The third case: The speech rate value is greater than the second speech rate threshold value.
[0188] In the embodiments of the present application, if it is determined that the speech rate value of the audio signal is greater than the second speech rate threshold value, then it is determined that the speech rate detection result is a high speech rate.
[0189] For example, assuming that the second speech rate threshold value is 0.15, if it is determined that the speech rate value rate_V of the audio signal is higher than 0.15, then it is determined that the speech rate detection result of the audio signal is a high speech rate.
[0190] In this way, determining the speech rate detection result based on the number of state transitions of the pitch period state, the number of audio frames, and the speech rate threshold value can improve the accuracy of determining the speech rate detection result, and provide a more accurate speech rate detection result for subsequent adjustment of the buffer length.
[0191] S23: Based on the speech rate detection result, each voice detection result, and each jitter value, in combination with a preset buffer length adjustment strategy, correspondingly adjust the current buffer length of the buffer.
[0192] In the embodiments of the present application, based on the voice detection result, a corresponding buffer length adjustment strategy is determined, and based on the speech rate detection result and each jitter value, in combination with the determined buffer length adjustment strategy, the current buffer length of the buffer is correspondingly adjusted.
[0193] Optionally, in the embodiments of the present application, a possible implementation manner is provided for correspondingly adjusting the current buffer length of the buffer. Refer to Figure 11 As shown, it is a schematic flowchart of the method for adjusting the buffer length in the embodiments of the present application, which specifically includes:
[0194] S231: Based on the reception time corresponding to each audio data packet, determine the audio data packets that meet the reception time condition from each audio data packet.
[0195] In the embodiments of the present application, a reception time condition is preset in advance. Since the receiving end records the reception time corresponding to each received audio data packet after receiving each audio data packet, based on the recorded reception time corresponding to each audio data packet, the audio data packets that meet the reception time condition are screened out from each audio data packet.
[0196] Among them, the reception time condition can be, for example, the latest reception time. Then, from each audio data packet, the audio data packet with the latest reception time is determined, that is, the last received audio data packet is used as the audio data packet that meets the reception time condition.
[0197] S232: If it is determined that each voice detection result does not include voice data, based on the non-voice control strategy and the target jitter value of the determined audio data packet, correspondingly adjust the current buffer length of the buffer.
[0198] In the embodiments of the present application, if it is determined that each voice detection result does not include voice data, it is determined that the buffer length adjustment strategy is the non-voice control strategy, and based on the non-voice control strategy and the target jitter value of the determined audio data packet, the current buffer length of the buffer is correspondingly adjusted.
[0199] For example, assuming that the voice detection results corresponding to 5 audio frames are 00000 respectively, it is determined that the voice activity detection (VAD) results corresponding to each audio frame of the audio signal do not include voice data, and the corresponding buffer length adjustment strategy is determined as the non-voice control strategy.
[0200] Optionally, in the embodiments of the present application, a possible implementation manner is provided for adjusting the buffer length based on a non-voice control strategy. Refer to Figure 12 As shown, it is a schematic flowchart of the method for adjusting the buffer length based on a non-voice control strategy in the embodiments of the present application, specifically including:
[0201] S2321: If the determined target jitter value of the audio data packet is less than the jitter value threshold, then use the preset first length as the current buffer length of the buffer.
[0202] In the embodiments of the present application, it is judged whether the target jitter value of the audio data packet is less than the jitter value threshold. If it is determined that the target jitter value of the audio data packet is less than the jitter value threshold, then use the preset first length as the current buffer length of the buffer.
[0203] For example, the jitter value threshold is THRD_1. If the determined target jitter value jitter_value of the audio data packet is less than the jitter value threshold THRD_1, then obtain the preset first length MIN_LEN, and use the preset first length MIN_LEN as the current buffer length len of the buffer, that is, len = MIN_LEN, which is the minimum buffer length.
[0204] It should be noted that the jitter value threshold in the embodiments of the present application is related to the sampling rate.
[0205] S2322: If it is determined that the target jitter value is not less than the jitter value threshold, then select any target length from the preset length range as the buffer length.
[0206] Among them, the length range is generated based on the first length and the second length, and the first length is less than the second length.
[0207] In the embodiments of the present application, it is judged whether the target jitter value of the audio data packet is less than the jitter value threshold. If it is determined that the target jitter value of the audio data packet is not less than the jitter value threshold, then select any target length from the preset length range as the current buffer length of the buffer.
[0208] For example, if it is determined that the target jitter value jitter_value is greater than or equal to the jitter value threshold THRD_1, then determine that the buffer length is the default length, that is, len = DEFAULT_LEN. The default length DEFAULT_LEN can be between the second length MAX_LEN and the first length MIN_LEN allowed by the buffer jitterbuffer. It can take the median value of the two, or any length. This application does not limit this in the embodiments.
[0209] Among them, THRD_1 can be, for example, 500 ms, MIN_LEN can be, for example, 1000 ms, and MAX_LEN can be, for example, 2000 ms. In the embodiments of the present application, there is no limitation on this.
[0210] In this way, when the audio signal does not contain voice data, even if packet loss occurs during the process of receiving audio data packets, or the cached data packets are compressed, it will not affect the voice call quality. Therefore, when the audio signal does not contain voice data, minimizing the buffer length can reduce unnecessary call latency.
[0211] S233: If it is determined that at least one voice detection result contains voice data, then based on the speech rate detection result and the target jitter value, combined with the voice control strategy, the buffer length is adjusted accordingly.
[0212] In the embodiments of the present application, if it is determined that each voice detection result does not contain voice data, then it is determined that the buffer length adjustment strategy is the voice control strategy, and based on the voice control strategy, the determined target jitter value of the audio data packet and the speech rate detection result, the current buffer length of the buffer is adjusted accordingly.
[0213] For example, assuming that the voice detection results corresponding to 5 audio frames are 00100 respectively, then it is determined that the corresponding buffer length adjustment strategy is the voice control strategy.
[0214] Optionally, in the embodiments of the present application, a possible implementation manner is provided for adjusting the buffer length based on the voice control strategy. Refer to Figure 13 As shown, it is a schematic flowchart of the method for adjusting the buffer length based on the voice control strategy in the embodiments of the present application, specifically including:
[0215] S2331: If it is determined that the target jitter value is less than the jitter value threshold, then any target length is selected from the preset length range as the buffer length.
[0216] Among them, the length range is generated based on the first length and the second length, and the first length is less than the second length.
[0217] In the embodiments of the present application, it is judged whether the target jitter value of the audio data packet is less than the jitter value threshold. If it is determined that the target jitter value of the audio data packet is not less than the jitter value threshold, then any target length is selected from the preset length range as the current buffer length of the buffer.
[0218] For example, if it is determined that the target jitter value jitter_value is greater than or equal to the jitter value threshold THRD_1, the buffer length is determined to be the default length, that is, len = DEFAULT_LEN. The default length DEFAULT_LEN can be between the second length MAX_LEN and the first length MIN_LEN allowed by the buffer jitterbuffer. It can take the median of the two or any length. This application embodiment does not limit this.
[0219] Among them, the first length is determined based on the empirical value in the experimental process, and the second length is also determined based on the empirical value in the experimental process. The first length is the minimum length to ensure that the buffer will not be overloaded, and the second length is the maximum length to ensure that there will be no additional call delay in the buffer.
[0220] S2332: If it is determined that the target jitter value is not less than the jitter value threshold, determine the buffer length based on the speech rate detection result and the target jitter value.
[0221] In the embodiment of this application, if it is determined that the target jitter value is not less than the jitter value threshold, determine the buffer length based on the speech rate detection result and the target jitter value.
[0222] In this way, when the audio signal contains voice data, if packet loss occurs during the process of receiving audio data packets, or the cached data packets are compressed, it will affect the voice call quality. Therefore, when the audio signal contains voice data, determining the buffer length based on the speech rate detection result and the jitter value can improve the accuracy of adjusting the buffer length, thereby avoiding the problem of call sound stuttering caused by the loss of valid audio data due to buffer overload and improving the fluency of audio signal playback.
[0223] Optionally, in the embodiment of this application, a possible implementation manner is provided to determine the buffer length. Refer to Figure 14 As shown, it is a flowchart of the method for determining the buffer length in the embodiment of this application, which specifically includes:
[0224] S2332-1: Determine the speech rate adjustment parameter of the audio signal based on the speech rate detection result.
[0225] In the embodiment of this application, based on the speech rate detection result and the correlation relationship between the speech rate detection result and the speech rate adjustment parameter, the speech rate adjustment parameter associated with the speech rate detection result is determined.
[0226] For example, when the speech rate detection result is a low speech rate, the speech rate adjustment parameter a is 0.8; when the speech rate detection result is a medium speech rate, the speech rate adjustment parameter a is 1; when the speech rate detection result is a high speech rate, the speech rate adjustment parameter a is 1.2. Therefore, the value of the speech rate adjustment parameter changes accordingly according to the speech rate detection result.
[0227] S2332-2: Determine the third length based on the speech rate adjustment parameter, the target jitter value, and a preset jitter value adjustment function.
[0228] In the embodiments of the present application, the third length is determined based on the speech rate adjustment parameter, the target jitter value, and the jitter value adjustment function.
[0229] Optionally, in the embodiments of the present application, a possible implementation manner is provided for determining the third length. Refer to Figure 15 As shown, it is a schematic flowchart of the method for determining the third length in the embodiments of the present application, specifically including:
[0230] S2332-2-1: Use the target jitter value as the variable of the preset jitter value adjustment function, and obtain a fourth length based on the target jitter value and the jitter value adjustment function.
[0231] In the embodiments of the present application, the target jitter value is used as the variable of the jitter value adjustment function, that is, the jitter value adjustment function is a function with the target jitter value as the variable. Then, a fourth length is obtained based on the target jitter value and the jitter value adjustment function.
[0232] For example, the fourth length is f(jitter_value), where f(jitter_value) is a jitter value adjustment function with the target jitter value jitter_value as the variable.
[0233] It should be noted that the jitter value adjustment parameter in the embodiments of the present application can be a monotonically increasing function, and the present application does not limit this.
[0234] S2332-2-2: Obtain the third length based on the speech rate adjustment parameter and the fourth length.
[0235] Among them, the speech rate adjustment parameter is positively correlated with the fourth length.
[0236] In the embodiments of the present application, the product between the jitter value adjustment function with the target jitter value as the variable and the speech rate adjustment parameter is calculated to obtain the third length. Therefore, the third length is positively correlated with the fourth length.
[0237] For example, the third length is a*f(jitter_value), where f(jitter_value) is a jitter value adjustment function with the target jitter value jitter_value as the variable, and a is the speech rate adjustment parameter.
[0238] S2332-3: Select the target length that meets the length condition from the third length and the second length as the buffer length.
[0239] In the embodiments of the present application, it is determined whether the third length and the second length meet the preset length condition respectively, the target length that meets the preset length condition is selected from the third length and the second length, and the selected target length is used as the buffer length.
[0240] Optionally, in the embodiments of the present application, a possible implementation manner is provided for selecting the buffer length. Refer to Figure 16 As shown, it is a flowchart of the method for selecting the buffer length in the embodiments of the present application, which specifically includes:
[0241] S2332-3-1: If it is determined that the third length is greater than the second length, then the second length is used as the target length.
[0242] In the embodiments of the present application, it is judged whether the third length is greater than the second length. If it is determined that the third length is greater than the second length, then the second length is used as the target length.
[0243] For example, len = min(MAX_LEN, a * f(jitter_value)).
[0244] Wherein, len is the buffer length, MAX_LEN is the second length, and a * f(jitter_value) is the third length.
[0245] When it is determined that the third length a * f(jitter_value) is greater than the second length MAX_LEN, then the second length MAX_LEN is used as the target length, that is, the buffer length.
[0246] S2332-3-2: If it is determined that the third length is not greater than the second length, then the third length is used as the target length.
[0247] In the embodiments of the present application, it is judged whether the third length is greater than the second length. If it is determined that the third length is not greater than the second length, then the third length is used as the target length.
[0248] For example, len = min(MAX_LEN, a * f(jitter_value)).
[0249] Wherein, len is the buffer length, MAX_LEN is the second length, and a * f(jitter_value) is the third length.
[0250] When it is determined that the third length a*f(jitter_value) is not greater than the second length MAX_LEN, the third length a*f(jitter_value) is used as the target length, that is, the buffer length.
[0251] In this way, by means of the jitter value adjustment function with the target jitter value as a variable and the speech rate adjustment parameter, the buffer length is determined, which can improve the accuracy of determining the buffer length. Moreover, by selecting the smaller target length from the third length and the second length as the buffer length, it is possible to minimize the buffer length while ensuring that the buffer is not overloaded, thereby avoiding end-to-end call delays.
[0252] In the embodiments of the present application, by adjusting the current buffer length of the buffer according to the respective voice detection results, speech rate detection results, and various jitter values, it is possible to avoid introducing unnecessary end-to-end call delays into the buffer. At the same time, through the adjustment of the buffer length, it is possible to reduce the problem of call sound stuttering caused by the loss of valid audio data due to buffer overload, thereby improving the overall call quality and subjective experience.
[0253] Based on the above embodiments, the following refers to the flowchart of another method for determining the buffer length in the embodiments of the present application. Figure 17 As shown, it is a logical schematic diagram of a method for determining the buffer length in the embodiments of the present application, which specifically includes:
[0254] S170: Determine whether the VAD corresponding to each audio frame is 0. If so, execute S171; if not, execute S174.
[0255] In the embodiments of the present application, VAD = 0 indicates that the audio data packet does not contain voice data, and VAD = 1 indicates that the audio data packet contains voice data.
[0256] S171: Determine whether jitter_value is less than THRD_1. If so, execute S172; if not, execute S173.
[0257] In the embodiments of the present application, when it is determined that none of the audio data packets contain voice data, it is determined whether the target jitter value jitter_value is less than the preset jitter value threshold THRD_1. If it is determined that the target jitter value jitter_value is less than the preset jitter value threshold THRD_1, the preset first length is used as the jitter value threshold and the adjusted buffer length len; if it is determined that the target jitter value jitter_value is not less than the preset jitter value threshold THRD_1, a length is arbitrarily selected from the preset length range as the buffer length len.
[0258] S172: len = MIN_LEN.
[0259] Among them, len represents the buffer length, and MIN_LEN represents the first length, that is, the minimum buffer length.
[0260] S173: len = DEFAULT_LEN.
[0261] Among them, len represents the buffer length, and DEFAULT_LEN is between the maximum length MAX_LEN allowed by the target jitter value jitter_value and the first length MIN_LEN, and the intermediate value between the maximum length MAX_LEN and the first length MIN_LEN can be selected.
[0262] S174: Determine whether jitter_value is less than THRD_1. If so, execute S175; if not, execute S176.
[0263] In the embodiments of the present application, when it is determined that one of the audio data packets in each audio data packet contains voice data, it is determined whether the target jitter value jitter_value is less than the preset jitter value threshold THRD_1. If it is determined that the target jitter value jitter_value is less than the preset jitter value threshold THRD_1, a length is randomly selected from the preset length range as the buffer length len. If it is determined that the target jitter value jitter_value is not less than the preset jitter value threshold THRD_1, the third length a * f(jitter_value) is determined based on the jitter value adjustment function f(jitter_value) with the target jitter value jitter_value as the variable and the speech rate adjustment parameter a, and the target length with the smaller value is selected from the second length MAX_LEN and the third length a * f(jitter_value) as the buffer length.
[0264] S175: len = DEFAULT_LEN.
[0265] In the embodiments of the present application, len represents the buffer length, and DEFAULT_LEN is between the maximum length MAX_LEN allowed by the target jitter value jitter_value and the first length MIN_LEN, and the intermediate value between the maximum length MAX_LEN and the first length MIN_LEN can be selected.
[0266] S176: len = min(MAX_LEN, a * f(jitter_value)).
[0267] In the embodiments of the present application, len represents the buffer length, MAX_LEN represents the second length, f(jitter_value) represents the jitter value adjustment function with the target jitter value jitter_value as the variable, and a represents the speech rate adjustment parameter.
[0268] Based on the above embodiments, a specific example is used below to elaborate in detail on the method for adjusting the buffer length in the embodiments of the present application. Refer to Figure 18 As shown, it is an example diagram of the method for adjusting the buffer length in the embodiments of the present application, specifically including:
[0269] First, the sender frames the audio signal X to obtain each audio frame x1, x2, x3, x4, x5 of the audio signal X.
[0270] Then, the sender performs voice detection on the audio frame x1 and obtains a voice detection result of 0 for the audio frame x1, performs voice detection on the audio frame x2 and obtains a voice detection result of 0 for the audio frame x2, performs voice detection on the audio frame x3 and obtains a voice detection result of 1 for the audio frame x3, performs voice detection on the audio frame x4 and obtains a voice detection result of 1 for the audio frame x4, performs voice detection on the audio frame x5 and obtains a voice detection result of 0 for the audio frame x5, and performs speech rate detection on the audio signal to obtain a speech rate detection result of medium speech rate for the audio signal X.
[0271] Then, the sender performs audio encoding on the audio frame x1 to obtain the audio data packet A1 and packs the voice detection result 0 into the audio data packet A1, performs audio encoding on the audio frame x2 to obtain the audio data packet A2 and packs the voice detection result 0 into the audio data packet A2, performs audio encoding on the audio frame x3 to obtain the audio data packet A3 and packs the voice detection result 1 into the audio data packet A3, performs audio encoding on the audio frame x4 to obtain the audio data packet A4 and packs the voice detection result 1 into the audio data packet A4, performs audio encoding on the audio frame x5 to obtain the audio data packet A5 and packs the voice detection result 0 and the speech rate detection result of medium speech rate into the audio data packet A5.
[0272] The sender sends the audio data packets A1, A2, A3, A4, and A5 to the receiver at preset time intervals respectively.
[0273] After the receiving end receives audio data packets A1, A2, A3, A4, and A5, it reads the voice detection results in each audio data packet respectively, obtains the voice detection result 0 in audio data packet A1, the voice detection result 0 in audio data packet A2, the voice detection result 1 in audio data packet A3, the voice detection result 1 in audio data packet A4, the voice detection result 0 in audio data packet A5, and the speech rate in the speech rate detection result, and performs jitter detection on audio data packet A5 to obtain the target jitter value 0.65 of audio data packet A5.
[0274] Finally, based on the voice detection result 0, the voice detection result 0, the voice detection result 1, the voice detection result 1, the voice detection result 0, the speech rate in the speech rate detection result, and the target jitter value 0.65, combined with the buffer length adjustment strategy, the receiving end determines that the current buffer length of the buffer is 1500 ms.
[0275] Based on the above embodiments, the speech rate detection process and the voice detection process in the embodiments of the present application can be deployed at the sending end. Refer to Figure 19 As shown, it is a schematic diagram of the detection process deployed at the sending end in the embodiments of the present application, specifically including:
[0276] First, the sending end frames the audio signal to obtain each audio frame of the audio signal, and respectively performs voice detection on each audio frame to obtain the voice detection result corresponding to each audio frame, and performs speech rate detection to obtain the speech rate detection result of the audio signal. Then, each audio frame is encoded to obtain the audio data packet corresponding to each audio frame, and the respective voice detection results and speech rate detection results are respectively packed into the corresponding audio data packets, and each audio data packet is sent to the receiving end at a preset time interval.
[0277] After the receiving end receives each audio data packet, it respectively parses and obtains the corresponding voice detection result and speech rate detection result from each audio data packet, and respectively performs jitter detection on each audio data packet to obtain the jitter value corresponding to each audio data packet, and stores each audio data packet in the buffer. Then, each audio data packet in the buffer is decoded to obtain the audio frame corresponding to each audio data packet, and each audio frame is played. At the same time, based on the speech rate detection result, each voice detection result, and the target jitter value determined from each jitter value, the buffer length is adjusted.
[0278] Based on the above embodiments, the speech rate detection process and the voice detection process in the embodiments of the present application can be deployed at the receiving end. Refer to Figure 20 As shown, it is a schematic diagram of the detection process deployed at the receiving end in the embodiments of the present application, specifically including:
[0279] First, the sending end frames the audio signal to obtain each audio frame of the audio signal. Then, each audio frame is encoded respectively to obtain an audio data packet corresponding to each audio frame, and each audio data packet is sent to the receiving end respectively at a preset time interval.
[0280] After receiving each audio data packet, the receiving end performs jitter detection on each audio data packet respectively to obtain a jitter value corresponding to each audio data packet, and stores each audio data packet in a buffer. Then, each audio data packet in the buffer is decoded respectively to obtain an audio frame corresponding to each audio data packet, and each audio frame is played. Meanwhile, voice detection is performed on each parsed audio frame respectively to obtain a voice detection result corresponding to each audio frame, and speech rate detection is performed to obtain a speech rate detection result of the audio signal. Based on the speech rate detection result, each voice detection result, and a target jitter value determined from each jitter value, the buffer length is adjusted.
[0281] Based on the same inventive concept as the method embodiment of the present application above, an apparatus for adjusting the buffer length is further provided in the embodiment of the present application. The principle of the apparatus for solving the problem is similar to that of the method in the above embodiment. Therefore, the implementation of the apparatus can refer to the implementation of the above method, and the repeated parts will not be described again.
[0282] Reference Figure 21 As shown, a schematic structural diagram of an apparatus for adjusting the buffer length in the embodiment of the present application includes a receiving module 211, a jitter detection module 212, a processing module 213, and an adjustment module 214.
[0283] The receiving module 211 is configured to receive each audio data packet sent by the sending end, where each audio data packet is obtained by encoding at least one audio frame of the audio signal;
[0284] The jitter detection module 212 is configured to obtain the jitter value detected respectively when receiving each audio data packet;
[0285] The processing module 213 is configured to obtain the voice detection result corresponding to each audio frame and the speech rate detection result of the audio signal; where the voice detection result represents whether the corresponding audio frame contains voice data; the speech rate detection result is a result obtained by detecting the voice content contained in the audio signal per unit time;
[0286] The adjustment module 214 is configured to adjust the current buffer length of the buffer correspondingly based on the speech rate detection result, each voice detection result, and each jitter value, in combination with a preset buffer length adjustment strategy.
[0287] In a possible embodiment, the adjustment module 214 is further configured to:
[0288] Based on the reception time corresponding to each audio data packet, determine, from each audio data packet, the audio data packets that meet the reception time condition;
[0289] If it is determined that each voice detection result does not include voice data, then based on the non-voice control strategy and the target jitter value of the determined audio data packet, adjust the current buffer length of the buffer accordingly;
[0290] If it is determined that at least one voice detection result includes voice data, then based on the speech rate detection result and the target jitter value, and in combination with the voice control strategy, adjust the buffer length accordingly.
[0291] In a possible embodiment, when adjusting the current buffer length of the buffer based on the non-voice control strategy and the target jitter value of the determined audio data packet, the adjustment module 214 is further configured to:
[0292] If the target jitter value of the determined audio data packet is less than the jitter value threshold, then use the preset first length as the current buffer length of the buffer;
[0293] If it is determined that the target jitter value is not less than the jitter value threshold, then select any target length from the preset length range as the buffer length, where the length range is generated based on the first length and the second length, and the first length is less than the second length.
[0294] In a possible embodiment, when adjusting the buffer length based on the speech rate detection result and the target jitter value, and in combination with the voice control strategy, the adjustment module 214 is further configured to:
[0295] If it is determined that the target jitter value is less than the jitter value threshold, then select any target length from the preset length range as the buffer length, where the length range is generated based on the first length and the second length, and the first length is less than the second length;
[0296] If it is determined that the target jitter value is not less than the jitter value threshold, then determine the buffer length based on the speech rate detection result and the target jitter value.
[0297] In a possible embodiment, when determining the buffer length based on the speech rate detection result and the target jitter value, the adjustment module 214 is further configured to:
[0298] Based on the speech rate detection result, determine the speech rate adjustment parameter of the audio signal;
[0299] Based on the speech rate adjustment parameter, the target jitter value, and the preset jitter value adjustment function, determine the third length;
[0300] Select, from the third length and the second length, the target length that meets the length condition as the buffer length.
[0301] In a possible embodiment, based on the speech rate adjustment parameter, the target jitter value, and a preset jitter value adjustment function, the third length is determined. The processing module 213 is further configured to:
[0302] Use the target jitter value as a variable of the preset jitter value adjustment function, and based on the target jitter value and the jitter value adjustment function, obtain a fourth length;
[0303] Based on the speech rate adjustment parameter and the fourth length, obtain the third length, where the speech rate adjustment parameter is positively correlated with the fourth length.
[0304] In a possible embodiment, when selecting a target length that meets the length condition from the third length and the second length, the processing module 213 is further configured to:
[0305] If it is determined that the third length is greater than the second length, use the second length as the target length;
[0306] If it is determined that the third length is not greater than the second length, use the third length as the target length.
[0307] In a possible embodiment, when obtaining the speech rate detection result of the audio signal, the processing module 213 is further configured to:
[0308] From each audio frame, determine each target audio frame whose voice detection result includes voice data;
[0309] Obtain the pitch period state corresponding to each target audio frame;
[0310] Based on each pitch period state, determine the number of state transitions of the audio signal;
[0311] According to the number of state transitions, the number of audio frames corresponding to each target audio frame, and a preset speech rate threshold value, determine the speech rate detection result of the audio signal.
[0312] In a possible embodiment, when obtaining the pitch period state corresponding to each target audio frame, the processing module 213 is further configured to:
[0313] For each target audio frame, perform the following operations respectively:
[0314] Perform pitch period detection on a target audio frame to obtain the pitch period value corresponding to the target audio frame;
[0315] Based on the pitch period value of a target audio frame and the pitch period value of the previous target audio frame corresponding to the target audio frame, determine the pitch period value difference;
[0316] Determine the pitch period state corresponding to a target audio frame according to the pitch period value difference and the difference threshold value.
[0317] In a possible embodiment, when determining the jitter value corresponding to each audio data packet respectively, the jitter detection module 212 is further configured to:
[0318] For each audio data packet, perform the following operations respectively:
[0319] Based on the reception time and transmission time of an audio data packet, and the reception time and transmission time of the previous audio data packet of an audio data packet, determine the time difference between an audio data packet and the previous audio data packet;
[0320] Based on the time difference, the smoothing coefficient, and the jitter value corresponding to the previous audio data packet, obtain the jitter value corresponding to an audio data packet.
[0321] For the convenience of description, the above parts are divided into each module (or unit) according to functions and described separately. Of course, when implementing the present application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.
[0322] After introducing the method and device for adjusting the buffer length according to the exemplary embodiments of the present application, next, introduce the device for adjusting the buffer length according to another exemplary embodiment of the present application.
[0323] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0324] In some possible embodiments, the device for adjusting the buffer length according to the present application may at least include a processor and a memory. Wherein, the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps in the method for adjusting the buffer length according to various exemplary embodiments of the present application described in this specification. For example, the processor may execute as Figure 2 shown in the steps.
[0325] After introducing the method and device for adjusting the buffer length according to the exemplary embodiments of the present application, next, introduce the electronic device according to another exemplary embodiment of the present application.
[0326] Based on the same inventive concept as the method embodiments of the present application, an electronic device is further provided in the embodiments of the present application. The principle of the electronic device to solve problems is similar to that of the method in the above embodiments. Therefore, the implementation of the electronic device can refer to the implementation of the above method, and the repeated parts will not be described again.
[0327] Refer to Figure 22 As shown, the electronic device 220 may at least include a processor 221 and a memory 222. Among them, the memory 222 stores program code. When the program code is executed by the processor 221, the processor 221 is caused to execute the steps in any of the above methods for adjusting the buffer length.
[0328] In some possible implementation manners, the electronic device according to the present application may at least include at least one processor and at least one memory. Among them, the memory stores program code. When the program code is executed by the processor, the processor is caused to execute the steps in the method for adjusting the buffer length according to various exemplary implementation manners of the present application described above in this specification. For example, the processor may execute as Figure 2 the steps shown in
[0329] In an exemplary embodiment, the present application further provides a storage medium including program code, such as the memory 222 including program code. The above program code can be executed by the processor 221 of the electronic device 220 to complete the above method for adjusting the buffer length. Optionally, the storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0330] Next, refer to Figure 23 to describe the electronic device 230 according to this implementation manner of the present application. Figure 23 The electronic device 230 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0331] As Figure 23 shown, the electronic device 230 is presented in the form of a general electronic device. The components of the electronic device 230 may include but are not limited to: the above at least one processing unit 231, the above at least one storage unit 232, and a bus 233 connecting different system components (including the storage unit 232 and the processing unit 231).
[0332] The bus 233 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a processor, or a local bus using any bus structure in a variety of bus structures.
[0333] The storage unit 232 may include a readable medium in the form of volatile memory, such as a random access memory (RAM) 2321 and / or a cache storage unit 2322, and may further include a read-only memory (ROM) 2323.
[0334] The storage unit 232 may also include a program / utility 2325 having a set (at least one) of program modules 2324. Such program modules 2324 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0335] The electronic device 230 may also communicate with one or more external devices 234 (such as a keyboard, a pointing device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 230, and / or may communicate with any device that enables the electronic device 230 to communicate with one or more other electronic devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 235. Also, the electronic device 230 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 236. As shown in the figure, the network adapter 236 communicates with other modules for the electronic device 230 through a bus 233. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 230, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0336] In some possible implementation manners, various aspects of the method for adjusting a buffer length provided in this application may also be implemented in the form of a program product, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps in the method for adjusting a buffer length according to various exemplary implementation manners of this application described above in this specification. For example, the electronic device may execute the steps as Figure 2 shown in.
[0337] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0338] The program product of the embodiments of the present application can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can run on a computing device. However, the program product of the present application is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with a command execution system, apparatus, or device.
[0339] The readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with a command execution system, apparatus, or device.
[0340] The program code contained on the readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0341] The program code for performing the operations of the present application can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages - such as Java, C++, etc., and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0342] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0343] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0344] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0345] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.
[0346] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. A method for adjusting the buffer length, characterized in that, including: receiving each audio data packet sent by a sending end, where each audio data packet is obtained by encoding at least one audio frame of an audio signal; respectively obtaining jitter values detected when receiving the audio data packets; acquiring a voice detection result corresponding to each audio frame and a speech rate detection result of the audio signal; wherein, the voice detection result indicates whether the corresponding audio frame contains voice data; the speech rate detection result is a result obtained by detecting the voice content included in the audio signal per unit time; based on the speech rate detection result, each voice detection result, and each jitter value, and in combination with a preset buffer length adjustment strategy, correspondingly adjusting the current buffer length of the buffer.
2. The method according to claim 1, wherein The correspondingly adjusting the current buffer length of the buffer based on the speech rate detection result, each voice detection result, and each jitter value, and in combination with a preset buffer length adjustment strategy includes: based on the reception time corresponding to each audio data packet, determining, from the audio data packets, audio data packets that meet the reception time condition; if it is determined that each voice detection result does not include voice data, correspondingly adjusting the current buffer length of the buffer based on a non-voice control strategy and the target jitter value of the determined audio data packets; if it is determined that at least one voice detection result includes voice data, correspondingly adjusting the buffer length based on the speech rate detection result and the target jitter value, in combination with a voice control strategy.
3. The method according to claim 2, characterized in that, The correspondingly adjusting the current buffer length of the buffer based on a non-voice control strategy and the target jitter value of the determined audio data packets includes: if the target jitter value of the determined audio data packet is less than a jitter value threshold, taking a preset first length as the current buffer length of the buffer; if it is determined that the target jitter value is not less than the jitter value threshold, selecting any target length from a preset length range as the buffer length, where the length range is generated based on the first length and a second length, and the first length is less than the second length.
4. The method according to claim 2, wherein The correspondingly adjusting the buffer length based on the speech rate detection result and the target jitter value, in combination with a voice control strategy includes: if it is determined that the target jitter value is less than a jitter value threshold, selecting any target length from a preset length range as the buffer length, where the length range is generated based on a first length and a second length, and the first length is less than the second length; if it is determined that the target jitter value is not less than the jitter value threshold, determining the buffer length based on the speech rate detection result and the target jitter value.
5. The method according to claim 4, characterized in that, The determining the buffer length based on the speech rate detection result and the target jitter value includes: determining a speech rate adjustment parameter of the audio signal based on the speech rate detection result; determining a third length based on the speech rate adjustment parameter, the target jitter value, and a preset jitter value adjustment function; selecting a target length that meets the length condition from the third length and the second length as the buffer length.
6. The method according to claim 5, wherein Determining a third length based on the speech rate adjustment parameter, the target jitter value, and a preset jitter value adjustment function specifically includes: Using the target jitter value as a variable of the preset jitter value adjustment function, and obtaining a fourth length based on the target jitter value and the jitter value adjustment function; Obtaining a third length based on the speech rate adjustment parameter and the fourth length, where the speech rate adjustment parameter is positively correlated with the fourth length.
7. The method according to claim 6, characterized in that Selecting a target length that meets the length condition from the third length and the second length specifically includes: If it is determined that the third length is greater than the second length, then using the second length as the target length; If it is determined that the third length is not greater than the second length, then using the third length as the target length.
8. The method according to any one of claims 1-7, characterized in that, Obtaining the speech rate detection result of the audio signal includes: Determining, from each audio frame, each target audio frame whose speech detection result is that it contains speech data; Obtaining the pitch period state corresponding to each of the target audio frames; Determining the number of state transitions of the audio signal based on each pitch period state; Determining the speech rate detection result of the audio signal according to the number of state transitions, the number of audio frames corresponding to each of the target audio frames, and a preset speech rate threshold value.
9. The method according to claim 8, wherein Obtaining the pitch period state corresponding to each of the target audio frames specifically includes: Performing the following operations respectively for each of the target audio frames: Performing pitch period detection on a target audio frame to obtain the pitch period value corresponding to the target audio frame; Determining a pitch period value difference based on the pitch period value of the target audio frame and the pitch period value of the previous target audio frame corresponding to the target audio frame; Determining the pitch period state corresponding to the target audio frame according to the pitch period value difference and a difference threshold value.
10. The method according to any one of claims 1 to 7, characterized in that, Respectively obtaining the jitter values detected when receiving each audio data packet includes: Performing the following operations respectively for each of the audio data packets: Determining the time difference between an audio data packet and the previous audio data packet based on the reception time and transmission time of the audio data packet, and the reception time and transmission time of the previous audio data packet of the audio data packet; Determining the jitter value detected when receiving the audio data packet based on the time difference, a smoothing coefficient, and the jitter value corresponding to the previous audio data packet.
11. An apparatus for adjusting the buffer length, characterized in that, Includes: A receiving module, configured to receive each audio data packet sent by a sending end, where each audio data packet is obtained by encoding at least one audio frame of an audio signal; A jitter detection module, configured to respectively obtain the jitter values detected when receiving each audio data packet; A processing module, configured to obtain the speech detection result corresponding to each audio frame and the speech rate detection result of the audio signal; where the speech detection result indicates whether the corresponding audio frame contains speech data; the speech rate detection result is a result obtained by detecting the speech content contained in the audio signal within a unit time; An adjustment module, configured to adjust the current buffer length of the buffer accordingly based on the speech rate detection result, each speech detection result, and each jitter value, in combination with a preset buffer length adjustment strategy.
12. The device according to claim 11, wherein The adjustment module is further configured to: Determine, from the audio data packets, the audio data packets that meet the reception time condition based on the reception time corresponding to each of the audio data packets; If it is determined that each speech detection result does not include speech data, adjust the current buffer length of the buffer accordingly based on a non-speech control strategy and the target jitter value of the determined audio data packets; If it is determined that at least one speech detection result includes speech data, adjust the buffer length accordingly based on the speech rate detection result and the target jitter value, in combination with a speech control strategy.
13. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of any one of the methods recited in claims 1 to 10.
14. A computer-readable storage medium, characterized in that, It includes program code, and when the program code runs on an electronic device, the program code is configured to cause the electronic device to execute the steps of any one of the methods recited in claims 1 to 10.
15. A computer program product, characterized in that, It includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; when a processor of an electronic device reads the computer instructions from the computer-readable storage medium, the processor executes the computer instructions, causing the electronic device to execute the steps of any one of the methods recited in claims 1 to 10.
Citation Information
Patent Citations
Method and device for adjusting jitter buffer
CN103685070A
Method of managing a jitter buffer, and jitter buffer using same
CN103988255A