Audio data processing method and device, equipment and storage medium

CN117061465BActive Publication Date: 2026-08-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-06
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

由于音频数据包在网络传输的过程中,可能会出现网络抖动的现象,该网络抖动的现象会影响音频数据包到达的时间,从而会导致音频数据播放卡顿与延时现象

Benefits of technology

[0019]本公开实施例提供了一种计算机程序产品或计算机程序,该计算机程序产品或计算机程序包括计算机指令,该计算机指令存储在计算机可读存储介质中。计算机设备的处理器从计算机可读存储介质读取该计算机指令,处理器执行该计算机指令,使得该计算机设备执行本公开任一实施例中的各种可选方式中提供的音频数据处理方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117061465B_ABST
    Figure CN117061465B_ABST
Patent Text Reader

Abstract

This disclosure provides an audio data processing method, apparatus, device, and storage medium, relating to the field of computer technology. The method includes: receiving a current audio frame of audio data, the content of which is user voice or music; obtaining a current network latency index based on the transmission and reception times of the current audio frame; determining the current actual storage capacity based on the current network latency index; obtaining a current instantaneous network jitter index based on the transmission and reception times; obtaining a current network jitter index based on the instantaneous network jitter index; determining a current target storage capacity based on the current network jitter index; buffering the current audio frame; and processing the audio frame in the buffer based on the current actual storage capacity and the current target storage capacity. By employing this disclosure, the current actual storage capacity and the current target storage capacity of the buffer can be adaptively adjusted based on the current network latency index and the current network jitter index, thereby improving the processing quality of audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an audio data processing method, apparatus, device, and storage medium. Background Technology

[0002] With the development of computer technology, terminal devices can provide increasingly higher quality audio services. For example, in instant messaging applications, the sending terminal can send corresponding audio data packets to the receiving terminal. The receiving terminal can then play the frames of the audio data packets in sequence. However, network jitter can occur during audio data packet transmission, affecting the arrival time and causing playback stuttering and delays. Therefore, audio data processing is necessary to improve its quality. Summary of the Invention

[0003] This disclosure provides an audio data processing method, apparatus, device, and storage medium, which can be used to improve the processing quality of audio data.

[0004] This disclosure provides an audio data processing method, comprising: receiving a current audio frame of audio data, the current audio frame including a transmission time of the current audio frame; obtaining a current network latency index based on the transmission time and reception time of the current audio frame; determining the current actual storage capacity of a buffer based on the current network latency index; obtaining a current instantaneous network jitter index based on the transmission time and reception time of the current audio frame, and the transmission time and reception time of the previous audio frame of the audio data; obtaining a current network jitter index based on the current instantaneous network jitter index; determining the current target storage capacity of the buffer based on the current network jitter index; buffering the current audio frame in the buffer; and processing audio frames in the buffer, the audio frames including the current audio frame, based on the current actual storage capacity and the current target storage capacity of the buffer.

[0005] This disclosure provides an audio data processing apparatus, comprising: a receiving module for receiving a current audio frame of audio data, the current audio frame including a transmission time; an acquisition module for obtaining a current network latency index based on the transmission time and reception time of the current audio frame; a determining module for determining the current actual storage capacity of a buffer based on the current network latency index; the acquisition module further for obtaining a current instantaneous network jitter index based on the transmission time and reception time of the current audio frame and the transmission time and reception time of the previous audio frame of the audio data; the acquisition module further for obtaining a current network jitter index based on the current instantaneous network jitter index; the acquisition module further for determining the current target storage capacity of the buffer based on the current network jitter index; a caching module for caching the current audio frame in the buffer; and a processing module for processing audio frames in the buffer, the audio frames including the current audio frame, based on the current actual storage capacity and the current target storage capacity of the buffer.

[0006] In an exemplary embodiment, the determining module is configured to obtain the smooth reception time of the current audio frame based on the current network latency index and the transmission time of the current audio frame; and to obtain the current actual storage capacity of the buffer based on the smooth reception time of the current audio frame and the actual playback time of the current audio frame.

[0007] In an exemplary embodiment, the determining module is configured to obtain a current residual jitter index based on the transmission and reception times of the current audio frame, the transmission and reception times of the previous audio frame, and the smooth reception time of the current audio frame; and to obtain the current actual storage capacity of the buffer based on the actual playback time of the current audio frame, the smooth reception time of the current audio frame, and the current residual jitter index.

[0008] In an exemplary embodiment, the acquisition module is configured to process the transmission and reception times of the current audio frame using a target filter to obtain the current network latency index;

[0009] The device further includes: a reset module, configured to obtain a first difference index between the reception time and transmission time of the current audio frame; obtain a second difference index between the current network latency index and the first difference index; if the second difference index is determined to be greater than a first threshold, obtain a cumulative difference index based on the second difference index; if the cumulative difference index is determined to be greater than the second threshold, reset the target filter.

[0010] In an exemplary embodiment, the acquisition module is configured to determine the interval in which the current instantaneous network jitter index is located; determine the count of the instantaneous network jitter index in each interval, wherein the instantaneous network jitter index includes the current instantaneous network jitter index; normalize the count of the instantaneous network jitter index in each interval to obtain the current network jitter probability density; and use the upper quantile of the current network jitter probability density as the current network jitter index.

[0011] In an exemplary embodiment, the acquisition module is further configured to: if it is determined that the number of packet loss compensations for the audio data within a first predetermined sliding window duration is greater than a first packet loss threshold, then increase α by a first step length, where α is a real number greater than 0 and less than 1; if it is determined that the number of packet loss compensations for the audio data within a second predetermined sliding window duration is less than or equal to a second packet loss threshold, then decrease α by a second step length; wherein the first step length is greater than the second step length.

[0012] In an exemplary embodiment, the interval includes a target interval; the acquisition module is configured to, if it is determined that the network jitter index of the audio data within a third predetermined sliding window duration is less than the minimum value of the network jitter index of the audio data within a fourth predetermined sliding window duration, then set the count of the instantaneous network jitter index in the target interval to zero; renormalize the count of the instantaneous network jitter index in each interval to obtain an updated current network jitter probability density; and use the upper quantile of the updated current network jitter probability density as the current network jitter index.

[0013] In an exemplary embodiment, the acquisition module is configured to perform a first anomaly detection on the current audio frame, the first anomaly detection including at least one of abnormal packet detection, duplicate packet detection, and late packet detection; and to obtain the current network jitter instantaneous index by using the transmission time and reception time of the current audio frame that passed the first anomaly detection, as well as the transmission time and reception time of the previous audio frame of the audio data.

[0014] In an exemplary embodiment, the processing module is configured to determine an upper limit index and a lower limit index of the current target storage capacity based on the current target storage capacity; the processing module is further configured to: accelerate the playback of the current audio frames in the buffer if the current audio frames are continuous, the current actual storage capacity is greater than the upper limit index of the current target storage capacity, and the audio data has not experienced packet loss compensation within a fifth predetermined time period; slow down the playback of the current audio frames in the buffer if the current audio frames are continuous and the current actual storage capacity is less than the lower limit index of the current target storage capacity; play the current audio frames in the buffer if the current audio frames are discontinuous and the current time is greater than the expected playback time of the current audio frames; and perform packet loss compensation on the audio data based on the current audio frames if the current audio frames are discontinuous and the current time is less than the expected playback time of the current audio frames.

[0015] In an exemplary embodiment, the acquisition module is further configured to obtain a current receiving end latency index based on the actual playback time and reception time of the current audio frame; obtain the difference between the current target storage capacity and the current residual jitter index; and obtain the estimated playback time of the current audio frame based on the transmission time of the current audio frame, the current network latency index, the current receiving end latency index, and the difference between the current target storage capacity and the current residual jitter index.

[0016] In an exemplary embodiment, the difference between the upper limit of the current target storage capacity and the current target storage capacity is greater than the difference between the lower limit of the current target storage capacity and the current target storage capacity.

[0017] This disclosure provides a computer device including a processor, a memory, and an input / output interface. The processor is connected to both the memory and the input / output interface. The input / output interface is used to receive and output data. The memory is used to store a computer program. The processor is used to invoke the computer program so that the computer device containing the processor executes the audio data processing method of any embodiment of this disclosure.

[0018] This disclosure provides a computer-readable storage medium storing a computer program adapted to be loaded and executed by a processor, such that a computer device having the processor performs an audio data processing method according to any embodiment of this disclosure.

[0019] This disclosure provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the audio data processing method provided in various alternative embodiments of this disclosure.

[0020] The technical solution provided in this disclosure, by processing the current audio frame, can obtain the current network latency index and the current network jitter index. Then, based on the current network latency index and the current network jitter index, the current actual storage capacity and the current target storage capacity of the buffer are obtained respectively. Thus, when the current audio frame is cached in the buffer, the audio frame in the buffer can be processed based on the obtained current actual storage capacity and the current target storage capacity. Since the current network latency index and the current network jitter index used to transmit the current audio frame are considered in the process of obtaining the current actual storage capacity and the current target storage capacity, it is possible to adaptively adjust the current actual storage capacity of the buffer based on the current transmission network conditions, improve the processing quality of audio data, and thus achieve a better balance between audio stuttering and latency. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of an implementation environment provided by an embodiment of this disclosure.

[0022] Figure 2 This is a schematic diagram of another implementation environment provided by the embodiments of this disclosure.

[0023] Figure 3 This is a flowchart of an audio data processing method provided in an embodiment of this disclosure.

[0024] Figure 4 This is a schematic diagram of a current network latency indicator provided in an embodiment of this disclosure.

[0025] Figure 5 This is a schematic diagram of a calculation method corresponding to a Kalman filtering process provided in an embodiment of this disclosure.

[0026] Figure 6 This is a schematic diagram of a function obtained after compensating the mean of the probability density of smooth reception time, as provided in an embodiment of this disclosure.

[0027] Figure 7 This is a flowchart of a method for obtaining the current actual storage capacity of a cache area according to an embodiment of this disclosure.

[0028] Figure 8This is a schematic diagram of a current network jitter index provided in an embodiment of this disclosure.

[0029] Figure 9 This is a schematic diagram of the current target storage capacity and the current actual storage capacity provided in an embodiment of this disclosure.

[0030] Figure 10 This is a flowchart of a method for playing audio data provided in an embodiment of this disclosure.

[0031] Figure 11 This is a flowchart of a method for inserting received audio data into a buffer area, as provided in an embodiment of this disclosure.

[0032] Figure 12 This is a diagram of an audio data processing architecture provided in an embodiment of this disclosure.

[0033] Figure 13 This is a schematic diagram of an audio data processing device provided in an embodiment of the present disclosure.

[0034] Figure 14 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure. Detailed Implementation

[0035] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0036] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0037] This disclosure provides an audio data processing method, such as... Figure 1 The diagram illustrates an implementation environment provided by an embodiment of this disclosure. This implementation environment may include terminal 11 and terminal 12.

[0038] Terminal 11 can collect audio data via a microphone or other device, which can be user voice or music. Terminal 11 can then send this audio data to terminal 12 in real time as audio frames. Terminal 12 can receive the audio data and process it using the methods provided in this embodiment. Optionally, after processing the audio data, terminal 12 can play the corresponding audio. For example, terminal 11 and terminal 12 can be connected via a network, which can be a wireless network or a wired network.

[0039] Or, such as Figure 2 As shown, terminals 11 and 12 can transmit audio through clients of applications including instant messaging applications. In an exemplary embodiment, terminals 11 and 12 are respectively connected to server 13. After acquiring audio data, terminal 11 can send the audio data to server 13 in real time in the form of audio frames. Server 13 receives the audio data and forwards it to terminal 12. Terminal 12 can receive the audio data sent by server 13 and process the audio data using the method provided in this embodiment. Optionally, after processing the audio data, terminal 12 can play the audio corresponding to the audio data. Exemplarily, both terminal 11 and terminal 12 can be connected to server 13 through a network, which can be a wireless network or a wired network.

[0040] Optionally, the aforementioned wireless or wired networks use standard communication technologies and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to Local Area Networks (LANs), Metropolitan Area Networks (MANs), Wide Area Networks (WANs), mobile, wired or wireless networks, private networks, or any combination of virtual private networks. In some embodiments, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network. Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Networks (VPNs), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, custom and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.

[0041] Terminal 11 and Terminal 12 can be various electronic devices, including but not limited to smartphones, tablets, laptops, desktop computers, wearable devices, smart voice interaction devices, vehicle terminals, smart home appliances, aircraft, augmented reality devices, virtual reality devices, etc.

[0042] Optionally, the clients of the applications installed on terminals 11 and 12 can be the same, or clients of the same type of application based on different operating systems. Depending on the terminal platform, the specific form of the application client can also differ; for example, the application client can be a mobile client, a PC client, etc.

[0043] Server 13 can be a server that provides various services, such as a backend management server that supports the devices operated by users using terminals 11 and 12. The backend management server can analyze and process received requests and other data, and then feed the processing results back to the terminals.

[0044] Optionally, the server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0045] Those skilled in the art will know that Figure 1 Terminal 11, Terminal 12 and Figure 2 The number of terminals 11, 12, and servers 13 in the figures is illustrative. Any number of terminals 11, 12, and servers 13 can be used as needed. This disclosure does not limit the number of terminals 11, 12, and servers 13.

[0046] The following detailed description of this exemplary implementation method is provided in conjunction with the accompanying drawings and embodiments.

[0047] First, this disclosure provides an audio data processing method that can be executed by any computer device with computing capabilities. The following example illustrates the application of this method to a terminal.

[0048] Figure 3 A flowchart of an audio data processing method according to an embodiment of this disclosure is shown, such as... Figure 3 As shown, the audio data processing method provided in this embodiment may include the following steps 301 to 308.

[0049] In step 301, the current audio frame of the audio data is received, and the current audio frame includes the transmission time of the current audio frame.

[0050] This disclosure does not limit the application scenarios of the audio data processing method. For example, the method can be applied to instant messaging applications. In this scenario, the method can be executed by a terminal. When a first user and a second user conduct a voice call through an instant messaging application, both the terminal corresponding to the first user and the terminal corresponding to the second user can process the audio data sent by the other party using the audio data processing method provided in this disclosure. As another example, the method can also be applied to live streaming applications, remote conferencing applications, etc. This disclosure does not limit the content of the audio data; the content of the audio data can be determined based on the application scenario. This disclosure also does not limit the current audio frame of the audio data; the current audio frame can be any frame in the audio data.

[0051] Taking an instant messaging application as an example, when a first user communicates with a second user, the terminal corresponding to the first user can collect audio data from the first user in real time and encapsulate the audio frames in the audio data frame by frame to obtain multiple audio data packets. Furthermore, the terminal corresponding to the first user can sequentially send these multiple audio data packets to the terminal corresponding to the second user. Each audio data packet can carry a sending timestamp, which indicates the sending time of the corresponding audio frame. The terminal corresponding to the second user can receive these multiple audio data packets.

[0052] In step 302, the current network latency index is obtained based on the transmission and reception times of the current audio frame.

[0053] For example, the instantaneous value of network latency is the difference between the reception time and the transmission time of a certain audio frame, which can be expressed by the following formula (1):

[0054]

[0055] In formula (1), Let s be the instantaneous network latency value of the k-th frame. k Let r be the transmission time of the audio frame of the kth frame. k is the reception time of the audio frame of the kth frame.

[0056] Network latency can be defined as the expected difference between the reception time and transmission time of the previous k audio frames (k can be an integer greater than or equal to zero), including the current audio frame (represented as the k-th frame). Network latency can be expressed by the following formula (2):

[0057] d net =E[r k -s k (2)

[0058] In formula (2), d net For network latency, the letter E represents the operation of seeking the expected value.

[0059] The current network latency metric is a smoothed estimate of the difference between the reception and transmission times of historical audio frames, including the current audio frame (denoted as frame k). In an exemplary embodiment, network latency may be caused by poor network conditions or slow terminal response due to large amounts of audio data.

[0060] Figure 4 This is a schematic diagram of a current network latency indicator provided in an embodiment of this disclosure. Figure 4 The function of transmission time describes the relationship between the frame number of an audio frame and its transmission time. The function of reception time describes the relationship between the frame number of an audio frame and its reception time. Therefore, the current network latency metric can be expressed as the average distance between the functions of transmission time and reception time. For example, the frame number of the audio frame is an integer greater than or equal to zero; for instance, the frame number of the first audio frame could be 0, and the frame number of the twenty-fifth audio frame could be 24.

[0061] In some embodiments, a target filter can be used to estimate the current network latency metric. Exemplarily, this target filter can be based on the Kalman filter algorithm, but this disclosure is not limited thereto. In the Kalman filter algorithm, the relationship between the state sequence and the observed variable sequence can be described by a state equation and an observation equation. In this embodiment, the state equation can be:

[0062] ω k+1 =ω k +ξ k (3)

[0063] The observation equation can be:

[0064] s k =[r k ,1]·ω k +η k (4)

[0065] In the above state equation (3), ω k (ω k It can be a 2×1 column vector, which can be used to describe the linear mapping relationship between the transmission time and reception time of an audio frame, and can be called ω. k Let ω be the system state variable for the k-th frame. The index k (k is an integer greater than or equal to zero) represents the k-th frame. k [2](ω k The value in the first column and second row of the table is the current network latency metric that needs to be calculated. kFor example, the ξ is a state transition noise. k It can be a Gaussian white noise with zero mean.

[0066] In the above observation equation (4), s k r is the transmission time of the audio frame of the k-th frame. k Let be the reception time of the audio frame of the k-th frame. [r] k [1] is the observation matrix, which can be represented by C. k To represent it, namely C k =[r k ,1]. η k For the purpose of observing noise, exemplarily, the η k It can be a Gaussian white noise with zero expectation. In actual calculations, we can assume η k With ξ k It does not change over time.

[0067] After determining the observation equation and the state equation, the Kalman filtering process can be expressed by the following formulas (5) to (11):

[0068] P 0,0 =Var(ω0) (5)

[0069] Q k =Var(ξ k (6)

[0070] r k =Var(η k (7)

[0071] P k,k-1 =P k-1,k-1 +Q k (8)

[0072]

[0073] P k,k =(IG) k C k )P k,k-1 (10)

[0074]

[0075]

[0076] Among them, P 0,0 Q k ω0 and ξ are respectively k The corresponding covariance. P 0,0 It can be used to indicate the initial state of the system. The corresponding covariance, where, For example, the prior state estimation covariance P corresponding to the k-th frame can be calculated using formula (8). k,k-1 That is, the P k,k-1 for The corresponding covariance, Let P be the prior state estimate corresponding to the k-th frame. For example, in this formula, P... k-1,k-1 Estimate the covariance of the posterior state corresponding to the (k-1)th frame. After obtaining P... k,k-1 Then, G can be calculated using formula (9). k Among them, G k For Kalman gain, G k The value can change at different times.

[0077] For example, in determining G k Then, the posterior state estimation covariance P corresponding to the k-th frame can be calculated using formula (10). k,k Furthermore, because this Kalman filtering process needs to run continuously until no more new audio frames are read, therefore the P... k,k This process needs to be repeated continuously. When the system enters state k+1, then P in formula (10) k,k That is, it becomes P k+1,k+1 P k,k-1 That is, it becomes P k+1,k Therefore, the covariance of the posterior state estimate corresponding to the current time can be continuously calculated using this formula (10).

[0078] For example, formula (11) is used to indicate the prior state estimate corresponding to the k-th frame. It equals the posterior state estimate corresponding to the (k-1)th frame. Formula (12) is used to indicate The calculation method, after obtaining After that, the current network latency metric can be obtained.

[0079] It should be noted that this disclosure does not require synchronization of the clocks of the terminal sending audio data and the terminal receiving audio data, because when using the current network latency index to obtain the current actual storage capacity, only the change in the current network latency index will affect the calculation of the current actual storage capacity, while the absolute value of the current network latency index will not affect the subsequent calculation.

[0080] In one possible implementation, the Kalman filtering process described above can be based on... Figure 5 The calculation method shown is implemented as follows. Figure 5 As shown, in,

[0081] In some embodiments, in addition to the Kalman filter algorithm, the target filter described above can also estimate the current network latency index using other filtering algorithms. This disclosure does not limit this, and the filtering algorithm can be, for example, a low-pass filter that uses the instantaneous value of network latency as input.

[0082] The aforementioned estimation of the current network latency using a target filter allows for tracking changes in the network state, resulting in a more accurate current network latency metric. This more accurate metric leads to a more precise determination of the actual current storage capacity. Consequently, better processing results can be achieved when processing audio data.

[0083] In an exemplary embodiment, the transmission and reception times of the current audio frame can be processed using a target filter to obtain a current network latency index. The method may further include: obtaining a first difference index between the reception and transmission times of the current audio frame; obtaining a second difference index between the current network latency index and the first difference index; if the second difference index is determined to be greater than a first threshold, obtaining a cumulative difference index based on the second difference index; and if the cumulative difference index is determined to be greater than the second threshold, resetting the target filter.

[0084] For example, the algorithm for resetting the filter described above can be called a cumulative sum control graph algorithm, which can be used to detect whether the target filter is abnormal. This disclosure does not limit the magnitude of the first threshold and the second threshold described above; the magnitude of both the first threshold and the second threshold can be set based on experience or application scenarios. For example, when the current network latency index... The first difference index (r) k -s k The second difference index between ) If the difference exceeds the set first threshold, the second difference indicator can be accumulated. The cumulative difference index is obtained; when the second difference index is less than or equal to the set first threshold, the current cumulative difference index can be cleared to zero. After the target filter obtains the current network latency index of the next frame, it recalculates the second difference index. If the second difference index is greater than the first threshold, it is accumulated again to obtain a new cumulative difference index. This continues until the cumulative difference index exceeds the set second threshold. At this point, the target filter is in an abnormal state, such as divergence, and therefore the target filter can be reset. For example, this can be achieved by setting the current ω... k The target filter is reset by resetting its value to the initial value ω0.

[0085] By resetting the target filter as described above, the abnormality of the target filter's state can be detected, and the filter can be reset promptly when it is in an abnormal state. This improves the robustness of the method for obtaining the current network latency metric.

[0086] In step 303, the current actual storage capacity of the cache is determined based on the current network latency metric.

[0087] For example, the buffer is an area for caching received voice data. This buffer can be used to delay the retrieval of voice data, allowing it to be retrieved and played from the buffer at a stable and smooth rate, thus enabling the user to hear clear and continuous voice. The current actual storage capacity is the amount of voice data actually cached in this buffer, also known as the actual storage level.

[0088] In an exemplary embodiment, determining the current actual storage capacity of the buffer based on the current network latency index may include: obtaining the smooth reception time of the current audio frame based on the current network latency index and the transmission time of the current audio frame; and obtaining the current actual storage capacity of the buffer based on the smooth reception time of the current audio frame and the actual playback time of the current audio frame.

[0089] In some embodiments, the smooth reception time of the current audio frame can be calculated using the following formula (13):

[0090]

[0091] Formula (13) above uses the k-th audio frame as the current audio frame. s is the smooth reception time of the current audio frame. k The transmission time of the current audio frame. This represents the current network latency metric. The smoothed reception time of the current audio frame can be used to estimate the reception time of the next audio frame. For example... Figure 4 As shown, the receiving time is smoothed, which means converting the jagged receiving time function into a smooth function.

[0092] In some embodiments, after obtaining the smooth reception time of the current audio frame, the expected value of the difference between the actual playback time and the smooth reception time of the current audio frame can be calculated. The expected value of the difference is the current actual storage capacity of the buffer. For example, let the audio frame of the k-th frame be the current audio frame, the current actual storage capacity of the buffer can be calculated using the following formula (14):

[0093]

[0094] Where, τ op represents the current actual storage capacity. k The actual playback time of the current audio frame. The smooth reception time of the current audio frame. This represents the current residual jitter metric. The value of E in the formula indicates the operation of calculating the expectation. In practice, the expectation can be estimated using a sliding window method to calculate the sample mean, or by using an IIR (Infinite Impulse Response) low-pass filter.

[0095] For example, in the application scenario of this disclosure embodiment, ideally, voice data is received once at uniform intervals. However, in reality, due to differences in transmission delays of different voice data, the time for receiving voice data often varies. This variation causes voice data to arrive out of order, a phenomenon known as network jitter. Network jitter may be caused by network congestion, timing drift, or routing changes. The aforementioned method for calculating the current actual storage capacity of the buffer is applicable to network jitter scenarios with symmetrical probability density, such as those exhibiting a Gaussian distribution. However, it produces significant errors for network jitter with asymmetrical probability density. For network jitter with asymmetrical probability density, such as a geometric distribution, the smooth reception time is often too short. Here, the concept of a current residual jitter metric can be introduced. This current residual jitter metric can be used to compensate for the smooth reception time, reducing the error in calculating the current actual storage capacity in asymmetric network jitter scenarios.

[0096] It is used to compensate for the mean of the probability density of smooth reception time, thereby eliminating the interference of late frames on smooth reception time and reducing the error in calculating the current actual storage capacity in non-Gaussian network jitter scenarios.

[0097] like Figure 6 As shown, the function obtained after compensating for the smooth reception time is as follows: Figure 6 The smooth reception time is shown as a function of residual jitter. Additionally, Figure 6 Functions of transmission time, reception time, and smooth reception time Figure 4 The same applies, so I will not repeat it here.

[0098] like Figure 7 As shown, in one possible implementation, the current actual storage capacity of the buffer is obtained based on the smooth reception time of the current audio frame and the actual playback time of the current audio frame, including the following steps 3031 to 3032.

[0099] In step 3031, the current residual jitter index is obtained based on the transmission and reception times of the current audio frame, the transmission and reception times of the previous audio frame, and the smooth reception time of the current audio frame.

[0100] In an exemplary embodiment, let the audio frame of the k-th frame be the current audio frame, and let the current residual jitter index be... The calculation can be performed using the following formula (15):

[0101]

[0102] In this formula, r k s represents the reception time of the current audio frame. k-1 r is the time when the previous audio frame was sent. k-1 The reception time of the previous audio frame, s k The transmission time of the current audio frame. The smooth reception time for the current audio frame.

[0103] In step 3032, the current actual storage capacity of the buffer is obtained based on the actual playback time of the current audio frame, the smooth reception time of the current audio frame, and the current residual jitter index.

[0104] For example, when the audio frame of the kth frame is the current audio frame, the current actual storage capacity of the buffer can be calculated by formula (14) in step 303 above, which will not be repeated here.

[0105] The method for determining the current actual storage capacity of a buffer provided in this disclosure reduces the error in calculating the current actual storage capacity in scenarios with asymmetric network jitter by first obtaining the current residual jitter index and then using the current residual jitter index to obtain the current actual storage capacity of the buffer. Furthermore, this method can track changes in the current actual storage capacity.

[0106] In step 304, the instantaneous network jitter index is obtained based on the transmission and reception times of the current audio frame and the transmission and reception times of the previous audio frame.

[0107] For example, the current instantaneous network jitter metric can be calculated by subtracting the difference between the reception and transmission times of the previous audio frame from the difference between the reception and transmission times of the current audio frame. Therefore, the current instantaneous network jitter metric can be calculated using formula (16):

[0108]

[0109] in, r is the instantaneous indicator of current network jitter. k s represents the reception time of the current audio frame. k r is the transmission time of the current audio frame. k-1 The reception time of the previous audio frame, s k-1This is the sending time of the previous audio frame.

[0110] In some embodiments, obtaining the instantaneous network jitter index based on the transmission and reception times of the current audio frame and the transmission and reception times of the previous audio frame of the audio data may include: performing a first anomaly detection on the current audio frame, the first anomaly detection including at least one of abnormal packet detection, duplicate packet detection, and late packet detection; and using the transmission and reception times of the current audio frame that has passed the first anomaly detection, as well as the transmission and reception times of the previous audio frame of the audio data, to obtain the instantaneous network jitter index.

[0111] For example, the above-described abnormal packet detection is used to remove abnormal packets. Specifically, the instantaneous network latency of the received k-th audio data packet is calculated and compared with the instantaneous network latency values ​​of the most recently received N (N can be a positive integer) audio data packets (k-1, k-2, ..., k-N+1). If the difference between the instantaneous network latency of the k-th audio data packet and the median of the most recently received N audio data packets is greater than the abnormal packet threshold, the k-th audio data packet is determined to be an abnormal packet. This abnormal packet threshold can be set based on experience or application scenarios, and this embodiment does not limit it. Taking an abnormal packet threshold of 150 milliseconds as an example, assuming one frame corresponds to 20 milliseconds, if the instantaneous network latency values ​​of the three most recently received audio data packets are 50 milliseconds, 70 milliseconds, and 90 milliseconds respectively, and an audio frame with an instantaneous network latency value of 300 milliseconds is received, then the audio data packet corresponding to this audio frame with an instantaneous network latency value of 300 milliseconds is an abnormal packet.

[0112] For example, the duplicate packet detection described above is used to remove duplicate packets. If an audio data packet corresponding to the same frame number is retransmitted, then the retransmitted audio data packet is a duplicate packet. For instance, if an audio data packet with frame number #5 is received, and then another audio data packet with frame number #5 is received again, then the second audio data packet with frame number #5 is a duplicate packet.

[0113] For example, the aforementioned late packet detection is used to remove late packets whose lateness time exceeds a lateness threshold. When an audio data packet corresponding to a certain frame number has been played, and then an audio data packet with a frame number preceding that frame number is received, the audio data packet with the frame number preceding that frame number is a late packet. For example, after playing the audio data packet with frame number #20, an audio data packet with frame number #13 is received. Since the audio data packet with frame number #13 should have been played before the audio data packet with frame number #20, the audio data packet with frame number #13 is a late packet. This embodiment of the disclosure does not limit the size of the lateness threshold; for example, the lateness threshold can be 1 second or 1.5 seconds. After removing late packets whose lateness time exceeds the lateness threshold, the late packet will not be played, but it can still be used to update the current network jitter index and the current network latency index.

[0114] This embodiment of the disclosure can remove voice data packets from the three scenarios described above through a first anomaly detection, thus preventing these voice data packets from being used in the calculation of the current network jitter instantaneous index. Therefore, this embodiment of the disclosure can improve the robustness of calculating the current network jitter index.

[0115] In step 305, the current network jitter index is obtained based on the current instantaneous network jitter index.

[0116] For example, the current network jitter metric is the result of estimating the network jitter metric of multiple audio frames, including the current audio frame, using the current instantaneous network jitter metric. Figure 8 As shown, the current network jitter metric can be represented by the horizontal distance between two dotted lines. Figure 8 The functions for sending time and receiving time are... Figure 4 The same applies, so I will not repeat it here.

[0117] In some embodiments, obtaining the current network jitter index based on the current instantaneous network jitter index may include: determining the interval in which the current instantaneous network jitter index is located; determining the count of the instantaneous network jitter index in each interval, wherein the instantaneous network jitter index includes the current instantaneous network jitter index; normalizing the count of the instantaneous network jitter index in each interval to obtain the current network jitter probability density; and using the upper quantile of the current network jitter probability density as the current network jitter index.

[0118] In one possible implementation, the current network jitter index can be represented by a random variable D. In this case, the current network jitter index can be calculated using the following formula (17).

[0119]

[0120] in, D is the current network jitter metric. α Let α be the upper quantile of the random variable D: α = P{D≤D α}

[0121] This disclosure does not limit the value of α in the above formula; the value of α can be set based on experience or application scenario.

[0122] In some embodiments, α can also be a variable parameter. In an exemplary embodiment, the audio data processing method provided in this disclosure may further include: if it is determined that the number of packet loss compensations for audio data within a first predetermined sliding window duration is greater than a first packet loss threshold, then α is increased by a first step length; if it is determined that the number of packet loss compensations for audio data within a second predetermined sliding window duration is less than or equal to a second packet loss threshold, then α is decreased by a second step length; wherein, the first step length is greater than the second step length.

[0123] For example, packet loss rate refers to the ratio of the number of audio data packets lost during the encapsulation and transmission of audio data packets to the total number of audio data packets sent. When the packet loss rate is high, it is possible that all audio data packets stored in the buffer have been read, but no new audio data packets have been received. In this case, audio can be recovered using packet loss compensation methods.

[0124] This embodiment of the disclosure does not limit the length of the first predetermined sliding window duration and the second predetermined sliding window duration. For example, the length of both the first and second predetermined sliding window durations can be set to 1 second or 5 seconds. This embodiment of the disclosure does not limit the magnitude of the first packet loss threshold and the second packet loss threshold. For example, when the length of the first predetermined sliding window is 1 second, the first packet loss threshold can be set to 5 or 10; when the length of the second predetermined sliding window is 5 seconds, the magnitude of the second packet loss threshold can be set to 1 or 2. This embodiment of the disclosure also does not limit the first step length and the second step length. For example, the values ​​of the first step length and the second step length are decimals greater than zero and less than 1. For example, the first step length can be 0.01 and the second step length can be 0.001. Or the first step length can be 0.005 and the second step length can be 0.002.

[0125] In one possible implementation, the length of the first predetermined sliding window can be 1 second, and the length of the second predetermined sliding window can be 5 seconds. The first packet loss threshold is set to 10, and the second packet loss threshold is set to 1. The first step size is set to 0.01, and the second step size is set to 0.001. In this case, if the number of packet loss compensations for audio data within 1 second is greater than 10, then the value of α is increased by 0.01; if the number of packet loss compensations for audio data within 5 seconds is less than or equal to 1, then the value of α is decreased by 0.001.

[0126] For example, setting the first step length to be greater than the second step length allows the network jitter metric to quickly track jitter changes during sudden jitter events, while maintaining stability during periodic, pulse-like jitter events. For instance, in a periodic, pulse-like weak network scenario, a significant network jitter occurs once per cycle, while jitter at other times tends to be milder. When a significant network jitter occurs, more packet loss compensation may occur. Because the first step length is set larger, the value of α increases more quickly, and the current network jitter metric also increases more rapidly. Therefore, setting the first step length larger allows the acquired current network jitter metric to track changes in network jitter. Subsequently, when the network jitter subsides and the network condition remains a persistently weak network environment, there may be less packet loss compensation. Because the second step length is set smaller, the value of α decreases more slowly, thus keeping the current network jitter metric relatively stable.

[0127] This embodiment of the disclosure, by setting the first step length to be greater than the second step length, enables the acquired current network jitter metric to track changes in network jitter when significant network jitter occurs. Furthermore, it keeps the current network jitter metric relatively stable when network jitter tends to level off. This results in obtaining a superior current network jitter metric.

[0128] In an exemplary embodiment, the interval includes a target interval; wherein, using the upper quantile of the current network jitter probability density as the current network jitter index may include: if it is determined that the network jitter index of the audio data within a third predetermined sliding window duration is less than the minimum value of the network jitter index of the audio data within a fourth predetermined sliding window duration, then the count of the instantaneous network jitter index in the target interval is set to zero; the count of the instantaneous network jitter index in each interval is renormalized to obtain an updated current network jitter probability density; and the upper quantile of the updated current network jitter probability density is used as the current network jitter index.

[0129] For example, the current network jitter probability density can be divided into multiple intervals for counting the instantaneous network jitter index falling within each interval. Setting the target interval to zero allows the instantaneous network jitter index in each interval to restart counting, thereby updating the obtained network jitter probability density. For example, the multiple intervals mentioned above may include 0 milliseconds to 20 milliseconds, 20 milliseconds to 40 milliseconds, 40 milliseconds to 60 milliseconds, ..., 1980 milliseconds to 2000 milliseconds. In this embodiment, the target interval can be an interval with a lower limit of the maximum value of the instantaneous network jitter index corresponding to the third predetermined sliding window duration plus 100 milliseconds, and an upper limit of 2000 milliseconds, but this disclosure is not limited to this. This disclosure does not limit the length of the third predetermined sliding window duration and the fourth predetermined sliding window duration. For example, the length of the third predetermined sliding window duration can be 10 seconds or 15 seconds, and the length of the fourth predetermined sliding window duration can be 20 seconds or 30 seconds. Taking a third predetermined sliding window duration of 10 seconds and a fourth predetermined sliding window duration of 20 seconds as an example. If the instantaneous network jitter index of the current audio frame up to 10 seconds ago, or the instantaneous network jitter index of the audio frame up to 30 seconds ago, exceeds 300 milliseconds, then the count of the instantaneous network jitter index in the target interval can be set to zero, and the count of the instantaneous network jitter index in each interval can be normalized again to obtain the updated current network jitter probability density.

[0130] In situations such as when the network state changes from a weak state to a normal state, network jitter may decrease rapidly. The present invention, through the method described above for updating the current network jitter probability density, can reduce the current network jitter index accordingly. Therefore, the present invention can obtain a more accurate current network jitter index.

[0131] In step 306, the current target storage capacity of the buffer is determined based on the current network jitter index.

[0132] For example, the current target storage capacity is the number of audio frames that need to be cached in the buffer based on the current network conditions. Subsequent processing of the audio frames in the buffer involves accelerating or decelerating the playback of the audio frames to bring the current actual storage capacity closer to the current target storage capacity.

[0133] In some embodiments, the current target storage capacity can be calculated using the following formula (18):

[0134]

[0135] Where, τ t For the current target storage capacity, This represents the current network jitter metric. x is a preset value used to indicate a threshold for network jitter. This embodiment of the disclosure does not limit the preset value x; for example, x can be 50 milliseconds or 60 milliseconds.

[0136] In an exemplary embodiment, when the current network jitter index is greater than the preset value x, it indicates that the network condition is weak, and the current network jitter index is used as the current target storage capacity of the buffer, thereby increasing the current target storage capacity; when the current network jitter index is less than the preset value x, it indicates that the network condition is normal, and the preset value x is used as the current target storage capacity of the buffer.

[0137] In step 307, the current audio frame is cached in the buffer.

[0138] In some embodiments, the current audio frame may carry a frame number, which can be used to cache the current audio frame at the corresponding position in the buffer. Therefore, it can be ensured that audio frames in the buffer are read in order of their frame numbers.

[0139] In step 308, audio frames in the buffer are processed according to the current actual storage capacity and the current target storage capacity of the buffer. The audio frames include the current audio frame.

[0140] In an exemplary embodiment, processing audio frames in the buffer based on the current actual storage capacity and the current target storage capacity may include: determining an upper limit index and a lower limit index of the current target storage capacity based on the current target storage capacity.

[0141] In one possible implementation, the current target storage capacity limit can be calculated using the following formula (19):

[0142] τ t,h =x h *τ t (19)

[0143] Where, τ t,h x represents the current target storage capacity limit. h The coefficient of the current target storage capacity limit indicator is not limited in this embodiment of the disclosure. It can be set based on experience and application scenarios. For example, the coefficient of the current target storage capacity limit indicator can be 1.25 or 1.5.

[0144] Alternatively, the current target storage capacity lower limit can be calculated using the following formula (20):

[0145] τ t,l =xl *τ t (20)

[0146] Where, τ t,l x is the lower limit of the current target storage capacity. l The coefficient of the current target storage capacity lower limit indicator is not limited in this embodiment of the disclosure. It can be set based on experience and application scenarios. For example, the coefficient of the current target storage capacity lower limit indicator can be 0.875 or 0.5.

[0147] For example, when the current network conditions are good, the difference between the upper and lower limits of the current target storage capacity can be set smaller, improving the precision of cache capacity control. When the current network conditions are poor, the difference between the upper and lower limits of the current target storage capacity can be set larger, reducing the precision of cache capacity control and thus avoiding frequent adjustments to the actual storage capacity.

[0148] In some embodiments, processing audio frames in the buffer based on the current actual storage capacity and the current target storage capacity of the buffer further includes at least one of the following four methods:

[0149] Method 1: If the current audio frames are consecutive, the current actual storage capacity is greater than the current target storage capacity limit, and no packet loss compensation occurs in the audio data within the fifth predetermined duration, then the playback of the current audio frame in the buffer is accelerated. For example, if the frame number of the currently playing audio frame is consecutive to the frame number of the audio frame currently being read from the buffer, then the current audio frames are consecutive. For instance, if the frame number of the currently playing audio frame is #16, and the audio frame currently being read from the buffer is #17, then the current audio frames are consecutive. This embodiment does not limit the length of the fifth predetermined duration; for example, the fifth predetermined duration can be 1 second or 1.25 seconds.

[0150] For example, accelerating the playback of the current audio frame in the buffer aims to reduce the current actual storage capacity. Using the formula (14) for calculating the current actual storage capacity, it can be seen that accelerating the playback of the current audio frame in the buffer reduces the actual playback time p of the audio frame. k Current actual storage capacity τ o This will also decrease. Therefore, accelerating the playback of the current audio frame in the buffer can bring the current actual storage capacity closer to the current target storage capacity limit.

[0151] In another possible implementation, if the current audio frames are consecutive and the current actual storage capacity is greater than the current target storage capacity limit, the playback of the current audio frames in the buffer can be accelerated. The aforementioned condition that "no packet loss compensation occurs within the fifth predetermined duration" is a preferred condition; adding this condition reduces the likelihood of making a decision to accelerate the playback of the current audio frames. This is because more frequent acceleration may cause more packet loss compensation.

[0152] Optionally, to avoid frequent acceleration, the difference between the upper limit of the current target storage capacity and the current target storage capacity can be set to be greater than the difference between the lower limit of the current target storage capacity. This reduces the likelihood that the current actual storage capacity exceeds the upper limit of the current target storage capacity, thus minimizing the need for frequent adjustments to the current actual storage capacity.

[0153] Method 2: If the current audio frames are consecutive and the current actual storage capacity is less than the lower limit of the current target storage capacity, then the playback of the current audio frame in the buffer will be slowed down. If the frame number of the currently playing audio frame is not consecutive with the frame number of the audio frame currently being read from the buffer, then the current audio frames are considered discontinuous. For example, if the frame number of the currently playing audio frame is #20, and the smallest frame number of the audio frame currently being read from the buffer is #22, then the current audio frames are considered discontinuous.

[0154] For example, slowing down the playback of the current audio frame in the buffer aims to increase the current actual storage capacity. From the above calculation formula (14) for the current actual storage capacity, it can be seen that slowing down the playback of the current audio frame in the buffer increases the actual playback time p of the audio frame. k Current actual storage capacity τ o This will also increase. Therefore, slowing down the playback of the current audio frame in the buffer can bring the current actual storage capacity closer to the lower limit of the current target storage capacity.

[0155] Method 3: If the current audio frame is not continuous and the current time is greater than the expected playback time of the current audio frame, then play the current audio frame in the buffer.

[0156] For example, the audio data processing method provided in this disclosure may further include: obtaining a current receiving end latency index based on the actual playback time and reception time of the current audio frame; obtaining the difference between the current target storage capacity and the current residual jitter index; and obtaining the estimated playback time of the current audio frame based on the transmission time of the current audio frame, the current network latency index, the current receiving end latency index, and the difference between the current target storage capacity and the current residual jitter index.

[0157] For example, the current receiver latency metric is used to indicate the average delay time between the actual playback time and the reception time of an audio frame. In an exemplary embodiment, the current receiver latency metric can be calculated using the following formula (21):

[0158]

[0159] in, p represents the current receiver latency metric. k r is the actual playback time of the current audio frame. k This represents the reception time of the current audio frame. After obtaining the current receiver latency, the estimated playback time can be calculated, for example, using the following formula (21):

[0160]

[0161] in, s is the estimated playback time of the current audio frame. k The transmission time of the current audio frame. τ is the current network latency metric. t For the current target storage capacity, The current residual jitter index, This represents the current receiver latency indicator.

[0162] In one possible implementation, if the frame number of the currently playing audio frame is #15, and the smallest audio frame number that can be read from the buffer is #17, then the current audio frames are not consecutive. In this case, audio frame #17 is the current audio frame, and the estimated playback time of this current audio frame can be calculated. Assume the calculated estimated playback time is 12:51:32.123. If the current time is 12:51:32.125, then the current time is greater than the estimated playback time of the current audio frame, so audio frame #17 can be played directly. If audio frame #16 is received subsequently, then audio frame #16 will not be played. For example, this operation in method three can be called a frame skipping operation.

[0163] Method 4: If the current audio frame is not continuous and the current time is less than the expected playback time of the current audio frame, then packet loss compensation is performed on the audio data based on the current audio frame.

[0164] For example, if the frame number of the currently playing audio frame is #15, and the smallest currently available audio frame in the buffer is #17, and the calculated estimated playback time is 12:51:32.123, then if the current time is 12:51:32.115, the current time is less than the estimated playback time of the current audio frame. Therefore, audio frame #15 can be used to compensate for packet loss in audio frame #16, and the estimated audio frame can be played.

[0165] In some embodiments, after the decision-making process described in methods one to four above, the smooth reception time of the current audio frame can be adjusted. The reception time r of the current audio frame k The value is compared to the smooth reception time of the current audio frame. The reception time r of the current audio frame k If the difference between the two values ​​exceeds a reference threshold, the current audio frame will not be used for calculating the receiver latency index and the current residual jitter index. This improves the accuracy of the receiver latency index and the current residual jitter index calculation. This embodiment does not limit the reference threshold; it can be limited based on experience or application scenarios.

[0166] For example, Figure 9 This diagram illustrates the relationship between the current target storage capacity and the current actual storage capacity. Figure 9 The horizontal lines represent the actual playback time function, and the dotted lines represent the smoothed reception time minus residual jitter function. The current target storage capacity is... Figure 9 The horizontal distance between the dashed line and the dotted line in the diagram. The actual current storage capacity is... Figure 9 The horizontal distance between the horizontal line and the dotted line in the diagram. Additionally, Figure 9 The functions for transmission time, reception time, and smoothed reception time are all related to... Figure 4 The same applies here, so it will not be repeated.

[0167] It should be noted that, as Figure 3 As shown, in this embodiment of the present disclosure, steps 302 and 303 may be executed first, followed by steps 304, 305, and 306. Alternatively, steps 304, 305, and 306 may be executed first, followed by steps 302 and 303. Furthermore, steps 302 and 303 may be executed simultaneously with steps 304, 305, and 306. This embodiment of the present disclosure does not limit the scope of the embodiments.

[0168] Furthermore, in some embodiments, step 307 may be executed first after step 301, followed by steps 302, 303, 304, 305, and 306. This disclosure does not limit the scope of the embodiments.

[0169] The method provided in this disclosure, by processing the current audio frame, can obtain the current network latency index and the current network jitter index. Then, based on the current network latency index and the current network jitter index, it obtains the current actual storage capacity and the current target storage capacity of the buffer. Thus, when the current audio frame is cached in the buffer, the audio frame in the buffer can be processed based on the obtained current actual storage capacity and the current target storage capacity. Since the current network latency index and the current network jitter index used to transmit the current audio frame are considered in the process of obtaining the current actual storage capacity and the current target storage capacity, it is possible to adaptively adjust the current actual storage capacity of the buffer based on the current transmission network conditions, improve the processing quality of audio data, and thus achieve a better balance between audio stuttering and latency.

[0170] like Figure 10 As shown in the embodiment of this disclosure, a method for playing audio data is provided, which includes the following steps.

[0171] Step 1001: Read the current audio frame from the buffer. This step can be found in step 307 above, and will not be repeated here.

[0172] Step 1002, Decision: Normal playback, speed up, slow down, frame skip, packet loss compensation. The speed up, slow down, frame skip, and packet loss compensation in this step correspond to the operations performed in methods one, two, three, and four of step 308 above, and will not be repeated here. The decision in this step also includes normal playback, which means that no operation is performed on the currently read audio frame, and subsequent frames can be played directly.

[0173] Step 1003, Decoding. Audio frames in the buffer can be stored in the form of audio data packets. After reading the current audio frame in the buffer, the audio data packet corresponding to the current audio frame can be decoded to obtain the audio frame.

[0174] Step 1004, Second Anomaly Detection. This second anomaly detection can be the smooth reception time of the current audio frame in step 308 above. The reception time r of the current audio frame kThe operation involves comparing the values, and / or, in step 305 above, if the network jitter index of the audio data within the third predetermined sliding window is less than the minimum value of the network jitter index of the audio data within the fourth predetermined sliding window, then the count of the instantaneous network jitter index in the target interval is set to zero; the count of the instantaneous network jitter index in each interval is renormalized to obtain the updated current network jitter probability density; and the operation of using the upper quantile of α of the updated current network jitter probability density as the current network jitter index is performed. This step can be referred to in steps 305 and 308 above, and will not be repeated here.

[0175] Step 1005: Calculate the receiver delay index. This step can be found in step 308 above, and will not be repeated here.

[0176] Step 1006: Calculate the current residual jitter index. This step can be found in step 303 above, and will not be repeated here.

[0177] Step 1007, DSP (Digital Signal Processing) Algorithm: Acceleration, Deceleration, Packet Loss Compensation. When the decision result of step 1012 above is acceleration, deceleration, or packet loss compensation, the corresponding acceleration, deceleration, or packet loss compensation operation can be executed through the DSP algorithm. The operation steps for acceleration, deceleration, or packet loss compensation can be found in step 308 above, and will not be repeated here.

[0178] Step 1008: Play audio frames.

[0179] like Figure 11 As shown in the figure, this disclosure provides a method for inserting received audio data into a buffer, the method comprising the following steps.

[0180] Step 1101: Receive audio data packets. For example, the audio data packets are sent by another terminal.

[0181] Step 1102: Parse the header of the audio data packet. After receiving the audio data packet, it is necessary to parse its header. For example, the header of the audio data packet may contain the sending timestamp corresponding to the audio data packet.

[0182] Step 1103, First Anomaly Detection. This step can be referred to step 304 above, and will not be repeated here. Optionally, after performing the first anomaly detection, the target filter reset operation in step 302 above can also be performed.

[0183] Step 1104: Calculate the current network jitter index. This step can be found in step 305 above, and will not be repeated here.

[0184] Step 1105: Calculate the current network latency metric. This step is the same as step 302 above and will not be repeated here. Optionally, the order of steps 1104 and 1105 can be interchanged; this embodiment does not limit the order of steps 1104 and 1105.

[0185] Step 1106: Store the audio frame in the buffer. This step can be found in step 307 above, and will not be repeated here.

[0186] For example, an audio data processing architecture provided in this disclosure embodiment can be as follows: Figure 12 As shown, the frequency data processing architecture may include: a receiving unit 1201, a decoding unit 1202, an audio anti-jitter unit 1203, a buffer unit 1204, a decision unit 1205, a digital signal processing unit 1206, a playback unit 1207, a current network jitter index calculation unit 1208, a current network latency calculation unit 1209, a current receiver latency index calculation unit 1210, and a current residual jitter index calculation unit 1211.

[0187] The receiving unit 1201 receives audio data packets. The decoding unit 1202 decodes the received audio data packets to obtain the audio frames within them. The audio jitter reduction unit 1203 connects to the buffer unit 1204, the decision unit 1205, the digital signal processing unit 1206, and the playback unit 1207 to reduce network jitter by processing the audio data. The buffer unit 1204 stores the received audio data packets. The decision unit 1205 connects to the current network jitter index calculation unit 1208, the current network latency calculation unit 1209, the current receiver latency index calculation unit 1210, and the current residual jitter index calculation unit 1211 to make decisions regarding the audio frames in the audio data packets. The digital signal processing unit 1206 executes the decisions made by the decision unit 1205, such as acceleration, deceleration, and packet loss compensation, using a digital signal processing algorithm. The playback unit 1207 plays the audio frames read from the buffer unit 1204.

[0188] The current network jitter calculation unit 1208 is used to calculate the current network jitter index. The current network latency calculation unit 1209 is used to calculate the current network latency index. The current receiver latency calculation unit 1210 is used to calculate the current receiver latency index. The current residual jitter calculation unit 1211 is used to calculate the current residual jitter index.

[0189] Figure 13 This is a schematic diagram of an audio data processing apparatus provided in an embodiment of this disclosure. Figure 13As shown, the audio data processing apparatus provided in this embodiment may include a receiving module 1301, an acquisition module 1302, a determining module 1303, a buffering module 1304, and a processing module 1305.

[0190] The receiving module 1301 can be used to receive the current audio frame of audio data, which includes the transmission time of the current audio frame.

[0191] The acquisition module 1302 can be used to obtain the current network latency index based on the sending and receiving time of the current audio frame.

[0192] The determination module 1303 can be used to determine the current actual storage capacity of the buffer based on the current network latency metric.

[0193] The acquisition module 1302 can also be used to obtain the instantaneous network jitter index based on the transmission and reception times of the current audio frame and the transmission and reception times of the previous audio frame of the audio data.

[0194] The acquisition module 1302 can also be used to obtain the current network jitter index based on the current instantaneous network jitter index.

[0195] The acquisition module 1302 can also be used to determine the current target storage capacity of the buffer based on the current network jitter index.

[0196] The caching module 1304 can be used to cache the current audio frame to the buffer area.

[0197] The processing module 1305 can be used to process audio frames in the buffer based on the current actual storage capacity and the current target storage capacity of the buffer. The audio frames include the current audio frame.

[0198] In an exemplary embodiment, the determining module 1303 can be used to obtain the smooth reception time of the current audio frame based on the current network latency index and the transmission time of the current audio frame; and to obtain the current actual storage capacity of the buffer based on the smooth reception time of the current audio frame and the actual playback time of the current audio frame.

[0199] In an exemplary embodiment, the determining module 1303 can be used to obtain the current residual jitter index based on the transmission and reception times of the current audio frame, the transmission and reception times of the previous audio frame, and the smooth reception time of the current audio frame; and to obtain the current actual storage capacity of the buffer based on the actual playback time of the current audio frame, the smooth reception time of the current audio frame, and the current residual jitter index.

[0200] In an exemplary embodiment, the acquisition module 1302 can be used to process the transmission and reception times of the current audio frame using the target filter to obtain the current network latency index;

[0201] The device may further include:

[0202] The reset module can be used to obtain a first difference index between the reception time and transmission time of the current audio frame; obtain a second difference index between the current network latency index and the first difference index; if it is determined that the second difference index is greater than the first threshold, then obtain a cumulative difference index based on the second difference index; if it is determined that the cumulative difference index is greater than the second threshold, then reset the target filter.

[0203] In an exemplary embodiment, the acquisition module 1302 can be used to determine the interval in which the current instantaneous network jitter index is located; determine the count of the instantaneous network jitter index in each interval, the instantaneous network jitter index including the current instantaneous network jitter index; normalize the count of the instantaneous network jitter index in each interval to obtain the current network jitter probability density; and use the upper quantile of the current network jitter probability density as the current network jitter index.

[0204] In an exemplary embodiment, the acquisition module 1302 can also be used to: if it is determined that the number of packet loss compensations for the audio data within the first predetermined sliding window duration is greater than the first packet loss threshold, then increase α by the first step length, where α is a real number greater than 0 and less than 1; if it is determined that the number of packet loss compensations for the audio data within the second predetermined sliding window duration is less than or equal to the second packet loss threshold, then decrease α by the second step length; wherein the first step length is greater than the second step length.

[0205] In an exemplary embodiment, the interval includes a target interval; the acquisition module 1302 can be used to: if it is determined that the network jitter index of the audio data within a third predetermined sliding window duration is less than the minimum value of the network jitter index of the audio data within a fourth predetermined sliding window duration, then set the count of the instantaneous network jitter index in the target interval to zero; renormalize the count of the instantaneous network jitter index in each interval to obtain the updated current network jitter probability density; and use the upper quantile of α of the updated current network jitter probability density as the current network jitter index.

[0206] In an exemplary embodiment, the acquisition module 1302 is used to perform a first anomaly detection on the current audio frame. The first anomaly detection includes at least one of abnormal packet detection, duplicate packet detection, and late packet detection. The current network jitter instantaneous index is obtained by using the transmission time and reception time of the current audio frame that has passed the first anomaly detection, as well as the transmission time and reception time of the previous audio frame of the audio data.

[0207] In an exemplary embodiment, the processing module 1305 is configured to determine the upper limit index of the current target storage capacity and the lower limit index of the current target storage capacity based on the current target storage capacity;

[0208] The processing module 1305 is further configured to: accelerate the playback of the current audio frame in the buffer if the current audio frames are continuous, the current actual storage capacity is greater than the upper limit of the current target storage capacity, and no packet loss compensation occurs in the audio data within a fifth predetermined time period; slow down the playback of the current audio frame in the buffer if the current audio frames are continuous and the current actual storage capacity is less than the lower limit of the current target storage capacity; play the current audio frame in the buffer if the current audio frames are discontinuous and the current time is greater than the expected playback time of the current audio frame; and perform packet loss compensation on the audio data based on the current audio frame if the current audio frames are discontinuous and the current time is less than the expected playback time of the current audio frame.

[0209] In an exemplary embodiment, the acquisition module 1302 can also be used to obtain the current receiving end delay index based on the actual playback time and reception time of the current audio frame; obtain the difference between the current target storage capacity and the current residual jitter index; and obtain the estimated playback time of the current audio frame based on the transmission time of the current audio frame, the current network delay index, the current receiving end delay index, and the difference between the current target storage capacity and the current residual jitter index.

[0210] In an exemplary embodiment, the difference between the current target storage capacity upper limit indicator and the current target storage capacity is greater than the difference between the current target storage capacity and the current target storage capacity lower limit indicator.

[0211] The apparatus provided in this embodiment can obtain the current network latency index and the current network jitter index by processing the current audio frame. Then, it can obtain the current actual storage capacity and the current target storage capacity of the buffer based on the current network latency index and the current network jitter index, respectively. In this way, when the current audio frame is cached in the buffer, the audio frame in the buffer can be processed based on the obtained current actual storage capacity and the current target storage capacity. Since the current network latency index and the current network jitter index used to transmit the current audio frame are considered in the process of obtaining the current actual storage capacity and the current target storage capacity, the current actual storage capacity of the buffer can be adaptively adjusted based on the current transmission network conditions, thereby improving the processing quality of audio data and achieving a better balance between audio stuttering and latency.

[0212] See Figure 14 , Figure 14 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure. Figure 14As shown, the computer device in this embodiment may include one or more processors 1401, a memory 1402, and an input / output interface 1403. The processor 1401, memory 1402, and input / output interface 1403 are connected via a bus 1404. The memory 1402 stores a computer program, which includes program instructions. The input / output interface 1403 receives and outputs data, such as for data interaction between the host machine and the computer device, or for data interaction between various virtual machines within the host machine. The processor 1401 executes the program instructions stored in the memory 1402.

[0213] The processor 1401 can perform the following operations: receive the current audio frame of audio data, the current audio frame including the transmission time of the current audio frame; obtain the current network latency index based on the transmission time and reception time of the current audio frame; determine the current actual storage capacity of the buffer based on the current network latency index; obtain the current instantaneous network jitter index based on the transmission time and reception time of the current audio frame, and the transmission time and reception time of the previous audio frame of audio data; obtain the current network jitter index based on the current instantaneous network jitter index; determine the current target storage capacity of the buffer based on the current network jitter index; buffer the current audio frame into the buffer; and process the audio frames in the buffer, including the current audio frame, based on the current actual storage capacity and the current target storage capacity of the buffer.

[0214] In some feasible implementations, the processor 1401 may be a central processing unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0215] The memory 1402 may include read-only memory and random access memory, and provides instructions and data to the processor 1401 and input / output interface 1403. A portion of the memory 1402 may also include non-volatile random access memory. For example, the memory 1402 may also store device type information.

[0216] In practice, the computer device can execute the implementation methods provided by each step in any of the above method embodiments through its built-in functional modules. For details, please refer to the implementation methods provided by each step in the figure shown in the above method embodiments, which will not be repeated here.

[0217] This disclosure provides a computer device including a processor, an input / output interface, and a memory. The processor retrieves a computer program from the memory and executes the steps of the method shown in any of the above embodiments.

[0218] This disclosure also provides a computer-readable storage medium storing a computer program adapted to be loaded by a processor and executed by the method for training an acoustic model provided in each step of any of the above embodiments. Specific implementations of each step in each of the above embodiments can be found therein and will not be repeated here. Furthermore, the beneficial effects of using the same method will not be repeated here either. For technical details not disclosed in the embodiments of the computer-readable storage medium involved in this disclosure, please refer to the description of the method embodiments of this disclosure. As an example, the computer program may be deployed to execute on a single computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed across multiple locations and interconnected via a communication network.

[0219] The computer-readable storage medium can be the audio data processing apparatus provided in any of the foregoing embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the computer device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0220] This disclosure also provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various alternative embodiments described above.

[0221] The terms "first," "second," etc., used in the specification, claims, and drawings of this disclosure are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.

[0222] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0223] The methods and related apparatus provided in this disclosure are described with reference to the method flowcharts and / or structural diagrams provided in this disclosure. Specifically, each block of the method flowchart and / or structural diagram, as well as combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable application display device to create a machine, such that the instructions, which execute via the processor of the computer or other programmable application display device, generate instructions for implementing the process... Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable application display device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable application display device to cause a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 A process or multiple processes and / or structures illustrate the steps of the functions specified in one or more boxes.

[0224] The above-disclosed embodiments are merely preferred embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Therefore, any equivalent variations made in accordance with the claims of this disclosure shall still fall within the scope of this disclosure.

Claims

1. An audio data processing method, characterized in that, include: The current audio frame from which audio data is received, wherein the current audio frame includes the transmission time of the current audio frame; Based on the transmission and reception times of the current audio frame, a filter is used to estimate the current network latency index, which is a smoothed estimate of the difference between the reception and transmission times of historical audio frames, including the current audio frame. Based on the current network latency metric and the transmission time of the current audio frame, the smooth reception time of the current audio frame is obtained; The current actual storage capacity of the buffer is determined based on the expected difference between the smooth reception time and the actual playback time of the current audio frame, wherein the current actual storage capacity is the amount of voice data actually buffered in the buffer. Based on the transmission and reception times of the current audio frame, and the transmission and reception times of the previous audio frame of the audio data, the instantaneous network jitter index is obtained. The current network jitter index is obtained based on the instantaneous current network jitter index. The current target storage capacity of the cache is determined based on the current network jitter index. The current audio frame is cached in the buffer area; Based on the current actual storage capacity and the current target storage capacity of the buffer, the audio frames in the buffer are processed, and the audio frames include the current audio frame.

2. The method as described in claim 1, characterized in that, The current actual storage capacity of the buffer is obtained based on the smooth reception time and the actual playback time of the current audio frame, including: The current residual jitter index is obtained based on the transmission and reception times of the current audio frame, the transmission and reception times of the previous audio frame, and the smooth reception time of the current audio frame. The current actual storage capacity of the buffer is obtained based on the actual playback time of the current audio frame, the smooth reception time of the current audio frame, and the current residual jitter index.

3. The method as described in claim 1, characterized in that, Based on the transmission and reception times of the current audio frame, the current network latency metric is obtained, including: The current network latency index is obtained by processing the transmission and reception times of the current audio frame using a target filter. The method further includes: Obtain a first difference index between the reception time and the transmission time of the current audio frame; Obtain a second difference index between the current network latency index and the first difference index; If it is determined that the second difference index is greater than the first threshold, then the cumulative difference index is obtained based on the second difference index; If the cumulative difference index is determined to be greater than the second threshold, the target filter is reset.

4. The method as described in claim 1, characterized in that, The current network jitter index is obtained based on the instantaneous current network jitter index, including: Determine the range in which the current instantaneous network jitter index is located; Determine the count of the instantaneous network jitter index within each interval, wherein the instantaneous network jitter index includes the current instantaneous network jitter index; The counts of instantaneous network jitter metrics within each interval are normalized to obtain the current network jitter probability density. The upper quantile of the current network jitter probability density is used as the current network jitter index.

5. The method as described in claim 4, characterized in that, Also includes: If it is determined that the number of packet loss compensations for the audio data within the first predetermined sliding window duration is greater than the first packet loss threshold, then the length is increased by α according to the first step, where α is a real number greater than 0 and less than 1. If it is determined that the number of packet loss compensations for the audio data within the second predetermined sliding window duration is less than or equal to the second packet loss threshold, then α is reduced by the second step size; Wherein, the length of the first step is greater than the length of the second step.

6. The method as described in claim 4, characterized in that, The interval includes the target interval; The upper quantile of the current network jitter probability density (α) is used as the current network jitter index, including: If it is determined that the network jitter index of the audio data within the third predetermined sliding window duration is less than the minimum value of the network jitter index of the audio data within the fourth predetermined sliding window duration, then the count of the instantaneous network jitter index within the target interval is set to zero. Renormalize the instantaneous network jitter index count within each interval to obtain the updated current network jitter probability density; The upper quantile of the updated current network jitter probability density is used as the current network jitter index.

7. The method as described in claim 4, characterized in that, Based on the transmission and reception times of the current audio frame, and the transmission and reception times of the previous audio frame of the audio data, the instantaneous network jitter index is obtained, including: The current audio frame is subjected to a first anomaly detection, which includes at least one of abnormal packet detection, duplicate packet detection, and late packet detection. The instantaneous index of current network jitter is obtained by using the transmission and reception times of the current audio frame detected by the first anomaly, as well as the transmission and reception times of the previous audio frame of the audio data.

8. The method as described in claim 1, characterized in that, Based on the current actual storage capacity and the current target storage capacity of the buffer, the audio frames in the buffer are processed, including: Determine the upper limit and lower limit of the current target storage capacity based on the current target storage capacity. Processing audio frames in the buffer based on the current actual storage capacity and the current target storage capacity of the buffer, further comprising at least one of the following: If the current audio frames are continuous, the current actual storage capacity is greater than the current target storage capacity upper limit, and the audio data does not experience packet loss compensation within the fifth predetermined time period, then the playback of the current audio frames in the buffer area is accelerated. If the current audio frames are consecutive and the current actual storage capacity is less than the current target storage capacity lower limit, then the playback of the current audio frames in the buffer area is slowed down. If the current audio frame is not continuous and the current time is greater than the expected playback time of the current audio frame, then the current audio frame in the buffer area is played. If the current audio frame is not continuous, and the current time is less than the expected playback time of the current audio frame, then packet loss compensation is performed on the audio data based on the current audio frame.

9. The method as described in claim 8, characterized in that, Also includes: Based on the actual playback time and reception time of the current audio frame, obtain the current receiver latency index; Obtain the difference between the current target storage capacity and the current residual jitter index; The estimated playback time of the current audio frame is obtained based on the transmission time of the current audio frame, the current network latency index, the current receiver latency index, and the difference between the current target storage capacity and the current residual jitter index.

10. The method as described in claim 8, characterized in that, The difference between the upper limit of the current target storage capacity and the current target storage capacity is greater than the difference between the lower limit of the current target storage capacity and the current target storage capacity.

11. An audio data processing apparatus, characterized in that, include: A receiving module is used to receive the current audio frame of audio data, wherein the current audio frame includes the transmission time of the current audio frame; The acquisition module is used to estimate the current network latency index using a filter based on the transmission and reception times of the current audio frame. The current network latency index is a smoothed estimate of the difference between the reception and transmission times of historical audio frames, including the current audio frame. The determining module is used to obtain the smooth reception time of the current audio frame based on the current network latency index and the transmission time of the current audio frame; The current actual storage capacity of the buffer is determined based on the expected difference between the smooth reception time and the actual playback time of the current audio frame, wherein the current actual storage capacity is the amount of voice data actually buffered in the buffer. The acquisition module is further configured to obtain the instantaneous network jitter index based on the transmission and reception times of the current audio frame and the transmission and reception times of the previous audio frame of the audio data. The acquisition module is further configured to obtain the current network jitter index based on the current instantaneous network jitter index; The acquisition module is further configured to determine the current target storage capacity of the cache based on the current network jitter index; A caching module is used to cache the current audio frame into the buffer area; The processing module is used to process audio frames in the buffer based on the current actual storage capacity and the current target storage capacity of the buffer, wherein the audio frames include the current audio frame.

12. A computer device, characterized in that, Includes processor, memory, and input / output interfaces; The processor is connected to the memory and the input / output interface respectively, wherein the input / output interface is used to receive data and output data, the memory is used to store computer programs, and the processor is used to call the computer programs so that the computer device executes the audio data processing method according to any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded and executed by a processor to cause a computer device having the processor to perform the audio data processing method according to any one of claims 1-10.

14. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the audio data processing method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Method and device for processing network jitter

    CN110875860A

  • Live broadcast chorus method and device, electronic equipment and storage medium

    CN110992920A