Audio decoding anti-jitter method and device, electronic equipment, storage medium and program
By adjusting the audio sampling rate to match the audio output rate, the jitter problem caused by instability in the audio decoding process was solved, achieving smooth audio data output and improved quality.
Patent Information
- Application Number
- CN202511412942.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-12-05
AI Technical Summary
In the audio signal transmission and playback chain, instability in the decoding stage can cause audio jitter, especially in unstable network environments, leading to abnormal audio playback and affecting user experience.
By determining the audio data reception stability index and buffer index, the audio sampling rate is adjusted to match the audio output rate, thus achieving smooth audio data output.
It improves the output quality of audio data, reduces playback anomalies, and enhances the user experience.
Smart Images

Figure CN121075342A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer application, and in particular to an audio decoding anti-jitter method and device, electronic equipment, storage medium and program. BACKGROUND
[0002] In the transmission and playback link of the audio signal, the decoding link is the core step of connecting the compressed code stream and the audible audio signal: the audio compressed code stream output by the upstream device is parsed into the pulse code modulation signal through the decoding module, and then is converted into the analog sound signal through the audio rendering module. In this process, the stability of the decoded code stream directly determines the fluency of the audio output. Once the code stream fluctuates or is abnormal, the "audio jitter" problem is easily caused, which is specifically manifested as the playback time axis of the audio signal being offset, the sampling rate and the playback clock being mismatched, the audio frame being lost or repeated, etc. It seriously damages the user's auditory experience, and in real-time voice communication, audio live broadcast and other scenarios, it may even cause information transmission interruption or business function failure. The technical causes of the instability of the decoded code stream at present mainly lie in the uncertainty interference of the transmission link (including the network instability and the instability of other audio transmission communication modes, including but not limited to serial communication, PCIE communication, radio frequency communication); taking the network instability as an example: in the scenarios such as streaming media playback and real-time voice which rely on network transmission, network bandwidth fluctuation, data packet loss, delay, out-of-order, signal attenuation, etc. will cause the time interval of the audio code stream reaching the decoding end to be uneven. For example, when the network is congested, the code stream transmission rate suddenly decreases, and the decoding module cannot continuously obtain data, causing "waiting", which causes audio stuttering; and after the network recovers, a large amount of accumulated code streams rush in, which causes the decoding module to overload, causing audio frame skipping, forming the jitter phenomenon of "stuttering-fast forwarding" alternation, and finally causing abnormal audio playback.
[0003] In summary, an optimization technology capable of adaptively coping with transmission fluctuation is urgently needed to realize stable audio decoding output in all scenarios. SUMMARY
[0004] The present application provides an audio decoding anti-jitter method, device, electronic equipment, storage medium and program, which adjusts the sampling rate of the decoded audio data according to the reception and buffering of the audio data before decoding, improves the matching degree of the audio data sampling point number and the audio output rate, enhances the smoothness of the audio output, prevents abnormal audio data output, and improves the audio playback quality.
[0005] According to an aspect of the present application, an audio decoding anti-jitter method is provided, wherein the method comprises:
[0006] determining the reception stability index and the audio buffering index of the audio data to be decoded;
[0007] The method for adjusting the audio sampling rate is determined based on the received stability index and the audio buffer index;
[0008] The audio sampling rate of the decoded audio data is adjusted according to the adjustment method described above.
[0009] According to another aspect of the present invention, an audio decoding anti-jitter device is provided, wherein the device comprises:
[0010] The indicator acquisition module is used to determine the reception stability indicator and audio buffer indicator of the audio data to be decoded;
[0011] The adjustment strategy module is used to determine the adjustment method of the audio sampling rate based on the receiving stability index and the audio buffer index;
[0012] The adjustment execution module is used to adjust the audio sampling rate of the decoded audio data according to the adjustment method.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the audio decoding anti-jitter method according to any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the audio decoding anti-jitter method according to any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer program product is provided, wherein the computer program product includes a computer program that, when executed by a processor, implements the audio decoding anti-jitter method according to any embodiment of the present invention.
[0019] The technical solution of this invention involves statistically analyzing the reception stability index and audio buffer index of the audio data to be decoded, determining the adjustment method for the audio sampling rate of the audio data based on the statistically obtained reception stability index and audio buffer index, and adjusting the decoded audio data according to the adjustment method. By changing the sampling frequency of the audio data, this invention can achieve dynamic scaling of the audio data, improve jitter caused by unstable audio buffering or transmission, achieve balanced output of audio data, reduce playback abnormalities of audio data, and enhance the user experience.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of an audio decoding anti-jitter method provided in Embodiment 1 of the present invention;
[0023] Figure 2 This is a flowchart of another audio decoding anti-jitter method provided in Embodiment 2 of the present invention;
[0024] Figure 3 This is a flowchart of another audio decoding anti-jitter method provided in Embodiment 3 of the present invention;
[0025] Figure 4 This is a flowchart of another audio decoding anti-jitter method provided in Embodiment 4 of the present invention;
[0026] Figure 5 This is a flowchart of another audio decoding anti-jitter method provided in Embodiment 5 of the present invention;
[0027] Figure 6 This is a flowchart of a receiving stability index calculation method provided in Embodiment 5 of the present invention;
[0028] Figure 7 This is a schematic diagram of the structure of an audio decoding anti-jitter device according to Embodiment Six of the present invention;
[0029] Figure 8This is a schematic diagram of the structure of an electronic device that implements the audio decoding anti-jitter method of the present invention. Detailed Implementation
[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0032] Example 1
[0033] Figure 1 This is a flowchart of an audio decoding anti-jitter method according to Embodiment 1 of the present invention. This embodiment is applicable to audio decoding scenarios. The method can be executed by an audio decoding anti-jitter device, which can be implemented in hardware and / or software. This device can be configured in a terminal device or server. Figure 1 As shown, the method includes:
[0034] Step 110: Determine the reception stability index and audio buffer index of the audio data to be decoded.
[0035] Audio data can be a digital carrier of sound signals. In scenarios such as movies, short videos, or video conferencing, audio data can be encapsulated together with video data. Audio data can include various formats, including but not limited to OPUS, OGG, or FLAC. Reception stability metrics are indicators that measure whether audio data is transmitted stably. These metrics can be obtained through statistical analysis of audio data latency, transmission interval, packet loss rate, and other indicators. Audio buffering metrics indicate the buffering status of audio data and can be determined by the audio data buffer settings and the real-time buffer size.
[0036] In this embodiment of the invention, after acquiring the audio data, the transmission stability index and audio buffer index of the audio data to be decoded can be statistically analyzed. For example, parameters such as latency, transmission interval, and packet loss rate of the audio data within a specific time period can be statistically analyzed. The weighted sum of the statistical results of different types of parameters can be used as the reception stability index. Of course, only parameters such as packet loss rate or transmission interval can be statistically analyzed as the reception stability index. The audio buffer index can include the quotient of buffer size and audio data frame or the real-time buffering number of audio data frames. In some embodiments, the reception stability index of audio data can also be determined based on the number of audio data received per unit time. For example, the number of audio data frames received per unit time can be used as the reception stability index.
[0037] Step 120: Determine the adjustment method for the audio sampling rate based on the reception stability index and the audio buffer index.
[0038] The audio sampling rate refers to the number of times a continuously changing analog audio signal is sampled per unit time. The unit of audio sampling rate is Hertz (Hz) or kilohertz (KHz), etc. The audio sampling rate is used in the audio data decoding process to reconstruct the analog audio signal from the audio data. The adjustment method can be a way to adjust the audio sampling rate, such as increasing or decreasing the sampling frequency based on the original audio sampling rate.
[0039] In this embodiment of the invention, the audio sampling rate can be adjusted based on the acquired stability index and audio buffer index to adjust the audio sampling rate of the decoded audio data, thereby changing the sampling rate during the audio data output process. By adjusting the sampling rate, the number of sampling points in the decoded audio data is changed, achieving a scaling effect on the audio data and smoothing the audio data output, thus enhancing the quality of the audio data output. For example, determining the adjustment method of the audio sampling rate based on the receiving stability index and audio buffer index may include pre-training a deep learning model. The deep learning model can predict the stability index and audio buffer index to obtain the corresponding increase or decrease in the audio sampling rate as the adjustment method. Alternatively, a mapping relationship can be constructed between the adjustment amount of the audio sampling rate change frequency and the receiving stability index, and between the frequency adjustment direction of the audio sampling rate and the audio buffer index. The acquired receiving stability index and audio buffer index can be substituted into the above mapping relationship to determine the frequency adjustment direction and the adjustment amount of the frequency change rate of the audio sampling rate as the adjustment method.
[0040] Step 130: Adjust the audio sampling rate of the decoded audio data according to the adjustment method.
[0041] Specifically, the audio sampling rate of the decoded audio data can be adjusted according to a determined adjustment method. After decoding, the number of sampling points in the audio data can be increased or decreased by changing the audio sampling rate, achieving a scaling effect. This allows the decoded audio data to be output to match the software and / or hardware output rate, thereby reducing playback anomalies and achieving smooth audio output. For example, the frequency of audio sampling rate changes can be increased or decreased according to a determined adjustment method. This frequency of change is achieved by increasing or decreasing the rate of change. This allows the audio signal to be restored according to the adjusted audio sampling rate, thus achieving smooth audio data transmission.
[0042] This invention, through statistical analysis of the reception stability index and audio buffer index of the audio data to be decoded, determines the adjustment method for the audio sampling rate of the audio data based on the statistically obtained reception stability index and audio buffer index. The audio sampling rate of the decoded audio data is then adjusted according to this adjustment method. This invention determines the stability of the received audio data based on the reception stability index and the buffer status of the received audio data to be decoded based on the audio buffer status. The sampling rate of the encoded audio data is adjusted according to the stability index and audio buffer status, ensuring that the number of sampling points matches the audio output rate. This achieves precise scaling control of the decoded audio data, improves abnormal audio data conditions, achieves smooth audio data output, enhances the output quality of the audio data, and improves the user experience.
[0043] Example 2
[0044] Figure 2 This is a flowchart of another audio decoding anti-jitter method provided by Embodiment 2 of the present invention. This embodiment is a concretization based on the above embodiments, describing the process of determining the receiving stability index and the audio buffer index. See [link to documentation]. Figure 2 The method provided in this embodiment of the invention specifically includes the following steps:
[0045] Step 210: Calculate the transmission interval of audio data frames according to a preset time interval, and use the standard deviation of each transmission interval as an indicator of reception stability.
[0046] The audio data frame can carry structured data blocks of audio data. The audio data frame includes audio information of fixed duration or fixed data amount within the audio data. The audio data can be divided into multiple audio data frames. The transmission interval can be the transmission interval between two adjacent audio data frames. This transmission interval can be determined by the time difference of the received audio data frames. It can be understood that the received audio data frame can be audio data encoded according to a certain specific format, and the audio data frame has not yet been decoded.
[0047] In this embodiment of the invention, the transmission interval of received audio data frames within a preset time interval can be statistically analyzed. For the statistically analyzed transmission interval of the audio data frames, a standard deviation can be calculated. The standard deviation of the transmission interval can be used as a reception stability indicator. The preset time interval can be configured as needed, and its granularity can be minutes, seconds, etc. In some embodiments of the invention, statistically analyzing the transmission interval of audio data frames according to a preset time interval and using the standard deviation of each transmission interval as a reception stability indicator includes:
[0048] The transmission intervals between each audio data frame received within a preset time interval are statistically analyzed, and a first ratio between the sum of each transmission interval and the total number of transmission intervals is determined. The square of the difference between each transmission interval and the ratio is determined, and the sum of the squares corresponding to each transmission interval is obtained. A second ratio between the sum and the total number of transmission intervals is obtained, and the square root of the second ratio is used as a reception stability index.
[0049] Step 220: At the statistical moment of the reception stability index, obtain the number of buffered audio data frames of the audio data, and use the number of buffers as the audio buffer index.
[0050] The statistical time can be the moment when the reception stability index is triggered to be statistically analyzed. This statistical time can be determined by a preset time interval. For example, the reception stability index can be statistically analyzed once every preset time interval. At the same time, the determination of the audio buffer index can also be triggered.
[0051] In this embodiment of the invention, when each reception stability indicator is statistically analyzed, the number of audio data frames stored in the buffer can also be counted for the audio data, and this number of buffers can be used as the audio buffering indicator for the audio data frames. It is understood that for each statistically analyzed reception stability indicator, the number of audio data frames buffered at the same time can be counted as the audio buffering indicator.
[0052] Step 230: Determine the adjustment method for the audio sampling rate based on the reception stability index and the audio buffer index.
[0053] Step 240: Adjust the audio sampling rate of the decoded audio data according to the adjustment method.
[0054] This invention, in its embodiments, uses the transmission interval of audio data frames within a preset time interval as a statistical indicator and determines the standard deviation of the transmission interval as a reception stability index. It also uses the number of buffered audio data frames obtained at the statistical time of the reception stability index as an audio buffering index. Based on the reception stability index and the audio buffering index, it determines the adjustment method for the audio sampling rate and adjusts the audio sampling rate of the decoded audio data accordingly. This invention quantifies the stability and buffering status of audio data reception by using the transmission interval of audio data frames and the number of buffered audio data frames. Based on the stability and buffering status, it determines the adjustment method for the audio sampling rate of the decoded audio data. This adjustment method ensures that the number of sampling points in the decoded audio data matches the audio output speed, improving the scaling effect of the decoded audio data, enhancing the matching degree between the audio data and the audio output speed, improving the smoothness of the audio data output, reducing abnormal playback of audio data, and improving the audio output quality.
[0055] Example 3
[0056] Figure 3 This is a flowchart of another audio decoding anti-jitter method provided in Embodiment 3 of the present invention. This embodiment is a concretization based on the above-described embodiments, describing the process of determining the audio sampling rate adjustment method. See [link to documentation]. Figure 3 The method provided in this embodiment of the invention specifically includes the following steps:
[0057] Step 310: Determine the reception stability index and audio buffer index of the audio data to be decoded.
[0058] Step 320: Determine the adjustment value of the frequency change rate of the audio sampling rate based on the changes of each receiving stability index over time.
[0059] The frequency change rate can be information indicating the magnitude of the change in the audio sampling rate. The larger the frequency change rate, the greater the degree of change in the audio sampling rate. A larger frequency change rate can include an increase in the degree of increase or decrease in the audio sampling rate value, while a smaller frequency change rate can include a decrease in the degree of increase or decrease in the audio sampling rate value. The frequency change rate can be the amount of change in the audio sampling rate per unit time.
[0060] In this embodiment of the invention, the collected reception stability index can be statistically analyzed to determine its changes over time. These changes can include whether the reception stability index tends to be stable or unstable over time. The adjustment value of the frequency change rate of the audio sampling rate can be determined according to these changes. For example, if the change is becoming more and more stable over time, the adjustment value should be smaller. Conversely, if the change is becoming more and more unstable over time, the adjustment value should be larger.
[0061] Step 330: Determine the adjustment direction of the frequency change rate of the audio sampling rate based on the fit between the audio buffer index and the preset standard range.
[0062] The preset standard range allows access to a pre-configured number of buffered audio data frames. This preset standard range can be configured based on parameters such as the audio data frame format and transmission link status. The matching condition indicates the degree to which the audio buffer index matches the preset standard range. Matching conditions can include the audio buffer index being less than, greater than, or within the preset standard range. The adjustment direction indicates how the frequency change rate of the audio sampling rate is changed. This adjustment direction can include increasing, decreasing, or remaining unchanged; that is, the frequency change rate of the audio sampling rate can be changed by increasing, decreasing, or not changing the audio sampling rate.
[0063] In this embodiment of the invention, the audio buffer index can be compared with a preset standard range to determine the fit between the audio buffer index and the preset standard range. If the fit is that the audio buffer index is less than the lower limit of the preset standard range, then the number of audio data frames to be buffered needs to be increased. In this case, the adjustment direction of the frequency change rate of the audio sampling rate can be determined to be increased, that is, the frequency change rate can be changed by increasing the audio sampling rate. If the fit is that the audio buffer index is greater than the upper limit of the preset standard range, then the number of audio data frames to be buffered needs to be reduced. In this case, the adjustment direction of the frequency change rate of the audio sampling rate can be determined to be decreased, that is, the frequency change rate can be changed by decreasing the audio sampling rate.
[0064] Step 340: Adjust the audio sampling rate of the decoded audio data according to the adjustment method.
[0065] In this embodiment of the invention, the audio sampling rate can be adjusted according to the determined adjustment method, including the frequency change rate and the adjustment direction. For example, the original audio sampling rate of the audio data is 50 Hz, the frequency change rate is 100 Hz / s, and the adjustment direction is to increase, that is, the audio sampling rate starts at 50 Hz and increases by 100 Hz per second, so the audio sampling rate after 1 second can be 150 Hz.
[0066] In this embodiment of the invention, by acquiring the reception stability index and audio buffer index of the audio data to be decoded, the adjustment value of the frequency change rate of the audio sampling rate is determined according to the change of the reception stability index over time. Based on the fit between the audio buffer index and a preset standard range, the adjustment direction of the frequency change rate of the audio sampling rate is determined. The audio sampling rate of the decoded audio data is adjusted according to the determined adjustment value and adjustment direction, so that the adjustment amount and adjustment method of the audio sampling rate precisely match the actual reception and buffering conditions of the audio data. This allows for precise adjustment of the audio sampling rate, ensuring that the number of sampling points after audio sampling rate processing matches the output rhythm of the audio data. This enables scalable processing of the audio data, avoiding audio playback abnormalities caused by too much or too little decoded audio data, and controlling the smooth output of the decoded audio data to improve audio playback quality.
[0067] Example 4
[0068] Figure 4 This is a flowchart of another audio decoding anti-jitter method provided in Embodiment 3 of the present invention. This embodiment is a concretization based on the above-described embodiments, describing the process of determining the audio sampling rate adjustment method. See [link to documentation]. Figure 4 The method provided in this embodiment of the invention specifically includes the following steps:
[0069] Step 410: Determine the reception stability index and audio buffer index of the audio data to be decoded.
[0070] Step 420: The difference between the reception stability index in this statistical analysis and the reception stability index in the previous statistical analysis is taken as the change.
[0071] In this embodiment of the invention, the reception stability index of the current statistics and the reception stability index of the previous statistics are obtained, and the difference between the reception stability index of the current statistics and the reception stability index of the previous statistics is calculated. It can be understood that the difference can be specifically the reception stability index of the current statistics minus the reception stability index of the previous statistics, and the obtained difference can be used as the change of the reception stability index.
[0072] Step 430: If the change is greater than 0, the adjustment value is determined to be the first preset adjustment value, where the first adjustment value is positive; if the change is less than 0, the adjustment value is determined to be the second preset adjustment value, where the second adjustment value is negative; if the change is equal to 0, the adjustment value is determined to be 0.
[0073] The first preset adjustment value and the second preset adjustment value can be pre-configured changes to the frequency change rate of the audio sampling rate. The first preset adjustment value can correspond to an increase in the audio sampling rate, while the second preset adjustment value can correspond to a decrease in the audio sampling rate. The first preset adjustment value and the second preset adjustment value can be configured based on experience.
[0074] In this embodiment of the invention, if the change is greater than 0, that is, the difference between the current reception stability index and the previous reception stability index is greater than 0, the reception stability index increases over time, indicating that the stability of audio data transmission deteriorates. In this case, the change in the frequency variation rate of the audio sampling rate can be increased to make the decoded audio data closer to the audio output rhythm by significantly changing the audio sampling rate. If the change is less than 0, the index difference is less than 0, the reception stability index decreases over time, indicating that the stability of audio data transmission improves. In this case, the change in the frequency variation rate of the audio sampling rate can be reduced to minimize the change in the audio sampling rate and maintain a smooth output state of the audio data. If the change is 0, the current frequency variation rate can be maintained, that is, the audio sampling rate can be adjusted within a specified range of change to keep the decoded audio data in a stable audio output state.
[0075] Step 440: Obtain the upper and lower limits of the preset standard range, and determine the relationship between the audio buffer index and the upper and lower limits as the fit condition.
[0076] Specifically, the upper and lower limits of a preset standard range can be obtained, and the determined audio buffer index can be compared with the upper and / or lower limits respectively. The relationship between the upper and lower limits of the preset standard range of the audio buffer index is used as the fit condition.
[0077] Step 450: If the matching condition is that the audio buffer index is less than the lower limit, then the adjustment direction of the frequency change rate is determined to be decreasing; if the matching condition is that the audio buffer index is greater than the upper limit, then the adjustment direction of the frequency change rate is determined to be increasing; if the matching condition is that the audio buffer index is greater than or equal to the lower limit and less than or equal to the upper limit, then the adjustment direction of the frequency change rate is determined to be unchanged.
[0078] Specifically, if the audio buffer index is less than the lower limit and the number of buffered audio data frames is less than the preset standard range, the number of buffered audio data frames can be increased by reducing the audio sampling rate. Therefore, the adjustment direction of the frequency change rate is determined to be decreasing, i.e., changing the frequency change rate by reducing the audio sampling rate. If the audio buffer index is greater than the upper limit and the number of buffered audio data frames is higher than the preset standard range, the number of buffered audio data frames can be decreased by increasing the audio sampling rate. Therefore, the adjustment direction of the frequency change rate is determined to be increasing, i.e., changing the frequency change rate by increasing the audio sampling rate. If the audio buffer index is greater than or equal to the lower limit and less than or equal to the upper limit, i.e., the audio buffer index is within the preset standard range, then the frequency change rate does not need to be changed; i.e., the adjustment direction remains unchanged.
[0079] In this embodiment of the invention, the obtained adjustment value of the frequency change rate and the adjustment direction can be used as the method for adjusting the audio sampling rate.
[0080] Step 460: Adjust the audio sampling rate of the decoded audio data according to the adjustment method.
[0081] In this embodiment of the invention, by collecting the reception stability index and audio buffer index of the audio data to be decoded, the difference between the current reception stability index and the previous reception stability index is used as the change of the reception stability index over time. According to the relationship between the change and 0, a first preset adjustment value or a second preset adjustment value is selected as the adjustment value for the frequency change rate of the audio sampling rate. The relationship between the audio buffer index and the upper and lower limits of the preset standard range is obtained as the fit. According to the fit, the adjustment direction of the frequency change rate of the audio sampling rate is determined to be decreasing, increasing, or remaining unchanged. The audio sampling rate of the decoded audio data is adjusted according to the adjustment direction and adjustment amount of the frequency change rate in the determined adjustment method. This invention allows for the adjustment of the audio sampling rate of the decoded audio data based on the determined adjustment amount and direction of the frequency change rate. This ensures that the adjustment amount and method of the audio sampling rate precisely match the actual reception and buffering conditions of the audio data, thereby precisely adjusting the audio sampling rate. This ensures that the number of sampling points in the decoded audio data after audio sampling rate processing matches the output rhythm of the audio data. It enables the scaling of audio data, avoids audio playback abnormalities caused by too much or too little decoded audio data, and controls the smooth output of decoded audio data, thus improving audio playback quality.
[0082] Furthermore, based on the above embodiments of the invention, it also includes: if the index difference corresponding to the change is greater than a preset difference, then the audio sampling rate of the audio data is not adjusted.
[0083] The preset difference can be the critical reception stability index value corresponding to the maximum change in the audio sampling rate, and this maximum change can correspond to the maximum tolerance for audio changes.
[0084] Specifically, the index difference can be compared with the preset difference. If the index difference is greater than the preset difference, the change in the audio sampling rate of the audio data is too large, resulting in obvious pitch distortion and affecting the audio output quality. Therefore, the audio sampling rate of the audio data can be left unadjusted.
[0085] Example 5
[0086] In this embodiment of the invention, the audio jitter problem caused by unstable transmission networks or excessive decoded audio data frames is solved by adjusting the audio sampling rate of the audio data. Specifically, this embodiment of the invention achieves the purpose of audio data length scaling by fine-tuning the audio sampling rate of the decoded audio data frames. This can improve the audio data playback abnormality problem caused by insufficient audio data or overflow of buffered data frames, and achieve a balanced audio output effect. This embodiment of the invention can statistically analyze the reception stability index and audio buffer index of the decoded audio data stream. The reception stability index is determined based on the standard deviation of the transmission time interval of the audio data frames, that is:
[0087]
[0088] Where n represents the total number of time intervals counted, x i The time interval length is obtained for the i-th statistical time interval, where i is an integer from 0 to n. The statistical time interval can be adjusted according to the actual situation. For example, 2 seconds can be used to count the time interval of audio data frame transmission.
[0089] In this embodiment of the invention, the transmission interval x between each audio data frame received within a preset time interval is statistically defined. i And determine the sum of each transmission interval. The first ratio between the transmission interval and the total number of transmission intervals n; determine the square of the difference between each transmission interval and the ratio, and obtain the sum of the squares corresponding to each transmission interval. ; Obtain the second ratio of the sum to the total number of transmission intervals, and use the square root of the second ratio, stdev, as the receiver stability index.
[0090] In other embodiments of the invention, the reception stability index may include the number of audio data frames received per unit time. Specifically, the number of buffered audio data frames before audio decoding can also be statistically analyzed as an audio buffer index. The direction of frequency adjustment of the audio sampling rate is determined based on the initially set upper and lower limits of the buffer and the audio buffer index. For example, if the number of buffered audio data frames before audio decoding is 8 frames, and the upper limit of the buffer is 5 frames and the lower limit is 3 frames, if the number of frames exceeds the upper limit, the audio sampling rate can be increased to reduce the buffer length of the audio data frames; conversely, if the number of frames is below the lower limit, the audio sampling rate can be decreased to increase the buffer length of the audio data frames.
[0091] In this embodiment of the invention, the audio sampling rate conversion can use the audio sampling rate of the audio data decoding or playback as the base sampling rate, with an adjustment range of ±5%. It is understood that the upper and lower limits of the adjustment range can be further limited according to actual conditions. The larger the adjustment range, the more obvious the pitch change of the output audio signal. This embodiment of the invention can adjust the frequency variation of the audio sampling frequency based on the actual reception stability index. The larger the reception stability index, the worse the stability, and the larger the required adjustment range of the sampling rate. The audio buffer index indicates the direction of the audio sampling rate frequency adjustment. If the audio buffer index is close to the upper limit, the audio sampling rate is increased; if it is not less than the upper limit, the audio sampling rate is adjusted to the maximum value. Conversely, if the audio buffer index is close to the lower limit, the audio sampling rate is decreased; if it is not greater than the lower limit, the audio sampling rate is adjusted to the minimum value.
[0092] This invention can be applied to scenarios with poor network conditions, such as real-time conferencing or live video streaming. See [link / reference]. Figure 5 This invention can be applied to the decoding end of audio decoding data. The data receiving module in the decoding end receives the audio data to be decoded and buffers it to balance the difference between data transmission and subsequent decoding speeds, preventing data interruption or overflow. Then, the audio decoding module performs format parsing and audio decoding on the audio data, replacing it with uncompressed raw audio data. The audio processing module optimizes and adapts the audio data based on the audio sampling frequency, improving output sound quality or adapting to hardware characteristics. Finally, the audio data processed by the audio processing module is mixed and played back. This audio decoding anti-jitter method may include the following steps:
[0093] The audio data receiver receives audio data and decodes it. If the audio decoding anti-jitter method provided in this embodiment of the invention is not enabled, the decoded audio data can be mixed and then played back.
[0094] If the audio decoding anti-jitter method provided in this embodiment of the invention is enabled, the statistical module can statistically analyze the audio data reception stability index and audio buffer index during the audio data reception and decoding process. Two parameters for adjusting the audio sampling rate—namely, the frequency change rate Δ and the adjustment direction—can then be determined using these indicators. The audio processing module can obtain the frequency change rate and adjustment direction generated by the statistical module and adjust the audio sampling rate of the decoded audio data according to these indicators, thus achieving anti-jitter analysis of the audio data. It can be understood that the statistical process for the reception stability index and audio buffer index can be executed in the data reception module; that is, the reception stability index and audio buffer index can be obtained by statistically analyzing the undecoded audio data.
[0095] Based on the above embodiments of the invention, the calculation process of the receiving stability index within the statistics module can be as follows: Figure 6 As shown, obtain the current timestamp ct; obtain the time interval dt = lt - ct between the previous audio data frame and the current audio data frame, and calculate the cumulative time interval st = st + dt; obtain the mean at of the previous time interval and calculate the variance var = pow(dt - at). 2 ); Determine the cumulative variance and sum = sum + var; Determine if the statistical time period has been reached. If not, continue to count the time intervals. If the statistical time period has been reached, calculate the standard deviation stdev = sqrt(sum / n), calculate the mean of the time interval at = st / n, and reset st, sum, and n.
[0096] In this embodiment of the invention, anti-jitter is achieved by adjusting the audio sampling rate of the decoded audio data. The audio sampling rate of the decoded audio data can be adjusted based on the reception stability and buffering situation of the audio data. Increasing or decreasing the audio sampling rate will cause the number of sampling points of the decoded audio data to be adjusted accordingly with the audio sampling rate, ensuring that the number of sampling points of the decoded audio data matches the audio output rate. For example, in the case of poor transmission stability and excessive buffering, the number of sampling points of the audio data can be increased by increasing the sampling rate by a large change, thereby speeding up the audio data output and reducing the amount of buffered audio data, thus releasing the accumulated decoding buffer. Conversely, when the transmission stability is good and the buffer is too small, the number of sampling points of the decoded audio data can be reduced by decreasing the sampling rate by a small change, thereby increasing the amount of buffered audio data and accumulating the accumulated audio data.
[0097] Example 6
[0098] Figure 7 This is a schematic diagram of the structure of an audio decoding anti-jitter device provided according to Embodiment Six of the present invention. Figure 7 As shown, the device includes:
[0099] The indicator acquisition module 510 is used to determine the reception stability indicator and audio buffer indicator of the audio data to be decoded.
[0100] The adjustment strategy module 520 is used to determine the adjustment method of the audio sampling rate based on the reception stability index and the audio buffer index.
[0101] The adjustment execution module 530 is used to adjust the audio sampling rate of the decoded audio data according to the adjustment method.
[0102] In this embodiment of the invention, the indicator acquisition module statistically analyzes the reception stability index and audio buffer index of the audio data to be decoded. The adjustment strategy module determines the adjustment method for the audio sampling rate of the audio data based on the statistically obtained reception stability index and audio buffer index. The adjustment execution module adjusts the audio sampling rate of the decoded audio data according to the adjustment method. This embodiment of the invention determines the stability of the received audio data by using the reception stability index and the buffer status of the received audio data to be decoded by using the audio buffer status. The sampling rate of the encoded audio data is adjusted according to the stability index and the audio buffer status, so that the number of sampling points of the audio data matches the audio output rate. This achieves precise scaling control of the decoded audio data, improves abnormal situations of the audio data, achieves smooth output of the audio data, enhances the output quality of the audio data, and enhances the user experience.
[0103] In some embodiments of the invention, the indicator acquisition module 510 includes:
[0104] The stability unit is used to count the transmission interval of audio data frames according to a preset time interval, and use the standard deviation of each transmission interval as the reception stability index.
[0105] The buffer unit is used to obtain the number of buffered audio data frames for the statistical time of the reception stability index, and uses the number of buffers as the audio buffer index.
[0106] In some embodiments of the invention, the adjustment strategy module 520 includes:
[0107] The adjustment unit is used to determine the adjustment value of the frequency change rate of the audio sampling rate based on the changes of various reception stability indicators over time.
[0108] The adjustment direction unit is used to determine the adjustment direction of the frequency change rate of the audio sampling rate based on the fit between the audio buffer index and the preset standard range.
[0109] In some embodiments of the invention, the adjustment unit is specifically used to: take the difference between the current reception stability index and the previous reception stability index as the change; if the change is greater than 0, then determine the adjustment value as a first preset adjustment value, wherein the first adjustment value is a positive value; if the change is less than 0, then determine the adjustment value as a second preset adjustment value, wherein the second adjustment value is a negative value; if the change is equal to 0, then determine the adjustment value as 0.
[0110] In some embodiments of the invention, the adjustment direction unit is specifically used to: obtain the upper limit and lower limit of a preset standard range, and determine the relationship between the audio buffer index and the upper and lower limits as a matching condition; if the matching condition is that the audio buffer index is less than the lower limit, then the adjustment direction of the frequency change rate is determined to be decreasing; if the matching condition is that the audio buffer index is greater than the upper limit, then the adjustment direction of the frequency change rate is determined to be increasing; if the matching condition is that the audio buffer index is greater than or equal to the lower limit and less than or equal to the upper limit, then the adjustment direction of the frequency change rate is determined to be unchanged.
[0111] In some embodiments of the invention, the adjustment unit is further configured to: determine that if the index difference corresponding to the change is greater than a preset difference, then not adjust the audio sampling rate of the audio data.
[0112] The audio decoding anti-jitter device provided in this embodiment of the invention can execute the audio decoding anti-jitter method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0113] Example 7
[0114] Figure 8 This is a schematic diagram of an electronic device implementing the audio decoding anti-jitter method of this invention. The electronic device is intended to represent various forms of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0115] like Figure 8 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0116] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0117] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as audio decoding anti-jitter methods.
[0118] In some embodiments, the audio decoding anti-jitter method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the audio decoding anti-jitter method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the audio decoding anti-jitter method by any other suitable means (e.g., by means of firmware).
[0119] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0120] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0121] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0122] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0123] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0124] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0125] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0126] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. An audio decoding method against jitter, characterized by, The method comprises: determining a receiving stability index and an audio buffering index of audio data to be decoded; determining an adjustment mode of an audio sampling rate according to the receiving stability index and the audio buffering index; adjusting the audio sampling rate of the decoded audio data according to the adjustment mode.
2. The method of claim 1, wherein, The determining of the receiving stability index and the audio buffering index of the audio data to be decoded comprises: statistically determining transmission intervals of audio data frames of the audio data according to a preset time interval, and taking a standard deviation of each transmission interval as the receiving stability index; acquiring a buffering quantity of the audio data frames of the audio data at a statistical time of the receiving stability index, and taking the buffering quantity as the audio buffering index.
3. The method of claim 1 or 2, wherein, The determining of the adjustment mode of the audio sampling rate according to the receiving stability index and the audio buffering index comprises: determining an adjustment value of a frequency variation rate of the audio sampling rate according to a variation condition of each receiving stability index over time; determining an adjustment direction of the frequency variation rate of the audio sampling rate based on a fitting condition of the audio buffering index and a preset standard range.
4. The method of claim 3, wherein, The determining of the adjustment value of the frequency variation rate of the audio sampling rate according to the variation condition of each receiving stability index over time comprises: taking an index difference between the current statistical receiving stability index and a previous statistical receiving stability index as the variation condition; if the variation condition is greater than 0, determining that the adjustment value is a first preset adjustment value, wherein the first adjustment value is a positive value; if the variation condition is less than 0, determining that the adjustment value is a second preset adjustment value, wherein the second adjustment value is a negative value; if the variation condition is equal to 0, determining that the adjustment value is 0.
5. The method of claim 3, wherein, The determining of the adjustment direction of the frequency variation rate of the audio sampling rate based on the fitting condition of the audio buffering index and the preset standard range comprises: acquiring an upper limit value and a lower limit value of the preset standard range, and determining a size relationship between the audio buffering index and the upper limit value and the lower limit value as the fitting condition; if the fitting condition is that the audio buffering index is less than the lower limit value, determining that the adjustment direction of the frequency variation rate is to be reduced; if the fitting condition is that the audio buffering index is greater than the upper limit value, determining that the adjustment direction of the frequency variation rate is to be increased; if the fitting condition is that the audio buffering index is greater than or equal to the lower limit value and less than or equal to the upper limit value, determining that the adjustment direction of the frequency variation rate is to be unchanged.
6. The method of claim 4, wherein, Further comprising: if the index difference corresponding to the variation condition is greater than a preset difference value, not adjusting the audio sampling rate of the audio data.
7. An audio decoding anti-shake apparatus characterized by comprising: The device comprises: an index acquisition module configured to determine a receiving stability index and an audio buffering index of audio data to be decoded; an adjustment strategy module configured to determine an adjustment mode of an audio sampling rate according to the receiving stability index and the audio buffering index; an adjustment execution module configured to adjust the audio sampling rate of the decoded audio data according to the adjustment mode.
8. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the audio decoding anti-shake method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing the processor to implement the audio decoding anti-shake method according to any one of claims 1-6 when executed.
10. A computer program product, characterised in that, The computer program product comprises a computer program which, when executed by a processor, implements the audio decoding anti-shake method according to any one of claims 1-6.