Network audio playing method and device, equipment, storage medium and program product

By using audio reduction playback technology based on the reduction ratio in voice room applications, the playback lag caused by network abnormalities is solved, and a smoother and more stable audio playback experience is achieved.

CN120034670APending Publication Date: 2025-05-23BIGO TECH PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510106154.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In voice room applications, when the jitter buffer reaches the start-up condition, the normal playback of the audio frame is directly performed, which can easily cause playback lag in abnormal network conditions and affect the audience's listening experience.

Method used

Through audio deceleration playback based on the deceleration ratio, the audio packets in the jitter buffer can be quickly accumulated to the target cache number, reducing the probability of playback lag, and adaptively lowering the deceleration ratio to reduce the audience's deceleration experience and balance the playback lag experience and deceleration experience.

Benefits of technology

It effectively reduces the probability of playback lag, improves the overall listening experience of users, and ensures the smoothness and stability of audio playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034670A_ABST
    Figure CN120034670A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a network audio playing method and device, equipment, a storage medium and a program product, and the method comprises the steps: receiving an audio data packet sent by a server, and storing the audio data packet in a jitter buffer area; under the condition that the cumulative number of the first audio data packets is smaller than the target cache number set by the jitter buffer area, extracting audio frames corresponding to the audio data packets from the jitter buffer area based on the deceleration proportion to perform deceleration playing; when the current cache number of the audio data packets in the jitter buffer area is greater than or equal to the historical cache number recorded at the previous frame taking moment, and the jitter value of the preset fractile corresponding to the plurality of audio data packets received in the preset time range in the jitter buffer area is less than or equal to the preset jitter threshold value, the jitter value of the preset fractile corresponding to the plurality of audio data packets is obtained; and reducing the speed reduction proportion, and performing speed reduction playing of the audio frame based on the adjusted speed reduction proportion. According to the scheme, the playing lagging experience and the speed reduction experience are effectively balanced, and the overall listening experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of audio technology, and in particular to a network audio playback method, device, equipment, storage medium, and program product. Background Art

[0002] With the continuous development of social media and instant messaging tools, voice rooms have gradually become a mainstream social interaction mode. It can provide a high-quality platform for anchors to share content and for audiences to interact and communicate with each other, so that audiences can quickly obtain the required information and build closer social ties by listening to the anchor's voice live broadcast or directly engaging in discussions on diverse topics and sharing personal insights. Among them, when the audience enters the voice room, in order to improve the audio output rate, it will request the server to send its cached audio data packets, so that the client's player can quickly meet the start-up conditions, and the number of audio data packets requested is the target cache number of the jitter buffer, which is conducive to resisting network jitter and maintaining audio continuity.

[0003] In actual application, the server may not have cached audio data packets or the cached audio data packets may not reach the target cache quantity of the jitter buffer. However, the related technology directly plays the audio frames normally when the jitter buffer reaches the start-up conditions, which can easily cause playback jams in the event of network anomalies, affecting the audience's listening experience. Summary of the invention

[0004] The embodiments of the present application provide a network audio playback method, device and equipment, which solves the problem that the related technology directly performs normal playback of audio frames when the jitter buffer reaches the start-up conditions, which is easy to cause playback jams in the case of network abnormalities, affecting the audience's listening experience. By slowing down the audio playback based on the deceleration ratio, the audio data packets in the jitter buffer are quickly accumulated to the target cache quantity, effectively reducing the probability of playback jams, and adaptively lowering the deceleration ratio to reduce the audience's perceptible deceleration experience, effectively balancing the playback jam experience and the deceleration experience, and improving the user's overall listening experience.

[0005] In a first aspect, an embodiment of the present application provides a network audio playback method, the method comprising:

[0006] Receive an audio data packet sent by a server, and store the audio data packet in a jitter buffer, wherein the audio data packet includes a first audio data packet sent by the server and cached in advance and a second audio data packet sent in real time;

[0007] In a case where the accumulated number of the first audio data packets is less than the target cache number set in the jitter buffer, determining a reference ratio value according to the accumulated number and the target cache number, setting a deceleration ratio to the reference ratio value, and extracting audio frames corresponding to the audio data packets from the jitter buffer for decelerated playback based on the deceleration ratio;

[0008] When the current cache quantity of audio data packets in the jitter buffer is greater than or equal to the historical cache quantity recorded at the previous frame taking moment, and the jitter values ​​of preset percentiles corresponding to multiple audio data packets received within a preset time range in the jitter buffer are less than or equal to a preset jitter threshold, the deceleration ratio is lowered, and the audio frames are played at a reduced speed based on the adjusted deceleration ratio, so that the normal playback of the audio frames is restored when the current cache quantity reaches the target cache quantity.

[0009] In a second aspect, an embodiment of the present application further provides a network audio playback device, the device comprising:

[0010] A data caching module, configured to receive an audio data packet sent by a server and store the audio data packet in a jitter buffer, wherein the audio data packet includes a pre-cached first audio data packet sent by the server and a second audio data packet sent in real time;

[0011] an audio deceleration playback module, configured to, when the cumulative number of the first audio data packets is less than the target cache number set in the jitter buffer, determine a reference ratio value according to the cumulative number and the target cache number, set the deceleration ratio to the reference ratio value, and extract the audio frame corresponding to the audio data packet from the jitter buffer for deceleration playback based on the deceleration ratio;

[0012] The deceleration ratio adjustment module is configured to reduce the deceleration ratio when the current cache quantity of audio data packets in the jitter buffer is greater than or equal to the historical cache quantity recorded at the previous frame taking moment, and the jitter values ​​of preset percentiles corresponding to multiple audio data packets received within a preset time range in the jitter buffer are less than or equal to a preset jitter threshold, and to slow down the playback of audio frames based on the adjusted deceleration ratio, so as to restore normal playback of audio frames when the current cache quantity reaches the target cache quantity.

[0013] In a third aspect, an embodiment of the present application further provides a network audio playback device, the device comprising:

[0014] one or more processors;

[0015] a storage device configured to store one or more programs,

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the network audio playback method described in the embodiment of the present application.

[0017] In a fourth aspect, an embodiment of the present application further provides a non-volatile storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the network audio playback method described in the embodiment of the present application when executed by a computer processor.

[0018] In a fifth aspect, an embodiment of the present application further provides a computer program product, which includes a computer program stored in a computer-readable storage medium, and at least one processor of the device reads and executes the computer program from the computer-readable storage medium, so that the device executes the network audio playback method described in the embodiment of the present application.

[0019] In an embodiment of the present application, an audio data packet sent by a server is received and stored in a jitter buffer, wherein the audio data packet includes a first audio data packet pre-cached and sent by the server and a second audio data packet sent in real time; when the cumulative number of the first audio data packet is less than the target cache number set in the jitter buffer, a reference ratio value is determined based on the cumulative number and the target cache number, the deceleration ratio is set to the reference ratio value, and audio frames corresponding to the audio data packets are extracted from the jitter buffer based on the deceleration ratio for decelerated playback; when the current cache number of audio data packets in the jitter buffer is greater than or equal to the historical cache number recorded at the previous frame acquisition moment, and the jitter values ​​of preset percentiles corresponding to multiple audio data packets received within a preset time range in the jitter buffer are less than or equal to a preset jitter threshold, the deceleration ratio is lowered, and the audio frames are decelerated and played based on the adjusted deceleration ratio, so that the normal playback of the audio frames is restored when the current cache number reaches the target cache number. In the above scheme, when the cumulative number of the first audio data packets is less than the target cache number set in the jitter buffer, it can be regarded that the number of audio data packets accumulated in the jitter buffer is insufficient. The reference ratio value for the initial configuration of the deceleration ratio can be determined based on the cumulative number and the target cache number. The audio frames corresponding to the audio data packets are extracted from the jitter buffer based on the deceleration ratio for decelerated playback, so that the audio data packets in the jitter buffer can be quickly accumulated to the target cache number, effectively reducing the probability of playback jamming. Moreover, when the current cache number of audio data packets in the jitter buffer is greater than or equal to the historical cache number recorded at the previous frame acquisition moment, and the jitter values ​​of preset percentiles corresponding to multiple audio data packets received within a preset time range in the jitter buffer are less than or equal to the preset jitter threshold, it can be regarded that the probability of the risk of playback jamming is low, and the deceleration ratio can be lowered to reduce the audience's perceptible deceleration experience, effectively balance the playback jamming experience and the deceleration experience, and improve the user's overall listening experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A flowchart of a network audio playback method provided in an embodiment of the present application;

[0021] Figure 2 A flowchart of a network audio playback method including a process of determining to increase a deceleration ratio provided in an embodiment of the present application;

[0022] Figure 3 A flowchart of another network audio playback method including a process of determining to increase the deceleration ratio provided in an embodiment of the present application;

[0023] Figure 4 A flowchart of a network audio playback method including a process of lowering a deceleration ratio provided in an embodiment of the present application;

[0024] Figure 5 A flowchart of a network audio playback method including a process of updating a first growth power and a reference ratio value provided in an embodiment of the present application;

[0025] Figure 6 A flowchart of a network audio playback method including a process of increasing a deceleration ratio provided in an embodiment of the present application;

[0026] Figure 7 A flowchart of a network audio playback method including a process of updating a second growth power and a reference ratio value provided in an embodiment of the present application;

[0027] Figure 8 A structural block diagram of a network audio playback device provided in an embodiment of the present application;

[0028] Figure 9 A schematic diagram of the structure of a network audio playback device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the embodiments of the present application, rather than to limit the embodiments of the present application. It should also be noted that, for ease of description, only parts related to the embodiments of the present application are shown in the accompanying drawings, rather than all structures.

[0030] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0031] The network audio playback method provided in the embodiment of the present application can perform deceleration playback of audio frames based on the set deceleration ratio before the audio data packets in the jitter buffer are accumulated to the target cache quantity, so that the jitter buffer can quickly accumulate audio data packets, effectively reduce the probability of playback jams, and adaptively reduce the deceleration ratio, which can reduce the deceleration experience perceivable by the audience and effectively balance the playback jams and deceleration experiences. Relevant application scenarios may include: voice live broadcast, voice social networking, voice teaching, etc. The several application scenarios listed above are only exemplary and explanatory. In actual applications, the network audio playback method can also be used in network audio playback in other scenarios, and the embodiment of the present application does not limit this. The network audio playback method, device, equipment, storage medium and program product provided in the embodiment of the present application are intended to solve the problem that the related technology directly plays the audio frame normally when the jitter buffer reaches the start-up conditions, which is easy to cause playback jams in the case of network abnormalities, affecting the audience's listening experience.

[0032] In the network audio playback method provided in the embodiment of the present application, the executor of each step can be a computer device, which refers to any electronic device with data calculation, processing and storage capabilities, such as PC (Personal Computer), tablet computer and other terminal devices, and the embodiment of the present application does not limit this.

[0033] Figure 1 A flowchart of a network audio playback method provided in an embodiment of the present application. The network audio playback method is applied to a client, such as Figure 1 As shown, the following steps are included:

[0034] Step S101: receiving an audio data packet sent by a server, and storing the audio data packet in a jitter buffer, wherein the audio data packet includes a pre-cached first audio data packet sent by the server and a second audio data packet sent in real time.

[0035] Among them, taking voice live broadcast as an example, when the audience enters the room, the server can first send the first pre-cached audio data packet to the client, so that the client's player can quickly reach the start-up condition, and continue to send the non-pre-cached second audio data packet to the client in real time. The client can cache the received first audio data packet and the second audio data packet to the jitter buffer, and when the audio data packets stored in the jitter buffer reach the preset start-up number, start to extract the audio frames corresponding to the audio data packets from the jitter buffer for playback. In order to improve the audio output rate, the preset start-up number can be less than the target cache number of the jitter buffer, which is dynamically set by the client to adapt to different application scenarios and real-time network conditions, and is used to resist network jitter and ensure the smoothness of audio playback. Therefore, after starting audio playback, the number of audio data packets in the jitter buffer may not reach the target cache number, and it is impossible to better resist network jitter and ensure audio continuity under network abnormal conditions. It is necessary to extract the audio frames corresponding to the audio data packets from the jitter buffer in the subsequent steps S102 and S103 for deceleration playback, so that the audio data packets in the jitter buffer can quickly accumulate to the target cache number.

[0036] Step S102: When the cumulative number of first audio data packets is less than the target cache number set in the jitter buffer, a reference ratio value is determined according to the cumulative number and the target cache number, a deceleration ratio is set to the reference ratio value, and audio frames corresponding to the audio data packets are extracted from the jitter buffer for decelerated playback based on the deceleration ratio.

[0037] Among them, if the cumulative number of the first audio data packet is less than the target cache number set in the jitter buffer, it can be regarded as necessary to perform decelerated playback of the audio frame to quickly accumulate the audio data packets in the jitter buffer. The decelerated playback is specifically to stretch the voice signal in the time domain, thereby lengthening the playback time of the audio data. Its key parameter is the deceleration ratio, which indicates the ratio of the actual playback speed of the audio frame to the original playback speed. For example, the value range of the deceleration ratio can be set to be greater than or equal to 1. When the deceleration ratio is 1, the audio is played at the original speed. When the deceleration ratio is 1.2, the actual audio playback speed is reduced by 20%. Then, playing 1s of audio data can increase the jitter buffer by 200ms of audio data. At this time, the human ear can clearly perceive the deceleration, which can be used as the maximum ratio limit of the deceleration ratio; when the deceleration ratio is 1.01, the audio playback speed is reduced by 1%, then playing 1s of audio data can increase the jitter buffer by 10ms of audio data. At this time, the human ear can basically not perceive the deceleration, which can be used as the minimum ratio limit of the deceleration ratio. In addition, the difference between the cumulative number of first audio data packets and the target cache number can reflect the risk of jamming in audio playback. If the difference between the cumulative number of first audio data packets and the target cache number is large, it can be preliminarily estimated that there is a risk of jamming in audio playback, and a relatively large benchmark ratio value can be determined. If the difference between the cumulative number of first audio data packets and the target cache number is small, it can be preliminarily estimated that there is no risk of jamming in audio playback, and a relatively small benchmark ratio value can be determined, wherein the benchmark ratio value can be used for initial configuration and adjustment of the deceleration ratio. In one embodiment, the benchmark ratio value is determined based on the cumulative number and the target cache number, and the product of the cumulative number and the target cache number and a preset coefficient is compared. When the cumulative number is less than the product of the target cache number and the preset coefficient, the benchmark ratio value is determined as the set maximum ratio limit. When the cumulative number is greater than or equal to the product of the target cache number and the preset coefficient, the benchmark ratio value is determined as the set minimum ratio limit. Among them, the value range of the preset coefficient is 0 to 1. For example, the preset coefficient is set to 0.5. If the target cache quantity is less than the reference quantity, it is preliminarily estimated that there is a risk of jamming in the audio playback, and the maximum proportional limit can be set accordingly as the initial value of the deceleration ratio. For example, the reference proportional value can be 1.2. If the target cache quantity is greater than the reference quantity, it is preliminarily estimated that there is no risk of jamming in the audio playback, and the minimum proportional limit can be set accordingly as the initial value of the deceleration ratio. For example, the reference proportional value can be 1.01. In one embodiment, the reference proportional value is determined based on the cumulative quantity and the target cache quantity, and a plurality of quantity ranges of the preset quantity can be evenly divided according to the target cache quantity. For example, if the target cache quantity is y and the preset quantity is 3, then the evenly divided quantity range can be 0-x. 1 、x 1 -x2 and x 2 -x 3 , the corresponding configured reference ratio values ​​can be arranged from large to small, respectively, 1.2, 1.11, and 1.01, and the corresponding reference ratio value can be selected according to the number range in which the accumulated number is located. Thus, the deceleration ratio can be initially set to the reference ratio value, and the audio frames corresponding to the audio data packets are extracted from the jitter buffer based on the deceleration ratio for decelerated playback, so that the jitter buffer can quickly accumulate audio data packets.

[0038] Step S103: When the current cache quantity of audio data packets in the jitter buffer is greater than or equal to the historical cache quantity recorded at the previous frame taking moment, and the jitter values ​​of preset percentiles corresponding to multiple audio data packets received within a preset time range in the jitter buffer are less than or equal to a preset jitter threshold, the deceleration ratio is lowered, and the audio frames are played at a reduced speed based on the adjusted deceleration ratio, so that the normal playback of the audio frames is restored when the current cache quantity reaches the target cache quantity.

[0039] Among them, after each time the audio frame corresponding to the audio data packet is extracted from the jitter buffer based on the deceleration ratio for deceleration playback, the risk of audio playback jamming can be synchronously detected, and the deceleration ratio can be adaptively adjusted. Since the audio data packets sent by the server in real time will be continuously stored in the jitter buffer, and the audio frames corresponding to the audio data packets will be continuously extracted from the jitter buffer for deceleration playback, the current cache quantity of the jitter buffer will be in dynamic change. If the current cache quantity of the audio data packets in the jitter buffer is greater than or equal to the historical cache quantity recorded at the previous frame-taking moment, it can be regarded that the consumption speed of the audio data packets in the dynamic buffer is slower than the cache speed, and the current cache quantity is in an increasing trend. The previous frame-taking moment is the time node when the audio frame was extracted from the jitter buffer for playback last time, and if the jitter values ​​of the preset percentiles corresponding to the multiple audio data packets received within the preset time range in the jitter buffer are less than or equal to the preset jitter threshold, it can be regarded as the network jitter degree is small. Therefore, it can be comprehensively determined that there is no risk of audio playback jamming at present, and the deceleration ratio can be lowered to maintain the jitter buffer to continue to accumulate audio data packets while reducing the deceleration experience that the audience can perceive. Among them, since the client calculates the corresponding jitter value each time it receives an audio data packet, the calculation formula of the jitter value is as follows: i =(recvTime i -sendTime i )-(recvTime i-1 -sendTime i-1 ) Among them, jitter i is the jitter value of the current audio data packet, recvTime iThe receiving time of the current audio data packet, sendTime i The sending time of the current audio data packet, recvTime i-1 The receiving time of the previous audio data packet, sendTime i-1 is the sending time of the previous audio data packet. Then, multiple jitter values ​​corresponding to multiple audio data packets received within a preset time range can be collected, and the multiple jitter values ​​can be sorted from small to large, so that the jitter value of the preset percentile can be determined to characterize the current network jitter level. For example, the preset percentile can be the 95th percentile, which is not limited in this application. The purpose of comparing the jitter value of the preset percentile with the preset jitter threshold is to smooth the jitter, eliminate the influence of abnormal fluctuations, and improve the stability of detecting the risk of jamming of audio playback. In one embodiment, the reduction of the deceleration ratio can be based on a fixed percentage or a decreasing percentage to reduce the current value of the deceleration ratio. In one embodiment, the reduction of the deceleration ratio can be based on a reference ratio value, and a preset mathematical model is used to determine the specific amplitude of each reduction, so as to achieve a gradual reduction in the deceleration ratio. The preset mathematical model can be an exponential decay model, a logarithmic decay model, etc., which is not limited in this application. Optionally, if the current cached number of audio data packets in the jitter buffer is less than the historical cached number recorded at the previous frame taking moment, or the jitter value of a preset percentile corresponding to multiple audio data packets received within a preset time range in the jitter buffer is greater than a preset jitter threshold, it can be determined that there is a risk of jamming in the current audio playback, and the deceleration ratio can be lowered. Optionally, after the audio frame is decelerated based on the adjusted deceleration ratio, the current cached number of audio data packets in the jitter buffer can be compared with the historical cached number recorded at the previous frame taking moment, and the jitter values ​​of a preset percentile corresponding to multiple audio data packets received within a preset time range in the jitter buffer can be compared with a preset jitter threshold, and the deceleration ratio can be adaptively adjusted based on the comparison results, and the audio frame can be decelerated based on the adjusted deceleration ratio until the current cached number of the jitter buffer reaches the target cached number, and normal playback of the audio frame is restored.

[0040] In the above, by receiving the audio data packets sent by the server, the audio data packets are stored in the jitter buffer, wherein the audio data packets include a first audio data packet pre-cached and sent by the server and a second audio data packet sent in real time; when the cumulative number of the first audio data packets is less than the target cache number set in the jitter buffer, a reference ratio value is determined according to the cumulative number and the target cache number, the deceleration ratio is set to the reference ratio value, and the audio frames corresponding to the audio data packets are extracted from the jitter buffer based on the deceleration ratio for decelerated playback; when the current cache number of audio data packets in the jitter buffer is greater than or equal to the historical cache number recorded at the previous frame acquisition time, and the jitter values ​​of preset percentiles corresponding to multiple audio data packets received within a preset time range in the jitter buffer are less than or equal to the preset jitter threshold, the deceleration ratio is lowered, and the audio frames are decelerated and played based on the adjusted deceleration ratio, so that the normal playback of the audio frames is restored when the current cache number reaches the target cache number. In the above scheme, when the cumulative number of the first audio data packets is less than the target cache number set in the jitter buffer, it can be regarded that the number of audio data packets accumulated in the jitter buffer is insufficient. The reference ratio value for the initial configuration of the deceleration ratio can be determined based on the cumulative number and the target cache number. The audio frames corresponding to the audio data packets are extracted from the jitter buffer based on the deceleration ratio for decelerated playback, so that the audio data packets in the jitter buffer can be quickly accumulated to the target cache number, effectively reducing the probability of playback jamming. Moreover, when the current cache number of audio data packets in the jitter buffer is greater than or equal to the historical cache number recorded at the previous frame acquisition moment, and the jitter values ​​of preset percentiles corresponding to multiple audio data packets received within a preset time range in the jitter buffer are less than or equal to the preset jitter threshold, it can be regarded that the probability of the risk of playback jamming is low, and the deceleration ratio can be lowered to reduce the audience's perceptible deceleration experience, effectively balance the playback jamming experience and the deceleration experience, and improve the user's overall listening experience.

[0041] Figure 2 A flowchart of a network audio playback method including a process of determining a deceleration ratio increase is provided in an embodiment of the present application. Figure 2 As shown, the following steps are included:

[0042] Step S201: receiving an audio data packet sent by a server, and storing the audio data packet in a jitter buffer, wherein the audio data packet includes a pre-cached first audio data packet sent by the server and a second audio data packet sent in real time.

[0043] Step S202: When the cumulative number of first audio data packets is less than the target cache number set in the jitter buffer, a reference ratio value is determined according to the cumulative number and the target cache number, the deceleration ratio is set to the reference ratio value, and audio frames corresponding to the audio data packets are extracted from the jitter buffer for decelerated playback based on the deceleration ratio.

[0044] Step S203: When the current cached number of audio data packets in the jitter buffer is greater than or equal to the historical cached number recorded at the previous frame taking moment, and the jitter values ​​of the preset percentiles corresponding to the multiple audio data packets received within the preset time range in the jitter buffer are less than or equal to the preset jitter threshold, the deceleration ratio is reduced.

[0045] Step S204: when the current buffered quantity of audio data packets in the jitter buffer is less than the historical buffered quantity recorded at the previous frame fetching time, the deceleration ratio is increased.

[0046] Among them, if the current cache number of audio data packets in the jitter buffer is less than the historical cache number recorded at the previous frame taking moment, it can be regarded that the consumption speed of the audio data packets in the jitter buffer is faster than the cache speed, and the current cache number is in a decreasing trend. It can be determined that there is a risk of jamming of audio playback at present, and the deceleration ratio can be increased to speed up the accumulation of audio data packets in the jitter buffer and alleviate the risk of jamming of audio playback. In one embodiment, the deceleration ratio can be increased by increasing the current value of the deceleration ratio based on a fixed percentage or an increasing percentage. In one embodiment, the deceleration ratio can be increased based on a reference ratio value, and a preset mathematical model is used to determine the specific amplitude of each increase, so as to achieve a gradual increase in the deceleration ratio. The preset mathematical model can be an exponential growth model, a logarithmic growth model, etc., which is not limited in this application.

[0047] Step S205: decelerate the playback of the audio frames based on the adjusted deceleration ratio, so as to restore the normal playback of the audio frames when the current buffer quantity reaches the target buffer quantity.

[0048] As mentioned above, when the current cache number of audio data packets in the jitter buffer is less than the historical cache number recorded at the previous frame taking moment, it is determined that there is a risk of audio playback jamming, and the deceleration ratio is increased. This can help the jitter buffer accumulate more audio data packets, reach the target cache number faster, and reduce the risk of audio playback jamming.

[0049] Figure 3 A flowchart of another network audio playback method including a process of determining a deceleration ratio increase is provided in an embodiment of the present application. Figure 3 As shown, the following steps are included:

[0050] Step S301: receiving an audio data packet sent by a server, and storing the audio data packet in a jitter buffer, wherein the audio data packet includes a pre-cached first audio data packet sent by the server and a second audio data packet sent in real time.

[0051] Step S302: When the cumulative number of the first audio data packets is less than the target cache number set in the jitter buffer, a reference ratio value is determined according to the cumulative number and the target cache number, the deceleration ratio is set to the reference ratio value, and the audio frames corresponding to the audio data packets are extracted from the jitter buffer for decelerated playback based on the deceleration ratio.

[0052] Step S303: When the current cached number of audio data packets in the jitter buffer is greater than or equal to the historical cached number recorded at the previous frame taking moment, and the jitter values ​​of the preset percentiles corresponding to the multiple audio data packets received within the preset time range in the jitter buffer are less than or equal to the preset jitter threshold, the deceleration ratio is reduced.

[0053] Step S304: when the current cached number of audio data packets in the jitter buffer is less than the historical cached number recorded at the previous frame taking time, or the jitter values ​​of the preset percentiles corresponding to multiple audio data packets received within a preset time range in the jitter buffer are greater than the preset jitter threshold, the deceleration ratio is increased.

[0054] Among them, if the jitter value of the preset percentile corresponding to multiple audio data packets received within the preset time range in the jitter buffer is greater than the preset jitter threshold, it can be regarded as a large degree of network jitter, and it can be determined that there is a risk of audio playback jamming. The deceleration ratio can be increased to speed up the accumulation of audio data packets in the jitter buffer and alleviate the risk of audio playback jamming.

[0055] Step S305: decelerate the playback of the audio frames based on the adjusted deceleration ratio, so as to restore the normal playback of the audio frames when the current buffer quantity reaches the target buffer quantity.

[0056] As mentioned above, when the current cached number of audio data packets in the jitter buffer is less than the historical cached number recorded at the previous frame taking moment, or the jitter value of a preset percentile corresponding to multiple audio data packets received within a preset time range in the jitter buffer is greater than a preset jitter threshold, it is determined that there is a risk of audio playback jamming, and the deceleration ratio is increased. This can speed up the accumulation of audio data packets in the jitter buffer, reach the target cache number faster, and reduce the risk of audio playback jamming.

[0057] Figure 4 A flowchart of a network audio playback method including a process of lowering the deceleration ratio provided in an embodiment of the present application. Figure 4 As shown, the following steps are included:

[0058] Step S401: receiving an audio data packet sent by a server, and storing the audio data packet in a jitter buffer, wherein the audio data packet includes a pre-cached first audio data packet sent by the server and a second audio data packet sent in real time.

[0059] Step S402: When the cumulative number of first audio data packets is less than the target cache number set in the jitter buffer, a reference ratio value is determined according to the cumulative number and the target cache number, the deceleration ratio is set to the reference ratio value, and the audio frames corresponding to the audio data packets are extracted from the jitter buffer for decelerated playback based on the deceleration ratio.

[0060] Step S403: When the current cached number of audio data packets in the jitter buffer is greater than or equal to the historical cached number recorded at the previous frame taking moment, and the jitter values ​​of the preset percentiles corresponding to the multiple audio data packets received within the preset time range in the jitter buffer are less than or equal to the preset jitter threshold, the set first growth power is increased, and an exponential reduction value is calculated based on the increased first growth power, the set exponential base and the amplification factor, a first candidate ratio value is obtained by subtracting the base ratio value from the exponential reduction value, and the deceleration ratio is set to the larger value between the first candidate ratio value and the set minimum ratio limit.

[0061] In this embodiment, an exponential decay model is used to reduce the deceleration ratio. The first growth power, the exponential base and the amplification factor can be parameters of the exponential decay model. For example, the exponential decay model is as follows: Among them, b 1 is the exponential base, t 1 is the first growth power, K 1 is the amplification factor. The exponential base and the first growth power can be used to control the decay speed. The amplification factor can be used to control the decay amplitude. The exponential base and the amplification factor are constants, and the first growth power can gradually increase with the increase of the number of reductions. For example, the initial value of the first growth power is set to 0, the exponential base is 3, and the amplification factor is 0.02. When the deceleration ratio is reduced for the first time, the first growth power can be increased by 1 to 1, and the corresponding exponential reduction value calculated is (3 1 -1)*0.02=0.04. When the deceleration ratio is reduced for the second time, the first growth power can be increased by 1 to 2. The corresponding calculated exponential reduction value is (3 2 -1)*0.02=0.16, and so on. Accordingly, the first candidate ratio value can be obtained by subtracting the base ratio value from the index reduction value, and the deceleration ratio is set to the larger value of the first candidate ratio value and the set minimum ratio limit. The specific calculation formula for calculating the current value of the deceleration ratio is as follows:

[0062]

[0063] Among them, SDR t is the current value of the deceleration ratio, SDR min is the minimum ratio limit, SDR 0 is the base ratio value, b 1 is the exponential base, t 1 is the first growth power, K 1 is the magnification factor.

[0064] For example, the base ratio value is set to 1.2, and the minimum ratio limit value is 1.01. The first calculated exponential reduction value is 0.04, and the first candidate ratio value is 1.2-0.04=1.16>1.01, and the deceleration ratio is set to 1.16. The second calculated exponential reduction value is 0.16, and the first candidate ratio value is 1.2-0.16=1.04>1.01, and the deceleration ratio is set to 1.04. By analogy, the first candidate ratio value obtained by subsequent calculations will be less than 1.01. Therefore, the deceleration ratio will be maintained at 1.01 for the third time and thereafter, maintaining the deceleration playback that the user basically does not perceive until the current cache quantity of the jitter buffer reaches the target cache quantity.

[0065] Step S404: decelerate the playback of the audio frames based on the adjusted deceleration ratio, so that the normal playback of the audio frames is restored when the current buffer quantity reaches the target buffer quantity.

[0066] In the above, by calculating the exponential adjustment value based on the increased first growth power, the set exponential base and the amplification factor, and determining the deceleration ratio based on the exponential adjustment value, the benchmark ratio value and the minimum ratio limit value, the exponential decay model can be effectively used to reduce the deceleration ratio, ensuring that the initial change of the adjustment is slow, which can effectively avoid over-response, and only increase significantly with the accumulation of the number of reductions, reducing the negative impact of fluctuations on the user's sound quality experience, and can effectively respond to changes in network status to ensure the real-time nature of the adjustment. In addition, compared with complex feedback control algorithms, the exponential growth has a smaller amount of calculation, higher flexibility, and is more suitable for voice room scenarios with higher real-time requirements.

[0067] Figure 5 A flowchart of a network audio playback method including a process of updating a first growth power and a reference ratio value provided by an embodiment of the present application. Figure 5 As shown, the following steps are included:

[0068] Step S501: receiving an audio data packet sent by a server, and storing the audio data packet in a jitter buffer, wherein the audio data packet includes a pre-cached first audio data packet sent by the server and a second audio data packet sent in real time.

[0069] Step S502: When the cumulative number of first audio data packets is less than the target cache number set in the jitter buffer, a reference ratio value is determined according to the cumulative number and the target cache number, the deceleration ratio is set to the reference ratio value, and audio frames corresponding to the audio data packets are extracted from the jitter buffer for decelerated playback based on the deceleration ratio.

[0070] Step S503: When the current cached number of audio data packets in the jitter buffer is greater than or equal to the historical cached number recorded at the previous frame taking moment, and the jitter values ​​of the preset percentiles corresponding to the multiple audio data packets received within the preset time range in the jitter buffer are less than or equal to the preset jitter threshold, the current freeze risk flag is set to a false value, the set first growth power is increased, and the exponential reduction value is calculated based on the increased first growth power, the set exponential base and the amplification factor, the base ratio value is subtracted from the exponential reduction value to obtain a first candidate ratio value, and the deceleration ratio is set to the larger value between the first candidate ratio value and the set minimum ratio limit.

[0071] Among them, if the current cached number of audio data packets in the jitter buffer is less than the historical cached number recorded at the previous frame taking moment, and the jitter value of the preset percentile corresponding to the multiple audio data packets received within the preset time range in the jitter buffer is greater than the preset jitter threshold, it can be regarded that there is a risk of jamming in the audio playback, and the current jamming risk flag can be set to a true value. If the current cached number of audio data packets in the jitter buffer is greater than or equal to the historical cached number recorded at the previous frame taking moment, and the jitter value of the preset percentile corresponding to the multiple audio data packets received within the preset time range in the jitter buffer is less than or equal to the preset jitter threshold, it can be regarded that there is no risk of jamming in the audio playback, and the current jamming risk flag can be set to a false value.

[0072] Step S504: When the current jam risk flag is inconsistent with the historical jam risk flag recorded when the deceleration ratio was last adjusted, the first growth power is reset to zero and the reference ratio value is updated to the current value of the deceleration ratio.

[0073] Among them, the historical jamming risk flag recorded when the deceleration ratio was adjusted last time may be a true value or a false value. Since the current jamming risk flag is a false value, if the historical jamming risk flag recorded when the deceleration ratio was adjusted last time is a true value, the two are inconsistent, indicating that the network condition has changed, and the adjustment direction of the deceleration ratio has changed from upward to downward, and will continue to be downward in the future. Therefore, the first growth power can be reset to zero, and the benchmark ratio value can be updated to the current value of the deceleration ratio, so that the subsequent calculation of the exponential downward adjustment value starts based on the smaller first growth power, and the first candidate ratio value is calculated based on the current value of the deceleration ratio corresponding to the moment of reference adjustment direction conversion, so that the deceleration ratio can be adjusted more accurately.

[0074] Step S505: decelerate the playback of the audio frames based on the adjusted deceleration ratio, so as to restore the normal playback of the audio frames when the current buffer quantity reaches the target buffer quantity.

[0075] As mentioned above, when the current jamming risk flag is inconsistent with the historical jamming risk flag recorded when the deceleration ratio was adjusted last time, the first growth power can be reset to zero, and the baseline ratio value can be updated to the current value of the deceleration ratio. During the deceleration ratio adjustment process, the baseline ratio value can be changed to adapt to network changes, which is more conducive to adjusting the deceleration ratio to balance the playback jamming experience and the deceleration experience.

[0076] Figure 6 A flowchart of a network audio playback method including a process of increasing the deceleration ratio provided in an embodiment of the present application. Figure 6 As shown, the following steps are included:

[0077] Step S601: receiving an audio data packet sent by a server, and storing the audio data packet in a jitter buffer, wherein the audio data packet includes a pre-cached first audio data packet sent by the server and a second audio data packet sent in real time.

[0078] Step S602: When the cumulative number of first audio data packets is less than the target cache number set in the jitter buffer, a reference ratio value is determined according to the cumulative number and the target cache number, the deceleration ratio is set to the reference ratio value, and the audio frames corresponding to the audio data packets are extracted from the jitter buffer for decelerated playback based on the deceleration ratio.

[0079] Step S603: When the current cached number of audio data packets in the jitter buffer is greater than or equal to the historical cached number recorded at the previous frame taking moment, and the jitter values ​​of the preset percentiles corresponding to the multiple audio data packets received within the preset time range in the jitter buffer are less than or equal to the preset jitter threshold, the deceleration ratio is reduced.

[0080] Step S604: When the current cache quantity of the audio data packets in the jitter buffer is less than the historical cache quantity recorded at the previous frame fetching moment, or when the jitter value of a preset quantile corresponding to multiple audio data packets received within a preset time range in the jitter buffer is greater than the preset jitter threshold, increase the set second growth power. Calculate an exponential increase value based on the increased second growth power, the set exponential base, and the amplification factor. Add the reference ratio value to the exponential increase value to obtain a second candidate ratio value, and set the deceleration ratio to the smaller value between the second candidate ratio value and the set maximum ratio limit.

[0081] Among them, in this embodiment, an exponential growth model is adopted to increase the deceleration ratio. The second growth power, the exponential base, and the amplification factor can be parameters of the exponential growth model. For example, the exponential growth model is as follows: Among them, b 2 is the exponential base, t 2 is the second growth power, K 2 is the amplification factor. The exponential base and the second growth power can be used to control the growth rate, and the amplification factor can be used to control the growth amplitude. The exponential base and the amplification factor are constants, while the second growth power can gradually increase as the number of increase times increases. For example, if the initial value of the second growth power is set to 0, the exponential base is 3, and the amplification factor is 0.02, then when the deceleration ratio is increased for the first time, the second growth power can be increased by 1 to 1, and the corresponding calculated exponential increase value is (3 1 -1)*0.02 = 0.04. When the deceleration ratio is increased for the second time, the second growth power can be increased by 1 to 2, and the corresponding calculated exponential increase value is (3 2 -1)*0.02 = 0.16, and so on. Correspondingly, adding the reference ratio value to the exponential increase value can obtain the second candidate ratio value, and setting the deceleration ratio to the smaller value between the second candidate ratio value and the set maximum ratio limit. The specific calculation formula for the current value of the deceleration ratio is as follows:

[0082]

[0083] Among them, SDR t is the current value of the deceleration ratio, SDR max is the maximum ratio limit, SDR 0 is the reference ratio value, b 2 is the exponential base, t 2 is the second growth power, K 2 is the amplification factor.

[0084] For example, the base ratio value is set to 1.01, and the maximum ratio limit is 1.2. The exponential increase value calculated for the second time is 0.04, and the second candidate ratio value is 1.01+0.04=1.05<1.2, and the deceleration ratio is set to 1.05. The exponential increase value calculated for the second time is 0.16, and the second candidate ratio value is 1.01+0.16=1.17<1.2, and the deceleration ratio is set to 1.17. By analogy, the second candidate ratio value calculated subsequently will be greater than 1.2. Therefore, the deceleration ratio will be maintained at 1.2 for the third time and thereafter to avoid the user's perception of a strong deceleration experience due to excessive deceleration, until the current cache quantity of the jitter buffer reaches the target cache quantity.

[0085] Step S605: decelerate the playback of the audio frames based on the adjusted deceleration ratio, so as to restore the normal playback of the audio frames when the current buffer quantity reaches the target buffer quantity.

[0086] In the above, by calculating the exponential increase value based on the increased second growth power, the set exponential base and the amplification factor, and determining the deceleration ratio based on the exponential increase value, the benchmark ratio value and the maximum ratio limit value, the exponential growth model can be effectively used to increase the deceleration ratio, and the deceleration ratio can be effectively adjusted in response to real-time changes in network status, which is suitable for voice room scenarios with high real-time requirements.

[0087] Figure 7 A flowchart of a network audio playback method including a process of updating a second growth power and a reference ratio value provided by an embodiment of the present application. Figure 7 As shown, the following steps are included:

[0088] Step S701: receiving an audio data packet sent by a server, and storing the audio data packet in a jitter buffer, wherein the audio data packet includes a pre-cached first audio data packet sent by the server and a second audio data packet sent in real time.

[0089] Step S702: When the cumulative number of first audio data packets is less than the target cache number set in the jitter buffer, a reference ratio value is determined according to the cumulative number and the target cache number, the deceleration ratio is set to the reference ratio value, and the audio frames corresponding to the audio data packets are extracted from the jitter buffer for decelerated playback based on the deceleration ratio.

[0090] Step S703: When the current cached number of audio data packets in the jitter buffer is greater than or equal to the historical cached number recorded at the previous frame taking moment, and the jitter values ​​of the preset percentiles corresponding to the multiple audio data packets received within the preset time range in the jitter buffer are less than or equal to the preset jitter threshold, the deceleration ratio is reduced.

[0091] Step S704: When the current cached number of audio data packets in the jitter buffer is less than the historical cached number recorded at the previous frame taking moment, or the jitter value of a preset percentile corresponding to multiple audio data packets received within a preset time range in the jitter buffer is greater than a preset jitter threshold, the current freeze risk flag is set to a true value, the set second growth power is increased, and an exponential increase value is calculated based on the increased second growth power, the set exponential base and the amplification factor. The base ratio value and the exponential increase value are added to obtain a second candidate ratio value, and the deceleration ratio is set to the smaller value of the second candidate ratio value and the set maximum ratio limit.

[0092] Among them, if the current cached number of audio data packets in the jitter buffer is less than the historical cached number recorded at the previous frame taking moment, and the jitter value of the preset percentile corresponding to multiple audio data packets received within the preset time range in the jitter buffer is greater than the preset jitter threshold, it can be regarded that there is a risk of stuttering in the audio playback, and the current stuttering risk flag can be set to a true value.

[0093] Step S705: When the current jam risk flag is inconsistent with the historical jam risk flag recorded when the deceleration ratio was last adjusted, the second growth power is reset to zero and the reference ratio value is updated to the current value of the deceleration ratio.

[0094] Among them, the historical jamming risk flag recorded when the deceleration ratio was adjusted last time may be a true value or a false value. Since the current jamming risk flag is a true value, if the historical jamming risk flag recorded when the deceleration ratio was adjusted last time was a false value, the two are inconsistent, indicating that the network condition has changed, and the adjustment direction of the deceleration ratio has changed from downward to upward, and will continue to be adjusted upward in the future. Therefore, the second growth power can be reset to zero, and the benchmark ratio value can be updated to the current value of the deceleration ratio, so that the subsequent calculation of the exponential increase value starts based on the smaller second growth power, and the second candidate ratio value is calculated based on the current value of the deceleration ratio corresponding to the moment of reference adjustment direction conversion, so that the deceleration ratio can be adjusted more accurately.

[0095] Step S706: decelerate the playback of the audio frames based on the adjusted deceleration ratio, so as to restore the normal playback of the audio frames when the current buffer quantity reaches the target buffer quantity.

[0096] As mentioned above, when the current stuttering risk flag is inconsistent with the historical stuttering risk flag recorded when the deceleration ratio was adjusted last time, the second growth power can be reset to zero, and the baseline ratio value can be updated to the current value of the deceleration ratio. During the deceleration ratio adjustment process, the baseline ratio value can be changed to adapt to network changes, which is more conducive to adjusting the deceleration ratio to balance the playback stuttering experience and the deceleration experience.

[0097] Figure 8This is a structural block diagram of a network audio playback device provided in an embodiment of the present application. The device is configured to execute the network audio playback method provided in the above embodiment and has the corresponding functional modules and beneficial effects of the execution method. Figure 8 As shown, the device comprises:

[0098] The data caching module 101 is configured to receive audio data packets sent by the server and store the audio data packets in a jitter buffer, wherein the audio data packets include a first audio data packet pre-cached and sent by the server and a second audio data packet sent in real time;

[0099] The audio deceleration playback module 102 is configured to determine a reference ratio value according to the cumulative number and the target buffer number when the cumulative number of the first audio data packet is less than the target buffer number set in the jitter buffer, set the deceleration ratio to the reference ratio value, and extract the audio frame corresponding to the audio data packet from the jitter buffer for deceleration playback based on the deceleration ratio;

[0100] The deceleration ratio adjustment module 103 is configured to reduce the deceleration ratio when the current cache number of audio data packets in the jitter buffer is greater than or equal to the historical cache number recorded at the previous frame fetching moment, and the jitter values ​​of preset percentiles corresponding to multiple audio data packets received within a preset time range in the jitter buffer are less than or equal to a preset jitter threshold, and to perform decelerated playback of audio frames based on the adjusted deceleration ratio, so as to restore normal playback of audio frames when the current cache number reaches the target cache number.

[0101] As described above, by receiving audio data packets sent by a server and storing the audio data packets in a jitter buffer, where the audio data packets include first audio data packets pre-cached and sent by the server and second audio data packets sent in real time; when the cumulative number of the first audio data packets is less than a target cache number set for the jitter buffer, a reference ratio value is determined according to the cumulative number and the target cache number, the deceleration ratio is set to the reference ratio value, and audio frames corresponding to the audio data packets are extracted from the jitter buffer based on the deceleration ratio for decelerated playback; when the current cache number of the audio data packets in the jitter buffer is greater than or equal to the historical cache number recorded at the previous frame extraction moment, and the jitter value of a preset quantile corresponding to multiple audio data packets received within a preset time range in the jitter buffer is less than or equal to a preset jitter threshold, the deceleration ratio is adjusted downward, and the audio frames are decelerated for playback based on the adjusted deceleration ratio, so as to resume normal playback of the audio frames when the current cache number reaches the target cache number. In the above solution, when the cumulative number of the first audio data packets is less than the target cache number set for the jitter buffer, it can be considered that the number of audio data packets accumulated in the jitter buffer is insufficient. A reference ratio value for initially configuring the deceleration ratio can be determined according to the cumulative number and the target cache number. Extracting audio frames corresponding to the audio data packets from the jitter buffer based on the deceleration ratio can enable the audio data packets in the jitter buffer to quickly accumulate to the target cache number, effectively reducing the probability of playback stuttering. Moreover, when the current cache number of the audio data packets in the jitter buffer is greater than or equal to the historical cache number recorded at the previous frame extraction moment, and the jitter value of a preset quantile corresponding to multiple audio data packets received within a preset time range in the jitter buffer is less than or equal to a preset jitter threshold, it can be considered that the probability of playback having a stuttering risk is relatively low, and the deceleration ratio can be adjusted downward to reduce the deceleration experience perceptible to the audience, effectively balancing the playback stuttering experience and the deceleration experience, and improving the overall listening experience of the user.

[0102] In a possible embodiment, it further includes a first deceleration ratio increase module configured to:

[0103] When the current cache number of the audio data packets in the jitter buffer is less than the historical cache number recorded at the previous frame extraction moment, increase the deceleration ratio.

[0104] In a possible embodiment, it further includes a second deceleration ratio increase module configured to:

[0105] When the jitter value of a preset quantile corresponding to multiple audio data packets received within a preset time range in the jitter buffer is greater than the preset jitter threshold, increase the deceleration ratio.

[0106] In a possible embodiment, the deceleration ratio adjustment module 103 is further configured to:

[0107] Increase the set first growth power, and calculate the exponential down-adjustment value based on the increased first growth power, the set exponential base and the amplification factor;

[0108] Subtracting the base ratio value from the index reduction value to obtain a first candidate ratio value;

[0109] The deceleration ratio is set to the larger value between the first candidate ratio value and the set minimum ratio limit value.

[0110] In a possible embodiment, it further includes a first flag configuration module configured to:

[0111] Set the current lag risk flag to a false value;

[0112] The first update module is configured as follows:

[0113] When the current jam risk flag is inconsistent with the historical jam risk flag recorded when the deceleration ratio was last adjusted, the first growth power is reset to zero and the base ratio value is updated to the current value of the deceleration ratio.

[0114] In a possible embodiment, the second deceleration ratio increasing module is further configured as follows:

[0115] Increasing the set second growth power, and calculating the exponential increase value based on the increased second growth power, the set exponential base and the amplification factor;

[0116] Adding the base ratio value to the index increase value to obtain a second candidate ratio value;

[0117] The deceleration ratio is set to a smaller value between the second candidate ratio value and the set maximum ratio limit value.

[0118] In a possible embodiment, it further includes a first flag configuration module configured to:

[0119] Set the current lag risk flag to true;

[0120] The second update module is configured as follows:

[0121] When the current jam risk flag is inconsistent with the historical jam risk flag recorded when the deceleration ratio was last adjusted, the second growth power is reset to zero and the base ratio value is updated to the current value of the deceleration ratio.

[0122] In a possible embodiment, the audio slow-down playback module 102 is further configured to:

[0123] Compare the product of the accumulated quantity and the target cache quantity with the preset coefficient;

[0124] When the accumulated amount is less than the product of the target cache amount and the preset coefficient, the base ratio value is determined as the set maximum ratio limit value;

[0125] When the accumulated number is greater than or equal to the product of the target cache number and the preset coefficient, the reference ratio value is determined as the set minimum ratio limit value.

[0126] Figure 9 A schematic diagram of the structure of a network audio playback device provided in an embodiment of the present application is shown in FIG. Figure 9 As shown, the device includes a processor 201, a memory 202, an input device 203 and an output device 204; the number of processors 201 in the device can be one or more. Figure 9 A processor 201 is taken as an example; the processor 201, memory 202, input device 203 and output device 204 in the device can be connected by a bus or other means. Figure 9 The example of connecting through a bus is taken. The memory 202, as a computer-readable storage medium, can be configured to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the network audio playback method in the embodiment of the present application. The processor 201 executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory 202, that is, implements the above-mentioned network audio playback method. The input device 203 can be configured to receive input digital or character information, and generate key signal input related to user settings and function control of the device. The output device 204 may include display devices such as display screens.

[0127] The embodiment of the present application also provides a non-volatile storage medium containing computer executable instructions, which are configured to execute a network audio playback method described in the above embodiment when executed by a computer processor, including: receiving an audio data packet sent by a server, and storing the audio data packet in a jitter buffer, wherein the audio data packet includes a first audio data packet pre-cached and sent by the server and a second audio data packet sent in real time; when the cumulative number of the first audio data packet is less than the target cache number set in the jitter buffer, determining a reference ratio value according to the cumulative number and the target cache number, setting the deceleration ratio to the reference ratio value, and extracting an audio frame corresponding to the audio data packet from the jitter buffer based on the deceleration ratio for decelerated playback; when the current cache number of audio data packets in the jitter buffer is greater than or equal to the historical cache number recorded at the previous frame fetching moment, and the jitter values ​​of the preset percentiles corresponding to the multiple audio data packets received within a preset time range in the jitter buffer are less than or equal to the preset jitter threshold, lowering the deceleration ratio, and decelerating the audio frame based on the adjusted deceleration ratio, so that the normal playback of the audio frame is restored when the current cache number reaches the target cache number.

[0128] It is worth noting that in the embodiment of the above-mentioned network audio playback device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not configured to limit the protection scope of the embodiments of the present application.

[0129] In some possible implementations, various aspects of the method provided in this application may also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is configured to enable the computer device to execute the steps of the method according to various exemplary embodiments of the present application described above in this specification. For example, the computer device may execute the network audio playback method described in the embodiment of this application. The program product may be implemented in any combination of one or more readable media.

Claims

1. A network audio playback method, characterized in that: include: Receive an audio data packet sent by a server, and store the audio data packet in a jitter buffer, wherein the audio data packet includes a first audio data packet sent by the server and cached in advance and a second audio data packet sent in real time; In a case where the accumulated number of the first audio data packets is less than the target cache number set in the jitter buffer, determining a reference ratio value according to the accumulated number and the target cache number, setting a deceleration ratio to the reference ratio value, and extracting audio frames corresponding to the audio data packets from the jitter buffer for decelerated playback based on the deceleration ratio; When the current cache quantity of audio data packets in the jitter buffer is greater than or equal to the historical cache quantity recorded at the previous frame taking moment, and the jitter values ​​of preset percentiles corresponding to multiple audio data packets received within a preset time range in the jitter buffer are less than or equal to a preset jitter threshold, the deceleration ratio is lowered, and the audio frames are played at a reduced speed based on the adjusted deceleration ratio, so that the normal playback of the audio frames is restored when the current cache quantity reaches the target cache quantity.

2. The network audio playback method according to claim 1, characterized in that: Before performing the slow-down playback of the audio frame based on the adjusted slow-down ratio, the method further includes: When the current buffered quantity of audio data packets in the jitter buffer is less than the historical buffered quantity recorded at the previous frame fetching moment, the deceleration ratio is increased.

3. The network audio playback method according to claim 1, characterized in that: Before performing the slow-down playback of the audio frame based on the adjusted slow-down ratio, the method further includes: When the jitter values ​​of preset percentiles corresponding to a plurality of audio data packets received within a preset time range in the jitter buffer are greater than a preset jitter threshold, the deceleration ratio is increased.

4. The network audio playback method according to claim 1, characterized in that: The step of lowering the deceleration ratio includes: Increasing the set first growth power, and calculating the exponential down-adjustment value based on the increased first growth power, the set exponential base and the amplification factor; Subtracting the base ratio value from the index reduction value to obtain a first candidate ratio value; The deceleration ratio is set to a larger value between the first candidate ratio value and a set minimum ratio limit value.

5. The network audio playback method according to claim 4, characterized in that: Before the first increasing power of the increasing setting, the method further comprises: Set the current lag risk flag to a false value; After setting the deceleration ratio to a larger value between the first candidate ratio value and the set minimum ratio limit value, the method further includes: When the current jam risk flag is inconsistent with the historical jam risk flag recorded when the deceleration ratio was last adjusted, the first growth power is reset to zero and the reference ratio value is updated to the current value of the deceleration ratio.

6. The network audio playing method according to any one of claims 2-3, characterized in that: The step of increasing the deceleration ratio includes: Increasing the set second growth power, and calculating the exponential increase value based on the increased second growth power, the set exponential base and the amplification factor; Adding the base ratio value to the index increase value to obtain a second candidate ratio value; The deceleration ratio is set to a smaller value between the second candidate ratio value and a set maximum ratio limit value.

7. The network audio playback method according to claim 6, characterized in that: Before the step of increasing the setting to a second increasing power, the method further comprises: Set the current lag risk flag to true; After setting the deceleration ratio to a smaller value between the second candidate ratio value and the set maximum ratio limit value, the method further includes: When the current jam risk flag is inconsistent with the historical jam risk flag recorded when the deceleration ratio was adjusted last time, the second growth power is reset to zero and the reference ratio value is updated to the current value of the deceleration ratio.

8. The network audio playback method according to claim 1, characterized in that: The determining of the reference ratio value according to the accumulated quantity and the target cache quantity includes: Compare the cumulative number and the product of the target cache number and a preset coefficient; When the accumulated number is less than the product of the target cache number and a preset coefficient, the reference ratio value is determined as the set maximum ratio limit value; When the accumulated number is greater than or equal to the product of the target cache number and a preset coefficient, the reference ratio value is determined as a set minimum ratio limit value.

9. A network audio playback device, characterized in that: include: A data caching module, configured to receive an audio data packet sent by a server and store the audio data packet in a jitter buffer, wherein the audio data packet includes a pre-cached first audio data packet sent by the server and a second audio data packet sent in real time; an audio deceleration playback module, configured to, when the cumulative number of the first audio data packets is less than the target cache number set in the jitter buffer, determine a reference ratio value according to the cumulative number and the target cache number, set the deceleration ratio to the reference ratio value, and extract the audio frame corresponding to the audio data packet from the jitter buffer for deceleration playback based on the deceleration ratio; The deceleration ratio adjustment module is configured to reduce the deceleration ratio when the current cache quantity of audio data packets in the jitter buffer is greater than or equal to the historical cache quantity recorded at the previous frame taking moment, and the jitter values ​​of preset percentiles corresponding to multiple audio data packets received within a preset time range in the jitter buffer are less than or equal to a preset jitter threshold, and to slow down the playback of audio frames based on the adjusted deceleration ratio, so as to restore normal playback of audio frames when the current cache quantity reaches the target cache quantity.

10. A network audio playback device, comprising: one or more processors; A storage device is configured to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the network audio playback method described in any one of claims 1 to 8.

11. A non-volatile storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the network audio playback method according to any one of claims 1 to 8 when executed by a computer processor.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the network audio playback method according to any one of claims 1 to 8 is implemented.