Audio and video decoding output method and device, equipment and storage medium

By receiving audio and video data frames in real time and determining the expected rendering time based on network jitter parameters and frame cache queue length, the audio and video playback problem caused by delay compensation calculation exceptions is solved, and more accurate rendering time calculation and higher fault tolerance are achieved, thereby avoiding increased delay or lag during audio and video playback.

CN120050466APending Publication Date: 2025-05-27GUANGZHOU MAILING INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311601648.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, calculation abnormalities in delay compensation may lead to abnormal audio and video decoding, resulting in increased delay during audio and video playback or screen stuttering.

Method used

By receiving audio and video data frames in real time and obtaining network jitter parameters, confirming the cache length compensation time based on the frame rate and the cache length in the frame cache queue, and then determining the expected rendering time of the audio and video data frames, ensuring that the data frames are sent to the decoder for decoding and output at the desired rendering time.

Benefits of technology

It improves the accuracy of the expected rendering time calculation, increases the error tolerance rate of abnormal calculation during delay compensation, and effectively controls the increase in delay or the sound and picture stuttering during audio and video playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050466A_ABST
    Figure CN120050466A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an audio and video decoding output method and device, equipment and a storage medium, and the audio and video decoding output method comprises the steps: receiving audio and video data frames in real time, caching the audio and video data frames to a frame buffer queue, and obtaining a network jitter parameter in the process of receiving the audio and video data frames in real time; and according to the network jitter parameter, the frame rate of the audio and video data frame and the cache length in the frame cache queue, determining the cache length compensation duration corresponding to the currently received audio and video data frame. And according to the cache length compensation duration and a preset reference compensation duration, determining the expected rendering time of the currently received audio and video data frame. And sending the corresponding audio and video data frames to a decoder for decoding and outputting at the expected rendering time. The expected rendering time is predicted according to the cache length of the actual audio and video frame to be output, the calculation accuracy of the expected rendering time is improved, the error-tolerant rate when calculation abnormity occurs in the compensation process is increased, and audio and video playing time delay increase or audio and video lagging is effectively controlled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of audio - video technology, and in particular, to an audio - video decoding and output method, apparatus, device, and storage medium. Background Art

[0002] With the continuous development of electronic technology, various forms of electronic devices have emerged and are applied to various fields of social production activities. For example, in the fields of entertainment, education, social networking, and office work, users can receive audio - video data in real - time through applicable electronic devices, directly play the audio - video without saving media files, and obtain information.

[0003] To ensure the normal playback of audio - video data in various network states, a buffer is usually set to buffer audio - video data frames, especially to ensure the smoothness of the video in a weak network situation. Specifically, Kalman filtering is used to estimate network jitter and the expected reception time, and then compensation is made by comprehensively considering the possible delays in the audio - video data transmission and decoding processes to obtain the expected rendering time for playing the audio - video data frame. After the expected rendering time arrives, the audio - video data frame is passed to the decoder for decoding.

[0004] When the inventor studied the existing decoding at the expected rendering time, it was found that the possible delay situations are relatively complex, and calculation anomalies in various delay compensations may lead to abnormal final decoding, increased delay during audio - video playback, or frame freezing. Summary of the Invention

[0005] The present invention provides an audio - video decoding and output method, apparatus, device, and storage medium to solve the technical problem that calculation anomalies in existing delay compensation may lead to abnormal final decoding, increased delay during audio - video playback, or frame freezing.

[0006] In a first aspect, embodiments of the present invention provide an audio - video decoding and output method, which includes:

[0007] Receiving audio - video data frames in real - time and caching them into a frame buffer queue, and obtaining network jitter parameters during the process of receiving audio - video data frames in real - time;

[0008] According to the network jitter parameters, the frame rate of the audio - video data frame, and the cache length in the frame buffer queue, determining the cache length compensation duration corresponding to the currently received audio - video data frame;

[0009] According to the cache length compensation duration and a preset reference compensation duration, determining the expected rendering time of the currently received audio - video data frame;

[0010] Sending the corresponding audio - video data frame to the decoder for decoding and output at the expected rendering time.

[0011] Second aspect, an embodiment of the present invention provides an audio - video decoding and output device; the audio - video decoding and output device includes:

[0012] A data receiving unit, configured to receive audio - video data frames in real - time, buffer them into a frame buffer queue, and obtain network jitter parameters during the process of receiving audio - video data frames in real - time;

[0013] A length compensation unit, configured to confirm the cache length compensation duration corresponding to the currently received audio - video data frame according to the network jitter parameters, the frame rate of the audio - video data frame, and the cache length in the frame buffer queue;

[0014] A rendering time confirmation unit, configured to confirm the expected rendering time of the currently received audio - video data frame according to the cache length compensation duration and a preset reference compensation duration;

[0015] A decoding and sending unit, configured to send the corresponding audio - video data frame to a decoder for decoding and output at the expected rendering time.

[0016] Third aspect, an embodiment of the present invention provides an electronic device, the electronic device includes:

[0017] One or more processors;

[0018] A memory, configured to store one or more computer programs;

[0019] When the one or more computer programs are executed by the one or more processors, the electronic device implements the audio - video decoding and output method as in the first aspect.

[0020] Fourth aspect, an embodiment of the present invention provides a computer - readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the audio - video decoding and output method as in the first aspect is implemented.

[0021] The above audio - video decoding and output method, device, equipment, and storage medium. The audio - video decoding and output method includes: receiving audio - video data frames in real - time and caching them into a frame buffer queue, and obtaining network jitter parameters during the process of receiving audio - video data frames in real - time; according to the network jitter parameters, the frame rate of the audio - video data frames, and the cache length in the frame buffer queue, determining the cache - length compensation duration corresponding to the currently received audio - video data frame; according to the cache - length compensation duration and a preset reference compensation duration, determining the expected rendering time of the currently received audio - video data frame; and sending the corresponding audio - video data frame to the decoder for decoding and output at the expected rendering time. By utilizing the relationship between the current cache length in the frame buffer queue and the appropriate cache - length range, the delay compensation is corrected, and the expected rendering time of the latest received audio - video data frame is predicted based on the cache length of the actually to - be - output audio - video data frames, improving the accuracy of the calculation of the expected rendering time, increasing the fault - tolerance rate when calculation anomalies occur during various delay compensation processes, and effectively controlling the increase in delay or audio - video stuttering during audio - video playback. Brief Description of the Drawings

[0022] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following - described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 It is a flowchart of a method for an audio - video decoding and output method provided by an embodiment of the present application;

[0024] Figure 2 It is a schematic diagram of the storage state of the cache queue in an audio - video decoding and output method provided by an embodiment of the present application;

[0025] Figure 3 It is a schematic structural diagram of an audio - video decoding and output device provided by an embodiment of the present application;

[0026] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0027] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings. It can be understood that the specific embodiments described herein are used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings, rather than all structures.

[0028] It should be noted that due to space limitations, the description of this application does not enumerate all optional implementation manners. After reading the description of this application, those skilled in the art should be able to think that as long as the technical features do not conflict with each other, any combination of technical features can constitute an optional implementation manner.

[0029] The following will describe each embodiment in detail.

[0030] During the real-time playback of audio and video data, in order to ensure the normal playback of audio and video data in various network states, a buffer is usually set to buffer audio and video data frames, especially to ensure the smoothness of the video in the case of a weak network. Specifically, the Kalman filter is used to estimate the network jitter and the expected reception time, and then the RTT (Round-Trip Time) feedback by RTCP (RTP Control Protocol) is used as compensation. Adding the decoding delay and the rendering delay to obtain the expected rendering time, the expected rendering time for playing the audio and video data frame is obtained, and the audio and video data frame is passed to the decoder for decoding after the rendering time arrives.

[0031] However, in the actual process of audio and video data transmission, there may be many abnormal situations. For example, RTCP packets may be lost when the network is poor, resulting in a lag in the feedback RTT; the feedback RTT is large and small due to network fluctuations, and the calculated compensation value is not optimal, or an abnormality occurs in a certain link calculation, resulting in an increase in video delay or frame freezing of the picture.

[0032] To solve the above technical problems, the embodiments of this application propose an audio and video decoding output method. By using the relationship between the current buffer length in the frame buffer queue and the appropriate buffer length range, the delay compensation is corrected, and the buffer length of the actually to-be-output audio and video data frame is used to predict the expected rendering time of the newly received audio and video data frame, which improves the accuracy of the expected rendering time calculation, increases the fault tolerance rate when calculation abnormalities occur in various delay compensation processes, and effectively controls the increase in delay or frame freezing of audio and video during playback.

[0033] Figure 1 is a method flowchart of an audio and video decoding output method provided by the embodiments of this application. As Figure 1 shown, the audio and video decoding output method includes but is not limited to steps S110 - step S140.

[0034] Step S110: Receive audio and video data frames in real time and buffer them into the frame buffer queue, and obtain the network jitter parameters during the process of receiving audio and video data frames in real time.

[0035] The audio-visual data in the embodiments of this application refers to the audio-visual data that is received and played in real time. It can be the audio-visual data pre-saved in the server and received from the server, or the audio-visual data collected and sent by the other electronic device during real-time communication and correspondingly received by the local electronic device. The types of audio-visual data include separate audio data, separate video data, and integrated audio-visual data. In the specific process of processing real-time audio-visual data, the receiving end for outputting and playing audio-visual data usually does not directly process each received frame immediately, but instead goes through a frame buffer queue for transition in the middle. Based on the frame buffer queue, the received audio-visual data frames are cached according to the network transmission status, and the output and played audio-visual data frames are read and output frame by frame from the frame buffer queue according to established processing rules, so as to avoid interference between reception and playback as much as possible. Network jitter refers to the delay variation that occurs to data packets during network communication. The parameter for quantitatively characterizing the state of network jitter is the network jitter parameter.

[0036] During the specific real-time playback of audio-visual data, the process of receiving audio-visual data is processed in units of frames, and obtaining the network jitter parameter during the network data transmission process can refer to the existing implementations of related technologies, and will not be repeated here. It should be understood that "real-time" in real-time reception and real-time playback refers to the synchronization of reception and playback as a whole, and does not mean that each frame is received and played without any time difference interval.

[0037] Step S120: Confirm the cache length compensation duration corresponding to the currently received audio-visual data frame according to the network jitter parameter, the frame rate of the audio-visual data frame, and the cache length in the frame buffer queue.

[0038] In the embodiments of this application, considering that the number of audio-visual data frames cached in the frame buffer queue is actually also indirectly affected by the expected rendering time. For example, at a certain moment, if the feedback RTT suddenly increases, then the expected rendering time also increases accordingly. Before the expected rendering time is reached, the audio-visual data frames will be cached in the queue. Since the audio-visual data frames are not consumed in time (i.e., output and played), the subsequent audio-visual data frames continuously enter the queue, resulting in the accumulation of audio-visual data frames and an increase in the output and playback delay of the audio-visual data. At this time, the compensation calculated based on the feedback RTT is not optimal. Based on this, in the embodiments of this application, the cache length factor is further introduced during delay compensation. Compared with the prior art that only compensates based on various state information before receiving the audio-visual data frames, adding the cache length in the frame buffer queue when receiving the audio-visual data frames as a compensation reference can adjust the expected rendering time when the state of audio-visual data frame accumulation fluctuates, such as when there are too many or too few accumulated audio-visual data frames, reduce the delay and reduce stuttering at the same time, and obtain an approximately optimal expected rendering time. The frame rate specifically refers to the theoretical frame rate of the audio-visual data frame.

[0039] In the specific processing process, there are three ways to actually confirm the cache length compensation duration according to different cache states.

[0040] The first is that when the cache length is greater than the maximum queue length corresponding to the frame rate, regardless of the network jitter state, the cache length compensation duration is reduced. When the cache length is greater than the maximum queue length corresponding to the frame rate, it is equivalent to the accumulation of audio and video data frames in the cache, and it is necessary to increase the processing speed of the data in the frame cache queue. That is, without affecting the user experience of the output playback of audio and video data frames, it is necessary to reduce the waiting duration of audio and video data frames in the frame cache queue. When the benchmark compensation duration cannot adjust the established pre-set result, it is adjusted by reducing the cache length compensation duration. In the specific processing process, the judgment conditions for whether the cache length is too long, normal or too short are not fixed, but are determined according to the corresponding frame rate. In comparison, the higher the frame rate, the more audio and video data frames usually need to be cached, and the larger the threshold value for judging the existence of accumulation, that is, the larger the maximum queue length. That is to say, the maximum queue length does not refer to the upper limit set for the cached number of frames, but refers to the upper limit for judging the normal cache state.

[0041] Among them, the maximum queue length is obtained by multiplying the frame rate, and can be expressed by the equation: maximum queue length = (target delay + additional buffer coefficient) * frame rate, where the target delay and the additional buffer coefficient are two constants, which can be determined according to experience, and are specifically used to ensure that the value of the maximum queue length is reasonable, and to avoid quickly reaching the maximum queue length or being difficult to reach the maximum queue length. When the cache length is greater than the maximum queue length corresponding to the frame rate, the reduced cache length compensation duration is confirmed according to the maximum queue length, cache length and frame rate. The specific confirmation method can be expressed by the equation: reduced cache length compensation duration = compensation coefficient * (maximum queue length - cache length) / frame rate, where the compensation coefficient is a constant confirmed based on experience or statistics.

[0042] The second is to increase the buffer length compensation time when the network jitter parameter exceeds the preset first threshold and the buffer length is less than the minimum queue length corresponding to the frame rate. That is, when the network jitter is more severe and the buffer length is less than the minimum queue length corresponding to the frame rate, it is equivalent to that there is insufficient margin for the buffered audio and video data frames, and it is necessary to reduce the processing speed of the data in the frame buffer queue, that is, to ensure that the output playback of the audio and video data frames does not affect the user experience, it is necessary to increase the waiting time of the audio and video data frames in the frame buffer queue. When the benchmark compensation time cannot be preset in advance for adjustment, it is adjusted by increasing the buffer length compensation time. In the specific processing process, the judgment condition of whether the buffer length is too long, normal or too short is not fixed, but is determined according to the frame rate. In contrast, the higher the frame rate, the more audio and video data frames need to be cached, and the corresponding threshold value for judging that the buffered audio and video data frames are insufficient is larger, that is, the larger the minimum queue length. That is to say, the minimum queue length does not refer to the lower limit set for the number of cached frames, but refers to the lower limit for judging that it is in a normal cache state.

[0043] The minimum queue length is obtained by multiplying the product of the frame rate and the network jitter parameter, and can be expressed by the equation: minimum queue length = frame rate * target delay * coefficient * network jitter parameter, where the target delay and coefficient are constants confirmed based on experience or statistics, and the network jitter parameter is specifically characterized by a ratio. When the network jitter parameter exceeds the preset first threshold and the buffer length is less than the minimum queue length corresponding to the frame rate, the increased buffer length compensation duration is confirmed based on the minimum queue length, buffer length and frame rate, and the specific confirmation method can be expressed by the equation: increased queue length compensation = compensation coefficient * (minimum queue length – buffer length) / frame rate, where the compensation coefficient is a constant confirmed based on experience or statistics.

[0044] The third is that when the network jitter parameter is less than the preset first threshold value, and / or the cache length is in the closed interval formed by the minimum queue length and the maximum queue length, the preset cache length compensation duration is maintained, that is, the network jitter is relatively stable, and the cached audio and video data frames are not too high or too low. Maintain the preset cache length compensation duration, that is, maintain the normal processing speed. The cache length compensation duration can be 0, or a duration greater than 0. In addition, the first threshold can be the jitter in a normal network obtained based on statistics or experience, or multiply the jitter in a normal network by a positive number that is not 1.

[0045] The corresponding states of the above three processing methods in the frame buffer queue can be referred to Figure 2 .exist Figure 2In this case, it is assumed that the audio-visual data frames are stored starting from the bottom up, and the network jitter is always in a state where it may trigger an adjustment of the cache length compensation duration. As a result, the audio-visual data frames in the frame buffer queue will continuously change. Figure 2 The minimum queue length corresponding to the frame rate of the audio-visual data frames cached in it is A, and the maximum queue length is B. During the continuous change of the audio-visual data frames in the frame buffer queue, the audio-visual data frames are first cached in the P1 section. At this time, the cache length is always less than the minimum queue length, which can be regarded as having relatively few cached audio-visual data frames available for output playback. Correspondingly, it is necessary to increase the cache length compensation duration to avoid all the previously received audio-visual data frames being output and played, and the lag being concentrated before the latest received audio-visual data frame. This is equivalent to spreading the lag over multiple previous audio-visual data frames to eliminate the possible long-term lag that may occur later, and waiting for the arrival of subsequent audio-visual data frames during the output playback process to continue the audio-visual data frames required for output playback. When the cache length reaches the minimum queue length and does not exceed the maximum queue length, that is, the cache length is in the P2 section, the cache length of the audio-visual data frames can be regarded as being in a relatively balanced state of reception and output, and the current output playback rhythm can be maintained. When the cache length exceeds the maximum queue length, that is, the cache length is in the P3 section, it means that the cache length of the audio-visual data frames has started to accumulate. At this time, the cache length compensation duration can be reduced to increase the output playback speed of the audio-visual data frames to avoid continuous accumulation causing lag. Combining Figure 2 And as can be seen from the above description, based on the embodiments of the present application, the accuracy of the expected rendering time calculation can be improved according to the flexible response to the actual processing progress of the cache state of the frame buffer queue, the error tolerance rate in the calculation anomalies that occur during various delay compensation processes can be increased, and the increase in delay or audio-visual lag during audio-visual playback can be effectively controlled.

[0046] It should be understood that for the definition of the three sections, there can be different representation methods in the specific implementation process. For example, above the maximum queue length and below the minimum queue length; or for another example, considering that the audio-visual frame data is usually processed in whole frames as a unit, two decimals can be set as the maximum queue length and the minimum queue length respectively. Correspondingly, the situation where the actual cache length is exactly the threshold value may not occur directly, and the section judgment can be based on being greater than or less than the maximum queue length and / or the minimum queue length.

[0047] Step S130: Confirm the expected rendering time of the currently received audio-visual data frame according to the cache length compensation duration and the preset reference compensation duration.

[0048] The reference compensation duration is the sum of the encoding delay compensation, the rendering delay compensation, the network jitter compensation, and the round-trip delay compensation. Specifically, the expected rendering time is generated with reference to the expected reception time, plus all the compensation durations. That is, based on the expected reception time, delaying all the compensation durations backward gives the expected rendering time. After confirming the cache length compensation duration, the calculation process of the specific expected rendering time can refer to the calculation process in related technologies.

[0049] Step S140: Send the corresponding audio-visual data frame to the decoder for decoding and output at the expected rendering time.

[0050] For the audio-visual data frames saved frame by frame in the frame buffer queue, the expected rendering time is generated correspondingly as in steps S110 - S130 when received. When finally decoded and output, they are sent to the decoder for decoding at the corresponding expected rendering time according to the existing decoding and output process of the audio-visual.

[0051] The above audio-visual decoding and output method includes: receiving audio-visual data frames in real time and caching them into the frame buffer queue, and obtaining the network jitter parameters during the process of receiving audio-visual data frames in real time; confirming the cache length compensation duration corresponding to the currently received audio-visual data frame according to the network jitter parameters, the frame rate of the audio-visual data frame, and the cache length in the frame buffer queue; confirming the expected rendering time of the currently received audio-visual data frame according to the cache length compensation duration and the preset reference compensation duration; sending the corresponding audio-visual data frame to the decoder for decoding and output at the expected rendering time. By utilizing the relationship between the current cache length in the frame buffer queue and the appropriate cache length range, the delay compensation is corrected, and the expected time of the latest received audio-visual data frame is predicted based on the cache length of the actually to-be-output audio-visual data frame, which improves the accuracy of the expected rendering time calculation, increases the fault tolerance rate when calculation anomalies occur in various delay compensation processes, and effectively controls the increase in delay or audio-visual stuttering during audio-visual playback.

[0052] Figure 3 This is a schematic structural diagram of an audio-visual decoding and output device provided by an embodiment of the present application. As Figure 3 shown, the audio-visual decoding and output device includes a data reception unit 210, a length compensation unit 220, a rendering time confirmation unit 230, and a decoding and sending unit 240.

[0053] Among them, the data receiving unit 210 is used to receive audio and video data frames in real time, cache them into the frame buffer queue, and obtain the network jitter parameters during the process of receiving audio and video data frames in real time; the length compensation unit 220 is used to confirm the cache length compensation duration corresponding to the currently received audio and video data frame according to the network jitter parameters, the frame rate of the audio and video data frame, and the cache length in the frame buffer queue; the rendering time confirmation unit 230 is used to confirm the expected rendering time of the currently received audio and video data frame according to the cache length compensation duration and the preset reference compensation duration; the decoding and sending unit 240 is used to send the corresponding audio and video data frame to the decoder for decoding and output at the expected rendering time.

[0054] Based on the above embodiment, the length compensation unit 220 includes:

[0055] The first compensation module is used to reduce the cache length compensation duration when the cache length is greater than the maximum queue length corresponding to the frame rate;

[0056] The second compensation module is used to increase the cache length compensation duration when the network jitter parameter exceeds the preset first threshold and the cache length is less than the minimum queue length corresponding to the frame rate;

[0057] The third compensation module is used to maintain the preset cache length compensation duration when the network jitter parameter is less than the preset first threshold and / or the cache length is within the closed interval formed by the minimum queue length and the maximum queue length.

[0058] Based on the above embodiment, the maximum queue length is obtained by multiplying the frame rate by a multiple.

[0059] Based on the above embodiment, the minimum queue length is obtained by multiplying the product of the frame rate and the network jitter parameter by a multiple.

[0060] Based on the above embodiment, the first compensation module includes:

[0061] The reduction amplitude confirmation sub-module is used to confirm the reduced cache length compensation duration according to the maximum queue length, the cache length, and the frame rate when the cache length is greater than the maximum queue length corresponding to the frame rate.

[0062] Based on the above embodiment, the second compensation module includes:

[0063] The increase amplitude confirmation sub-module is used to confirm the increased cache length compensation duration according to the minimum queue length, the cache length, and the frame rate when the network jitter parameter exceeds the preset first threshold and the cache length is less than the minimum queue length corresponding to the frame rate.

[0064] Based on the above embodiments, the reference compensation duration is the sum of the encoding delay compensation, the rendering delay compensation, the network jitter compensation, and the round-trip delay compensation.

[0065] The audio-video decoding output device provided by the embodiments of the present application is included in an electronic device and can be used to execute the corresponding audio-video decoding output method provided in the above embodiments, and has the corresponding functions and beneficial effects.

[0066] It should be noted that in the embodiments of the above audio-video decoding output device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.

[0067] Figure 4 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 4 shown, the electronic device includes a processor 310 and a memory 320. In a common product form, it may further include an input device 330, an output device 340, and a communication device 350; the number of processors 310 in the electronic device may be one or more, Figure 4 taking one processor 310 as an example; the processor 310, the memory 320, the input device 330, the output device 340, and the communication device 350 in the electronic device may be connected by a bus or other means, Figure 4 taking the connection by a bus as an example.

[0068] The memory 320, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the audio-video decoding output method in the embodiments of the present application. The processor 310 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 320, that is, realizes the above audio-video decoding output method.

[0069] The memory 320 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device. In addition, the memory 320 may include high-speed random access memory, and may further include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 320 may further include a memory remotely set relative to the processor 310, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.

[0070] The input device 330 can be used to receive network configuration information. The output device 340 may include a display device such as a display screen.

[0071] The above electronic device can be used to execute any audio-video decoding and output method, and has corresponding functions and beneficial effects.

[0072] An embodiment of the present invention also provides a storage medium containing computer-executable instructions. The computer-executable instructions are used to perform related operations in the audio-video decoding and output method provided in any embodiment of the present application when executed by a computer processor, and have corresponding functions and beneficial effects.

[0073] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product.

[0074] Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1A step that specifies a function in one or more boxes.

[0075] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory. The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0076] Computer readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0077] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0078] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. Audio and video decoding output method, characterized in that, comprising: Receiving audio and video data frames in real time and caching them into a frame buffer queue, and obtaining network jitter parameters during the process of receiving the audio and video data frames in real time; According to the network jitter parameters, the frame rate of the audio and video data frames, and the cache length in the frame buffer queue, determining the cache length compensation duration corresponding to the currently received audio and video data frame; According to the cache length compensation duration and a preset reference compensation duration, determining the expected rendering time of the currently received audio and video data frame; Sending the corresponding audio and video data frame to a decoder for decoding and output at the expected rendering time.

2. The audio and video decoding output method according to claim 1, characterized in that, The step of determining the cache length compensation duration corresponding to the currently received audio and video data frame according to the network jitter parameters, the frame rate of the audio and video data frames, and the cache length in the frame buffer queue includes: When the cache length is greater than the maximum queue length corresponding to the frame rate, reducing the cache length compensation duration; When the network jitter parameter exceeds a preset first threshold and the cache length is less than the minimum queue length corresponding to the frame rate, increasing the cache length compensation duration; When the network jitter parameter is less than the preset first threshold, and / or the cache length is within the closed interval formed by the minimum queue length and the maximum queue length, maintaining the preset cache length compensation duration.

3. The audio and video decoding output method according to claim 2, characterized in that, The maximum queue length is obtained by multiplying the frame rate by a certain multiple.

4. The audio and video decoding output method according to claim 2, characterized in that, The minimum queue length is obtained by multiplying the product of the frame rate and the network jitter parameter by a certain multiple.

5. The audio and video decoding output method according to claim 2, characterized in that, The step of reducing the cache length compensation duration when the cache length is greater than the maximum queue length corresponding to the frame rate includes: When the cache length is greater than the maximum queue length corresponding to the frame rate, the reduced cache length compensation duration is determined according to the maximum queue length, the cache length, and the frame rate.

6. The audio and video decoding output method according to claim 2, characterized in that, The step of increasing the cache length compensation duration when the network jitter parameter exceeds a preset first threshold and the cache length is less than the minimum queue length corresponding to the frame rate includes: When the network jitter parameter exceeds a preset first threshold and the cache length is less than the minimum queue length corresponding to the frame rate, the increased cache length compensation duration is determined according to the minimum queue length, the cache length, and the frame rate.

7. The audio and video decoding output method according to any one of claims 1-6, characterized in that, The reference compensation duration is the sum of encoding delay compensation, rendering delay compensation, network jitter compensation, and round-trip delay compensation.

8. Audio and video decoding output device, characterized in that, comprising: A data receiving unit, configured to receive audio and video data frames in real time, buffer them into a frame buffer queue, and obtain network jitter parameters during the process of receiving the audio and video data frames in real time; A length compensation unit, configured to confirm the cache length compensation duration corresponding to the currently received audio and video data frame according to the network jitter parameters, the frame rate of the audio and video data frame, and the cache length in the frame buffer queue; A rendering time confirmation unit, configured to confirm the expected rendering time of the currently received audio and video data frame according to the cache length compensation duration and a preset reference compensation duration; A decoding and sending unit, configured to send the corresponding audio and video data frame to a decoder for decoding and output at the expected rendering time.

9. An electronic device, characterized in that, it includes: one or more processors; a memory, configured to store one or more computer programs; when the one or more computer programs are executed by the one or more processors, enabling the electronic device to implement the audio and video decoding and output method according to any one of claims 1-7.

10. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, it implements the audio and video decoding and output method according to any one of claims 1-7.