Video playback method, apparatus and system, device and storage medium

By adopting a coding structure of bidirectional predictive coding frames in real-time communication and outputting coded frames immediately when generated, combined with frame sequence adjustment and rendering at the receiving end, the problem of poor video compression effect in the existing technology is solved, and the cost of video transmission bandwidth is reduced and smooth playback is achieved.

WO2025208998A1PCT designated stage Publication Date: 2025-10-09BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/072854
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-02
Filing Date
2025-01-16
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

In existing real-time communication video coding methods, the use of bidirectional predictive coding frames leads to poor video compression effect and increases the video transmission bandwidth cost.

Method used

When the current delay of the receiving end and the preset allowed delay are met, a coding structure containing bidirectional predictive coding frames is used to encode the video frames, and bidirectional predictive coding frames are output immediately when generated. The receiving end adjusts the frame sequence and renders it to ensure smooth video playback.

Benefits of technology

It improves video compression, reduces video transmission bandwidth costs, and optimizes encoding delay without affecting the real-time communication experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025072854_09102025_PF_FP_ABST
    Figure CN2025072854_09102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are video playback methods, apparatus and system, a device and a storage medium. A video playback method comprises: if it is detected that a bidirectional predictive coding condition is currently satisfied, determining a target image group coding structure, there being at least one bidirectional predictive coded frame between two adjacent target frames in the target image group coding structure; controlling a coder to code, on the basis of the target image group coding structure, the current video frame sequence and output same, the bidirectional predictive coded frame being output during generation; and transmitting the current coded frame sequence to a receiver, so that the receiver decodes the current coded frame sequence, performs frame sequence adjustment on an obtained first video frame sequence to obtain a second video frame sequence, renders the second video frame sequence on the basis of a video frame collection time interval, and plays back a rendered third video frame sequence. Thus, without affecting the experience of real-time communications, the present disclosure uses coding structures comprising bidirectional predictive coded frames for coding, thus improving video compression effects.
Need to check novelty before this filing date? Find Prior Art

Description

Video playback method, device, system, equipment and storage medium

[0001] This application claims priority to Chinese Patent Application No. 202410396709.1 filed on April 2, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field

[0002] The embodiments of the present disclosure relate to a video playback method, apparatus, system, device, and storage medium. Background Art

[0003] With the rapid development of computer technology, real-time transmission of video data is often required to achieve real-time communications (RTC). Currently, in real-time communication scenarios, video frames after key frames are usually encoded as forward predictive coding frames (i.e., P frames) to reduce coding latency and thus ensure the user's real-time communication experience. However, compared with bidirectional predictive coding frames (i.e., B frames), the video compression effect of existing coding methods is poor, which in turn increases the bandwidth cost of video transmission. Summary of the Invention

[0004] The present disclosure provides a video playback method, apparatus, system, device and storage medium, which utilize a coding structure including bidirectional predictive coding frames for encoding without affecting the real-time communication experience, thereby improving the video compression effect and reducing the video transmission bandwidth cost.

[0005] In a first aspect, an embodiment of the present disclosure provides a video playback method, comprising:

[0006] If it is detected that a bidirectional predictive coding condition is currently satisfied based on a current delay of the receiving end and a preset allowed delay, determining a target group of pictures coding structure, wherein at least one bidirectional predictive coding frame exists between two adjacent target frames in the target group of pictures coding structure, wherein the target frames include key frames or forward predictive coding frames;

[0007] Controlling the encoder to encode and output the currently acquired current video frame sequence based on the target group of pictures coding structure, thereby obtaining a current coded frame sequence output by the encoder, wherein the bidirectionally predicted coded frames in the current coded frame sequence are output during generation;

[0008] The current encoded frame sequence is sent to the receiving end so that the receiving end decodes the current encoded frame sequence, adjusts the frame sequence of the first video frame sequence obtained by decoding, obtains a second video frame sequence having the same frame sequence as the current video frame sequence, and renders the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence, and plays the rendered third video frame sequence.

[0009] In a second aspect, the embodiments of the present disclosure further provide a video playback method, including:

[0010] Receiving a current coded frame sequence sent by a transmitting end, wherein the current coded frame sequence is obtained by the transmitting end controlling an encoder to encode and output a currently acquired current video frame sequence based on a target group of pictures coding structure, wherein the target group of pictures coding structure is determined based on a current delay and a preset allowed delay of the receiving end and when it is detected that a bidirectional predictive coding condition is currently satisfied, wherein at least one bidirectional predictive coding frame exists between two adjacent target frames in the target group of pictures coding structure, wherein the target frames include key frames or forward predictive coding frames;

[0011] Decoding the current coded frame sequence, and adjusting the frame sequence of the first video frame sequence obtained by decoding to obtain a second video frame sequence having the same frame sequence as the current video frame sequence;

[0012] Based on the video frame capture time interval in the current video frame sequence, the second video frame sequence is rendered, and the rendered third video frame sequence is played.

[0013] In a third aspect, the present disclosure further provides a video playback device, including:

[0014] a coding structure determining module, configured to, if it is detected based on a current delay of the receiving end and a preset allowed delay that a bidirectional predictive coding condition is currently satisfied, determine a target group of pictures coding structure, wherein at least one bidirectional predictive coding frame exists between two adjacent target frames in the target group of pictures coding structure, wherein the target frames include key frames or forward predictive coding frames;

[0015] an encoding control module, configured to control an encoder to encode and output a currently acquired current video frame sequence based on the target group of pictures encoding structure, thereby obtaining a current coded frame sequence output by the encoder, wherein bidirectionally predicted coded frames in the current coded frame sequence are output during generation;

[0016] The coding frame sequence sending module is used to send the current coding frame sequence to the receiving end so that the receiving end decodes the current coding frame sequence, adjusts the frame sequence of the first video frame sequence obtained by decoding, obtains a second video frame sequence with the same frame sequence as the current video frame sequence, and renders the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence, and plays the rendered third video frame sequence.

[0017] In a fourth aspect, an embodiment of the present disclosure further provides a video playback device, comprising:

[0018] a coded frame sequence receiving module, configured to receive a current coded frame sequence sent by a transmitting end, wherein the current coded frame sequence is obtained by the transmitting end controlling an encoder to encode and output a currently acquired current video frame sequence based on a target group of pictures coding structure, wherein the target group of pictures coding structure is determined based on a current delay and a preset allowed delay of the receiving end, and when it is detected that a bidirectional predictive coding condition is currently satisfied, wherein at least one bidirectional predictive coding frame exists between two adjacent target frames in the target group of pictures coding structure, wherein the target frames include key frames or forward predictive coding frames;

[0019] A coded frame decoding module is used to decode the current coded frame sequence and adjust the frame sequence of the first video frame sequence obtained by decoding to obtain a second video frame sequence having the same frame sequence as the current video frame sequence;

[0020] The rendering and playing module is used to render the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence, and play the rendered third video frame sequence.

[0021] In a fifth aspect, the embodiment of the present disclosure further provides a video playback system, the system comprising a transmitting end and a receiving end; wherein,

[0022] The sending end is used to implement the video playback method provided in the first aspect;

[0023] The receiving end is used to implement the video playback method provided in the second aspect.

[0024] In a sixth aspect, an embodiment of the present disclosure further provides an electronic device, comprising:

[0025] one or more processors;

[0026] a storage device for storing one or more programs,

[0027] When the one or more programs are executed by the one or more processors, the one or more processors implement the video playback method as described in any one of the embodiments of the present disclosure.

[0028] In a seventh aspect, an embodiment of the present disclosure further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the video playback method as described in any one of the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0030] FIG1 is a flow chart of a video playback method provided by an embodiment of the present disclosure;

[0031] FIG2 is an example diagram of a reference relationship of a bidirectional predictive coding frame involved in an embodiment of the present disclosure;

[0032] FIG3 is an example diagram of the entire link of a real-time video communication according to an embodiment of the present disclosure;

[0033] FIG4 is an example diagram of a video frame sequence processing process involved in an embodiment of the present disclosure;

[0034] FIG5 is a flow chart of a video playback method provided by an embodiment of the present disclosure;

[0035] FIG6 is a schematic structural diagram of a video playback device provided by an embodiment of the present disclosure;

[0036] FIG7 is a structural diagram of a video playback device provided by an embodiment of the present disclosure;

[0037] FIG8 is a structural diagram of a video playback system provided by an embodiment of the present disclosure; and

[0038] FIG9 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0039] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0040] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0041] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0042] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0043] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0044] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0045] Figure 1 is a flow chart of a video playback method provided by an embodiment of the present disclosure. This embodiment of the present disclosure is applicable to playing videos in real-time communication scenarios, where real-time communication scenarios may include business scenarios such as live broadcasting, remote video conferencing, or cloud gaming. The method can be executed by a video playback device, which can be implemented in the form of software and / or hardware and integrated into the transmitting end. Optionally, the transmitting end can be implemented by an electronic device, which can be a mobile terminal, a PC, or a server.

[0046] As shown in FIG1 , the video playback method specifically includes the following steps:

[0047] S110. If it is detected that the bidirectional prediction coding condition is currently met based on the current delay of the receiving end and the preset allowed delay, a target image group coding structure is determined, in which there is at least one bidirectional prediction coding frame between two adjacent target frames in the target image group coding structure, wherein the target frame includes a key frame or a forward prediction coding frame.

[0048] The receiving end can be a terminal corresponding to the transmitting end, used to receive and play videos. There can be one or more receiving ends. For example, in a remote video conferencing scenario, the transmitting end refers to the conference speaker, and the receiving end refers to the audience. The current delay refers to the time between video capture and playback. The current delay can change dynamically with changes in network status and coding structure. The preset allowable delay can refer to the maximum allowable video delay preset based on business needs and scenarios. For example, a maximum delay that is imperceptible to the user can be used as the preset allowable delay to ensure real-time communication. The bidirectional predictive coding condition can refer to the conditions that must be met for encoding using a bidirectional predictive coding structure. The bidirectional predictive coding structure refers to the group of pictures coding structure that contains bidirectional predictive coded frames. The target group of pictures coding structure can refer to the bidirectional predictive coding structure currently being used. The key frame (i.e., I-frame) is the first frame in each group of pictures. I-frame encoding and decoding only utilize the key frame's own information, without reference to other video frames. Forward predictive coded frames (i.e., P-frames) require reference to the previous frame to generate the encoding. Bidirectionally predicted coded frames (i.e., B frames) need to be coded and generated by referring to both the previous and next frames of the current frame. B frames have higher compression efficiency than P frames. For example, the target image group coding structure can be: IBPBPBP, in which there is only one B frame between an I frame and a P frame, and between two adjacent P frames. Alternatively, the target image group coding structure can also be: IBBPBPBBP, in which there are two B frames between an I frame and a P frame, and between two adjacent P frames. The more B frames there are between an I frame and a P frame, and between two adjacent P frames, the higher the compression efficiency, but the encoding and decoding time will also be longer.

[0049] Specifically, the transmitting end can detect whether the bidirectional predictive coding conditions are currently met based on the current delay of the receiving end and the preset allowed delay in real time or periodically. For example, if the current delay of each receiving end is less than the preset allowed delay, it means that each receiving end can appropriately increase the delay within the allowed range to avoid affecting the user's real-time communication experience. At this time, it can be determined that the bidirectional predictive coding conditions are currently met. If there is at least one receiving end whose current delay is greater than or equal to the preset allowed delay, it is determined that the bidirectional predictive coding conditions are not currently met. The process of detecting whether the bidirectional predictive coding conditions are currently met can be implemented in at least two ways as follows:

[0050] As an implementation, "detecting that a bidirectional predictive coding condition is currently satisfied based on a current delay of the receiving end and a preset allowable delay" in S110 may include: obtaining a current delay sent by at least one receiving end; and determining that the bidirectional predictive coding condition is currently satisfied if the current delay of each receiving end is less than the preset allowable delay. The receiving end transmits the current delay to the transmitting end, so that the transmitting end detects in real time whether the bidirectional predictive coding condition is currently satisfied based on the received current delay.

[0051] As another implementation method, "based on the current delay of the receiving end and the preset allowed delay, detecting that the bidirectional predictive coding condition is currently met" in S110 may include: sending the preset allowed delay to at least one receiving end, so that each receiving end compares the current delay with the received preset allowed delay, determines and returns the comparison result; if each comparison result received is that the current delay is less than the preset allowed delay, then it is determined that the bidirectional predictive coding condition is currently met. The sending end sends the preset allowed delay to each receiving end, and the receiving end compares the received preset allowed delay with the current delay and returns the comparison result to the sending end. The sending end detects in real time whether the bidirectional predictive coding condition is currently met based on the received comparison result. This end-to-end feedback method in which the receiving end feeds back information to the sending end in real time can ensure that the sending end starts bidirectional predictive coding at an appropriate time, thereby avoiding affecting the user's real-time communication experience.

[0052] Specifically, when the transmitter detects that the bidirectional predictive coding conditions are currently met, it indicates that bidirectional predictive coding is currently allowed to be enabled. At this time, a target image group coding structure containing at least one bidirectional predictive coding frame can be determined. For example, referring to Figure 2, the determined target image group coding structure is IBPBPBP, in which each B frame needs to be encoded with reference to the previous and next frames, and each P frame needs to be encoded with reference to the previous I frame or P frame. Therefore, the B frame needs to wait until the encoding of the referenced P frame is completed before it can be encoded. For example, after encoding the first frame (I frame), the third frame (P frame) needs to be encoded first, and then the second frame (B frame) needs to be encoded based on the first and third frames, resulting in a certain delay in B frame encoding. To address this, it is necessary to optimize the modules in the entire link of real-time video communication to reduce the delay caused by B frame encoding.

[0053] As shown in Figure 3, in a real-time communication scenario, the entire video real-time communication chain includes: an acquisition module, a pre-processing module, an encoder, a transmission module, a network transmission module, a receive buffer module, a decoder, a post-processing module, and a rendering module. This embodiment can further reduce the delay caused by B-frame encoding by optimizing the encoder and decoder, achieving optimal delay.

[0054] S120. Control the encoder to encode and output the currently acquired current video frame sequence based on the target image group coding structure to obtain a current coded frame sequence output by the encoder, wherein the bidirectionally predicted coded frames in the current coded frame sequence are output during generation.

[0055] The current video frame sequence may refer to at least one video frame currently captured by the transmitter, that is, the video frame that the receiver needs to play in real time. The current coded frame sequence includes the coded frames obtained after encoding each video frame in the current video frame sequence. The coded frames can be I frames, P frames, or B frames. Since video frames are input into the encoder one by one for encoding, the encoder is controlled to output each coded frame immediately after it is generated, without having to wait until the next frame is input. This avoids the increased delay in B-frame output caused by the encoder's single-input, single-output method, thereby reducing the delay of the entire link.

[0056] Specifically, by optimizing the encoder to control the encoder to encode the current video frame sequence currently collected according to the target image group encoding structure, and output it immediately when each encoded frame is generated, the bidirectional predictive encoded frame can be output immediately when it is generated without waiting, thereby quickly obtaining the current encoded frame sequence output by the encoder.

[0057] Exemplarily, S120 may include: controlling the encoder to determine the frame type corresponding to each current video frame in the currently captured current video frame sequence based on the target image group coding structure; encoding the current video frame based on the frame type, and outputting the coded frame when generating the coded frame to obtain the current coded frame sequence output by the encoder; wherein the encoder continuously outputs the forward predicted coded frame and at least one bidirectional predicted coded frame adjacent to the forward predicted coded frame.

[0058] Specifically, each current video frame in the currently captured current video frame sequence is sequentially input into the encoder. The encoder determines the frame type corresponding to the currently input current video frame based on the target group of pictures coding structure, encodes the current video frame according to the frame type, generates an encoded frame corresponding to the current video frame, and controls the generated encoded frame to be output immediately. For example, if the frame type corresponding to the current video frame is an I-frame, the current video frame is directly intra-encoded to generate an I-frame. If the frame type corresponding to the current video frame is a B-frame, the reference frame before the current video frame (e.g., an I-frame) and the reference frame after the current video frame (e.g., a P-frame) are determined. Since the subsequent reference frame has not yet been encoded, it is necessary to wait. If the frame type corresponding to the current video frame is a P-frame, the reference frame before the current video frame (e.g., an I-frame or a P-frame) is determined. Since the previous reference frame has already been encoded, the current video frame can be encoded based on the previous reference frame to obtain a P-frame. After the P-frame is encoded, the encoded P-frame is used as the reference frame after the B-frame to encode the current video frame to obtain a B-frame. The encoder can encode I-frames or P-frames directly without waiting, so that they can be output immediately when they are generated. When encoding B-frames, the encoder needs to wait for the subsequent P-frame to be encoded. After the P-frame encoding is completed, it can encode immediately, so that B-frames can be generated immediately after the P-frame is generated. At this time, the encoder is controlled to output the generated B-frame immediately, without waiting for the next video frame to be input before outputting the B-frame. This allows the P-frame and the adjacent B-frame to be output continuously without a certain interval between outputs.

[0059] It should be noted that when the encoder uses hardware encoding or high-performance software encoding, the encoding time per frame is usually very small. If the encoding time is ignored, the encoder will output B frames at the same time as P frames, thus avoiding B frame output delay.

[0060] Referring to Figure 4 , the captured current video frame sequence includes five video frames and is encoded according to the target group of pictures coding structure IBPBPBP. Specifically, the first frame (i.e., the video frame captured at 0 ms) corresponds to an I-frame type, the second frame corresponds to a B-frame type, the third frame corresponds to a P-frame type, the fourth frame corresponds to a B-frame type, and the fifth frame corresponds to a P-frame type. Based on the frame type corresponding to each frame, each video frame is encoded into a corresponding coded frame, and each coded frame is output as it is generated. For example, the third frame (P-frame) and the second frame (B-frame) are output consecutively. If encoding time is not considered, the third frame (P-frame) is input at approximately 132 ms, and both the third frame (P-frame) and the second frame (B-frame) are output at approximately 132 ms. This allows the second frame (B-frame) to be output promptly upon generation, eliminating the need for output at 198 ms, thus avoiding B-frame output delays.

[0061] S130. Send the current encoded frame sequence to the receiving end so that the receiving end decodes the current encoded frame sequence, adjusts the frame sequence of the first video frame sequence obtained by decoding, obtains a second video frame sequence having the same frame sequence as the current video frame sequence, renders the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence, and plays the rendered third video frame sequence.

[0062] The frame order of the current video frame sequence refers to the order in which the video frames are captured, i.e., the order in which they are played. The frame order of the first video frame sequence refers to the decoding order of the video frames, i.e., the encoding order. Due to the presence of B frames, the frame order of the first video frame sequence is different from the frame order of the current video frame sequence. For example, referring to FIG4 , the frame order of the current video frame sequence is IBPBPBP, and the frame order of the first video frame sequence is IPBPBPB. The video frame capture time interval refers to the interval at which a video frame is captured. For example, the video frame capture time interval in FIG4 is 66 ms.

[0063] Specifically, the transmitter sends the current coded frame sequence output by the encoder to the receiver. The receiver inputs the current coded frame sequence into the decoder. The decoder decodes and outputs the corresponding video frames in sequence according to the input coded frame order, thereby obtaining a first video frame sequence output by the decoder. The first video frame sequence has the same temporal positional relationship as the current coded frame sequence, namely, the same frame order and frame interval, as shown in Figure 4. In other words, the decoder only performs decoding operations and does not adjust the timing information of the input and output sequences, thereby improving decoding efficiency. To ensure normal video playback, the frame order of the first video frame sequence needs to be adjusted to obtain a second video frame sequence with the same frame order as the current video frame sequence. As shown in Figure 4, the frame order of the second video frame sequence is IBPBPBP. It should be noted that when adjusting the frame order, only the order of the frames needs to be adjusted, and the relative time intervals between frames do not need to be adjusted. Furthermore, no P-frame buffering occurs, allowing for rapid B-frame output. Since the frame intervals in the adjusted second video frame sequence are uneven, it is necessary to uniformly render the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence. For example, the interval between every two frames in the rendered third video frame sequence is the video frame acquisition time interval, so that the third video frame sequence can be played smoothly and video playback freezes can be avoided.

[0064] Exemplarily, after S130, it also includes: if it is detected that the bidirectional predictive coding conditions are not currently met based on the current delay of the receiving end and the preset allowed delay, the encoder is controlled to encode the current video frame sequence currently collected based on a non-bidirectional predictive coding structure, such as IPPPPP, and the current encoded frame sequence output by the encoder is sent to the receiving end, so that the receiving end decodes the current encoder, and renders and plays the decoded video frame sequence.

[0065] Specifically, after the sending end turns on bidirectional predictive coding, if it detects that the bidirectional predictive coding conditions are not currently met, it needs to turn off bidirectional predictive coding and use the non-bidirectional predictive coding structure for encoding. In this way, bidirectional predictive coding is only turned on when the delay requirements are met, realizing dynamic adjustment of the coding structure to avoid affecting the user's real-time communication experience.

[0066] The technical solution of the embodiment of the present disclosure is to determine the target image group coding structure containing at least one bidirectional predictive coding frame by detecting that the bidirectional predictive coding conditions are currently met based on the current delay of the receiving end and the preset allowed delay, and to control the encoder to encode and output the current video frame sequence currently captured based on the target image group coding structure, thereby obtaining the current coded frame sequence output by the encoder, wherein the encoder immediately outputs the bidirectional predictive coding frame when generating the bidirectional predictive coding frame, thereby reducing the coding delay of the bidirectional predictive coding frame. The current coded frame sequence is sent to the receiving end so that the receiving end decodes the current coded frame sequence, adjusts the frame sequence of the first video frame sequence obtained by decoding, and obtains a second video frame sequence with the same frame sequence as the current video frame sequence, so as to ensure that the video is played normally according to the acquisition order. Based on the video frame acquisition time interval in the current video frame sequence, the second video frame sequence is evenly rendered, so that the rendered third video frame sequence can be played smoothly, avoiding playback freezes. By reusing a coding structure containing bidirectionally predicted coding frames when the conditions for bidirectionally predicted coding are detected, the coding structure is dynamically adjusted to avoid impacting the real-time communication experience. Using bidirectionally predicted coding frames for encoding improves video compression, thereby reducing video transmission bandwidth costs. Furthermore, by controlling the encoder to output generated bidirectionally predicted coding frames promptly, eliminating the need for a waiting period, the encoding latency of the bidirectionally predicted coding frames is reduced, further enhancing the usability of the bidirectionally predicted coding structure in real-time communication scenarios.

[0067] On the basis of the above-mentioned technical solutions, "based on the current delay of the receiving end and the preset allowed delay, detecting that the bidirectional predictive coding conditions are currently met" in S110 may include: subtracting the current delay of the receiving end from the preset allowed delay to obtain the delay difference corresponding to the receiving end; if the delay difference is greater than or equal to the preset delay, it is determined that the bidirectional predictive coding conditions are currently met.

[0068] Among them, the preset delay is determined in advance based on the video frame acquisition time interval. The preset delay can be used to characterize the reference delay that may be increased by using B-frame encoding. For example, by optimizing the encoder and decoder, when encoding using a coding structure in which there is a bidirectional predictive coding frame between two adjacent target frames, the delay can be reduced from more than 2 frames to about 1 frame. For example, if the video frame acquisition time interval is 66ms, the delay using a coding structure including 1 B frame can be reduced from more than 132ms to about 66ms. Accordingly, the video frame acquisition time interval can be used as the preset delay.

[0069] Specifically, the preset allowed delay is subtracted from the current delay of each receiving end to obtain the delay difference corresponding to each receiving end. If the delay difference corresponding to each receiving end is greater than or equal to the preset delay, it indicates that the delay caused by the bidirectional predictive coding frame is currently acceptable with a high probability, and it can be determined that the bidirectional predictive coding conditions are currently met. If there is at least one receiving end whose corresponding delay difference is less than the preset delay, it is determined that the bidirectional predictive coding conditions are not currently met. By comparing the delay difference with the preset delay, it is possible to more accurately decide whether to turn on bidirectional predictive coding, thereby avoiding the situation of repeatedly turning on and off bidirectional predictive coding, thereby saving equipment resources.

[0070] Exemplarily, the transmitting end may make a decision based on the current delay sent by the receiving end. For example, the current delay sent by at least one receiving end is obtained; the current delay sent by the receiving end is subtracted from the preset allowed delay to obtain the delay difference corresponding to the receiving end; if the delay difference is greater than or equal to the preset delay, it is determined that the bidirectional predictive coding condition is currently met. Alternatively, the transmitting end may also make a decision by sending the preset allowed delay and the preset delay to each receiving end. For example, the preset allowed delay and the preset delay are sent to at least one receiving end, so that each receiving end subtracts the current delay from the received preset allowed delay to obtain the delay difference, and compares the delay difference with the received preset delay to determine and return the comparison result; if each comparison result received is that the delay difference is greater than or equal to the preset delay, it is determined that the bidirectional predictive coding condition is currently met.

[0071] On the basis of the above-mentioned technical solutions, "determining the target picture group coding structure" in S110 may include: if the current picture group coding structure is a non-bidirectional prediction coding structure, determining the target picture group coding structure from at least one candidate picture group coding structure; if the current picture group coding structure is a bidirectional prediction coding structure, determining the target picture group coding structure from at least two candidate picture group coding structures based on the current picture group coding structure.

[0072] The current group of pictures coding structure refers to the group of pictures coding structure currently used by the encoder at the transmitting end. The current group of pictures coding structure can be a non-bidirectional predictive coding structure or a bidirectional predictive coding structure. A non-bidirectional predictive coding structure refers to a group of pictures coding structure that does not contain B frames, such as IPPPPP. A bidirectional predictive coding structure refers to a group of pictures coding structure that contains B frames, such as IBPBPBP. A candidate group of pictures coding structure refers to a bidirectional predictive coding structure that is allowed to be selected. The number of candidate group of pictures coding structures can be one or more. Different numbers of bidirectional predictive coding frames exist between two adjacent target frames in different candidate group of pictures coding structures. For example, there are two candidate group of pictures coding structures: IBPBPBP and IBBPBBPBP.

[0073] Specifically, when the conditions for bidirectional predictive coding are currently met, it is possible to detect whether the current GOP coding structure is a non-bidirectional predictive coding structure or a bidirectional predictive coding structure. If the current GOP coding structure is a non-bidirectional predictive coding structure, this indicates that bidirectional predictive coding has not yet been enabled. In this case, one can be selected from all candidate GOP coding structures as a target GOP coding structure in order to enable bidirectional predictive coding. If the current GOP coding structure is a bidirectional predictive coding structure (in which case there will be at least two candidate GOP coding structures), this indicates that bidirectional predictive coding has already been enabled, and further delay increases, i.e., an increase in the number of B-frame encodings, are permitted. In this case, an optimal one can be selected from all candidate GOP coding structures as the target GOP coding structure.

[0074] Exemplarily, determining a target image group coding structure from at least one candidate image group coding structure may include: if there is only one candidate image group coding structure, determining the candidate image group coding structure as the target image group coding structure, wherein there is only one bidirectionally predicted coding frame between two adjacent target frames in the candidate image group coding structure; if there are at least two candidate image group coding structures, determining the candidate image group coding structure with the largest or smallest number of bidirectionally predicted coding frames between two adjacent target frames as the target image group coding structure.

[0075] Specifically, when the current picture group coding structure is a non-bidirectional prediction coding structure, if there is only one candidate picture group coding structure, and the candidate picture group coding structure is a coding result containing a B frame, then the candidate picture group coding structure can be directly determined as the target picture group coding structure, so the coding structure used each time is switched between the non-bidirectional prediction coding structure and the candidate picture group coding structure, that is, it can only decide whether to turn on bidirectional prediction coding.

[0076] If the current GOP coding structure is not a bidirectional predictive coding structure, and there are at least two candidate GOP coding structures, the candidate GOP coding structure with the largest number of B frames (i.e., the largest latency) can be determined as the target GOP coding structure, so that the bidirectional predictive coding structure with the largest number of B frames is preferentially used, further improving compression efficiency. Alternatively, the candidate GOP coding structure with the smallest number of B frames (i.e., the smallest latency) can be determined as the target GOP coding structure, so that the bidirectional predictive coding structure with the smallest number of B frames is preferentially used, avoiding the situation where the latency requirement cannot be met due to an excessive number of B frames.

[0077] Exemplarily, determining a target image group coding structure from at least two candidate image group coding structures based on a current image group coding structure may include: determining a target number based on a current number of bidirectionally predicted coding frames existing between two adjacent target frames in the current image group coding structure, wherein the target number is greater than the current number; and determining the target image group coding structure based on the candidate number and target number of bidirectionally predicted coding frames existing between two adjacent target frames in the candidate image group coding structure.

[0078] Specifically, when the current GOP coding structure is a bidirectional predictive coding structure, a target number greater than the current number can be determined, such as by adding 1 to the current number as the target number, and a candidate GOP coding structure with the target number as the target GOP coding structure is determined as the target GOP coding structure, thereby gradually increasing the number of B frames in the coding structure to use as many B frames as possible for encoding without affecting the user experience, thereby further improving video compression efficiency. It should be noted that if there is no candidate GOP coding structure with the target number as the target number, it indicates that the number of B frames in the current GOP coding structure has reached the maximum. In this case, there is no need to switch the coding structure, and encoding can continue using the current GOP coding structure as the target GOP coding structure.

[0079] It should be noted that if, based on the current delay of the receiving end and the preset allowed delay, it is detected that the bidirectional predictive coding conditions are not currently met, and the current picture group coding structure is a bidirectional predictive coding structure, it indicates that the delay requirements cannot be met due to excessive number of B frames in the current picture group coding structure. At this time, a target number smaller than the current number can be determined, such as the result of subtracting 1 from the current number as the target number, and the candidate picture group coding structure with the target number as the candidate number is determined as the target picture group coding structure, thereby gradually reducing the number of B frames in the coding structure when the delay requirements are not met, so as to achieve a balance between compression efficiency and playback delay.

[0080] Figure 5 is a flowchart of a video playback method provided by an embodiment of the present disclosure. This embodiment of the present disclosure is applicable to playing videos in real-time communication scenarios, where real-time communication scenarios may include business scenarios such as live broadcasting, remote video conferencing, or cloud gaming. The method can be executed by a video playback device, which can be implemented in software and / or hardware and integrated into the receiving end. Optionally, the sending end can be implemented by an electronic device, which can be a mobile terminal, PC, or server.

[0081] As shown in FIG5 , the video playback method specifically includes the following steps:

[0082] S210. Receive a current coded frame sequence sent by a transmitting end, wherein the current coded frame sequence is obtained by the transmitting end controlling the encoder to encode and output the currently acquired current video frame sequence based on a target image group coding structure, the target image group coding structure is determined based on the current delay and the preset allowed delay of the receiving end, and is determined when it is detected that a bidirectional prediction coding condition is currently met, and there is at least one bidirectional prediction coding frame between two adjacent target frames in the target image group coding structure, wherein the target frame includes a key frame or a forward prediction coding frame.

[0083] The transmitting end may be the terminal corresponding to the receiving end, used to capture and transmit the generated video. For example, in a remote video conferencing scenario, the transmitting end refers to the conference speaker, and the receiving end refers to the conference audience. The current coded frame sequence includes the coded frames obtained after encoding each video frame in the current video frame sequence. The coded frames can be I-frames, P-frames, or B-frames.

[0084] Specifically, the process of determining the current coded frame sequence in the transmitting end can refer to the relevant description of the above embodiment, which will not be repeated here. The transmitting end sends the current coded frame sequence output by the encoder in real time to the receiving end through the network.

[0085] S220: Decode the current coded frame sequence, and adjust the frame sequence of the first video frame sequence obtained by decoding to obtain a second video frame sequence having the same frame sequence as the current video frame sequence.

[0086] The coded frames in the current coded frame sequence correspond one-to-one to the video frames in the first video frame sequence. The frame sequence of the current video frame sequence refers to the order in which the video frames were acquired, i.e., the order in which they were played. The frame sequence of the first video frame sequence refers to the order in which the video frames were decoded, i.e., the order in which they were encoded. Due to the presence of B frames, the frame sequence of the first video frame sequence is different from the frame sequence of the current video frame sequence. For example, referring to FIG4 , the frame sequence of the current video frame sequence is IBPBPBP, while the frame sequence of the first video frame sequence is IPBPBPB.

[0087] Specifically, the receiving end inputs the received current coded frame sequence into the decoder. The decoder decodes and outputs the corresponding video frames in sequence according to the input coded frame order, thereby obtaining a first video frame sequence output by the decoder. The first video frame sequence has the same temporal positional relationship as the current coded frame sequence, i.e., the same frame order and frame interval, as shown in FIG4 . In other words, the decoder only performs decoding operations and does not adjust the timing information of the input and output sequences, thereby improving decoding efficiency. To ensure normal video playback, the frame order of the first video frame sequence needs to be adjusted to obtain a second video frame sequence with the same frame order as the current video frame sequence. As shown in FIG4 , the frame order of the second video frame sequence is IBPBPBP. It should be noted that when adjusting the frame order, only the order of the frames needs to be adjusted, and the relative time intervals between frames do not need to be adjusted. Furthermore, no P-frame buffering is performed, allowing for rapid B-frame output. Therefore, the relative time intervals between the forward predictive coded frames and the bidirectional predictive coded frames in the first video frame sequence are the same as the relative time intervals between the forward predictive coded frames and the bidirectional predictive coded frames in the second video frame sequence. For example, as shown in FIG4 , the P frame in the first video frame sequence is placed after the adjacent B frame, and the relative time interval between the P frame and the B frame remains unchanged.

[0088] S230: Render the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence, and play the rendered third video frame sequence.

[0089] The video frame acquisition time interval refers to the time interval at which a video frame is acquired. For example, the video frame acquisition time interval in FIG4 is 66 ms.

[0090] Specifically, since the frame intervals in the adjusted second video frame sequence are uneven, it is necessary to uniformly render the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence. For example, the interval between every two frames in the rendered third video frame sequence is the video frame acquisition time interval, so that the third video frame sequence can be played smoothly, avoiding video playback freezes.

[0091] Exemplarily, S230 may include: adjusting the rendering time interval between two adjacent target video frames in the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence, so that the adjusted rendering time interval is the video frame acquisition time interval; rendering each video frame in the adjusted second video frame sequence, and playing the rendered third video frame sequence.

[0092] The target video frame may include a video frame corresponding to a forward predictive coding frame or a video frame corresponding to a bidirectional predictive coding frame. The video frame corresponding to the forward predictive coding frame refers to a video frame decoded from a P frame. The video frame corresponding to the bidirectional predictive coding frame refers to a video frame decoded from a B frame.

[0093] Specifically, referring to FIG4 , the rendering interval between each set of BP frames in the second video frame sequence is adjusted to the video frame acquisition interval, i.e., 66 ms. For example, the rendering time of the P frames in each set of BP frames is delayed after the video frame acquisition interval, so that the rendering intervals between the B frames and P frames in the adjusted second video frame sequence are equal. Each video frame in the adjusted second video frame sequence is rendered, thereby achieving a uniform rendering effect for the B frames and P frames, thereby ensuring the continuity of video playback.

[0094] It should be noted that it is only necessary to adjust the rendering time interval between the B frames and P frames in the second video frame sequence, and there is no need to adjust the rendering time interval between the I frames and the B frames. In other words, there is no need to delay the rendering time of the I frame so that the first frame can be played quickly, so that the user can quickly see the first frame and improve the video viewing experience.

[0095] On the basis of the above technical solutions, before S220, it can also include: smoothing the decoding start time of two adjacent forward predicted coding frames in the current coding frame sequence so that the decoding start time intervals between the two adjacent forward predicted coding frames after smoothing are equal; based on the decoding start time of the forward predicted coding frame after smoothing, synchronously updating the decoding start time of the bidirectional predicted coding frame adjacent to the forward predicted coding frame so that the temporal position relationship between the forward predicted coding frame and the bidirectional predicted coding frame before and after smoothing remains unchanged.

[0096] Specifically, referring to Figures 3 and 4 , the receive buffer module at the receiving end is optimized so that, after receiving the current coded frame sequence, only the decoding start time of two adjacent P frames in the current coded frame sequence is smoothed. For example, the decoding start time of the P frame is delayed so that the decoding start time intervals between the two adjacent P frames after smoothing are equal. This achieves smoothing between P frames under network jitter and ensures that the P frames are smoothly and stably delivered to the decoder for decoding without any jitter, thereby achieving network jitter resistance. After the decoding start time of the P frame is delayed, all B frames adjacent to the P frame are synchronously delayed so that the temporal position relationship between the P and B frames before and after smoothing remains unchanged. In other words, the receive buffer module at the receiving end only smooths the decoding start time of two adjacent P frames, eliminating the need to smooth the decoding start time of PB frames. Furthermore, the module supports the rapid output of PB frames without hoarding P frames, further reducing the delay increased by B frame encoding.

[0097] FIG6 is a schematic structural diagram of a video playback device provided by an embodiment of the present disclosure. As shown in FIG6 , the device specifically includes: a coding structure determination module 310 , a coding control module 320 and a coding frame sequence sending module 330 .

[0098] Among them, the coding structure determination module 310 is used to determine the target image group coding structure if it is detected that the bidirectional prediction coding condition is currently met based on the current delay of the receiving end and the preset allowed delay, wherein there is at least one bidirectional prediction coding frame between two adjacent target frames in the target image group coding structure, wherein the target frame includes a key frame or a forward prediction coding frame; the coding control module 320 is used to control the encoder to encode and output the current video frame sequence currently captured based on the target image group coding structure, and obtain the current coding frame sequence output by the encoder, wherein the bidirectional prediction coding frame in the current coding frame sequence is output when it is generated; the coding frame sequence sending module 330 is used to send the current coding frame sequence to the receiving end so that the receiving end decodes the current coding frame sequence, adjusts the frame sequence of the first video frame sequence obtained by decoding, obtains a second video frame sequence with the same frame sequence as the current video frame sequence, and renders the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence, and plays the rendered third video frame sequence.

[0099] The technical solution provided by the embodiment of the present disclosure is to determine the target image group coding structure containing at least one bidirectional predictive coding frame by detecting that the bidirectional predictive coding conditions are currently met based on the current delay of the receiving end and the preset allowed delay, and to control the encoder to encode and output the current video frame sequence currently captured based on the target image group coding structure, thereby obtaining the current coded frame sequence output by the encoder, wherein the encoder immediately outputs the bidirectional predictive coding frame when generating the bidirectional predictive coding frame, thereby reducing the coding delay of the bidirectional predictive coding frame. The current coded frame sequence is sent to the receiving end so that the receiving end decodes the current coded frame sequence, adjusts the frame sequence of the first video frame sequence obtained by decoding, and obtains a second video frame sequence with the same frame sequence as the current video frame sequence, so as to ensure that the video is played normally according to the acquisition order. Based on the video frame acquisition time interval in the current video frame sequence, the second video frame sequence is evenly rendered so that the rendered third video frame sequence can be played smoothly to avoid playback freezes. By reusing a coding structure containing bidirectionally predicted coding frames when the conditions for bidirectionally predicted coding are detected, the coding structure is dynamically adjusted to avoid impacting the real-time communication experience. Using bidirectionally predicted coding frames for encoding improves video compression, thereby reducing video transmission bandwidth costs. Furthermore, by controlling the encoder to output generated bidirectionally predicted coding frames promptly, eliminating the need for a waiting period, the encoding latency of the bidirectionally predicted coding frames is reduced, further enhancing the usability of the bidirectionally predicted coding structure in real-time communication scenarios.

[0100] Based on the above technical solution, the coding structure determination module 310 is configured to:

[0101] The preset allowed delay is subtracted from the current delay of the receiving end to obtain the delay difference corresponding to the receiving end; if the delay difference is greater than or equal to the preset delay, it is determined that the bidirectional predictive coding conditions are currently met, wherein the preset delay is determined in advance based on the video frame acquisition time interval.

[0102] Based on the above technical solutions, the coding structure determination module 310 includes:

[0103] a first determining unit, configured to determine a target GOP coding structure from at least one candidate GOP coding structure if the current GOP coding structure is a non-bidirectional prediction coding structure;

[0104] a second determining unit, configured to determine a target GOP coding structure from at least two candidate GOP coding structures based on the current GOP coding structure if the current GOP coding structure is a bidirectional predictive coding structure;

[0105] There are different numbers of bidirectional prediction coding frames between two adjacent target frames in different candidate image group coding structures.

[0106] On the basis of the above technical solutions, the first determining unit is specifically configured to:

[0107] If there is only one candidate image group coding structure, the candidate image group coding structure is determined as the target image group coding structure, wherein there is only one bidirectional prediction coding frame between two adjacent target frames in the candidate image group coding structure; if there are at least two candidate image group coding structures, the candidate image group coding structure with the largest or smallest number of bidirectional prediction coding frames between two adjacent target frames is determined as the target image group coding structure.

[0108] On the basis of the above technical solutions, the second determining unit is specifically configured to:

[0109] Based on the current number of bidirectionally predicted coding frames existing between two adjacent target frames in the current picture group coding structure, a target number is determined, wherein the target number is greater than the current number; based on the to-be-selected number of bidirectionally predicted coding frames existing between two adjacent target frames in the to-be-selected picture group coding structure and the target number, a target picture group coding structure is determined.

[0110] Based on the above technical solutions, the encoding control module 320 is specifically configured to:

[0111] The encoder is controlled to determine, based on the target image group coding structure, a frame type corresponding to each current video frame in the currently acquired current video frame sequence; the current video frame is encoded based on the frame type, and the encoded frame is output when the encoded frame is generated, to obtain a current encoded frame sequence output by the encoder; wherein the encoder continuously outputs a forward predicted encoded frame and at least one bidirectionally predicted encoded frame adjacent to the forward predicted encoded frame.

[0112] The video playback device provided in the embodiments of the present disclosure can execute the video playback method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0113] FIG7 is a schematic structural diagram of a video playback device provided by an embodiment of the present disclosure. As shown in FIG7 , the device specifically includes: a coding frame sequence receiving module 410 , a coding frame decoding module 420 , and a rendering and playback module 430 .

[0114] Among them, the coding frame sequence receiving module 410 is used to receive the current coding frame sequence sent by the sending end, wherein the current coding frame sequence is obtained by the sending end by controlling the encoder to encode and output the current video frame sequence currently collected based on the target image group coding structure, and the target image group coding structure is determined based on the current delay of the receiving end and the preset allowed delay, and it is detected when the bidirectional prediction coding condition is currently met. There is at least one bidirectional prediction coding frame between two adjacent target frames in the target image group coding structure, wherein the target frame includes a key frame or a forward prediction coding frame; the coding frame decoding module 420 is used to decode the current coding frame sequence, and adjust the frame sequence of the first video frame sequence obtained by decoding to obtain a second video frame sequence with the same frame sequence as the current video frame sequence; the rendering and playback module 430 is used to render the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence, and play the rendered third video frame sequence.

[0115] The technical solution provided by the embodiment of the present disclosure is to determine the target image group coding structure containing at least one bidirectional predictive coding frame by detecting that the bidirectional predictive coding conditions are currently met based on the current delay of the receiving end and the preset allowed delay, and to control the encoder to encode and output the current video frame sequence currently captured based on the target image group coding structure, thereby obtaining the current coded frame sequence output by the encoder, wherein the encoder immediately outputs the bidirectional predictive coding frame when generating the bidirectional predictive coding frame, thereby reducing the coding delay of the bidirectional predictive coding frame. The current coded frame sequence is sent to the receiving end so that the receiving end decodes the current coded frame sequence, adjusts the frame sequence of the first video frame sequence obtained by decoding, and obtains a second video frame sequence with the same frame sequence as the current video frame sequence, so as to ensure that the video is played normally according to the acquisition order. Based on the video frame acquisition time interval in the current video frame sequence, the second video frame sequence is evenly rendered so that the rendered third video frame sequence can be played smoothly to avoid playback freezes. By reusing a coding structure containing bidirectionally predicted coding frames when the conditions for bidirectionally predicted coding are detected, the coding structure is dynamically adjusted to avoid impacting the real-time communication experience. Using bidirectionally predicted coding frames for encoding improves video compression, thereby reducing video transmission bandwidth costs. Furthermore, by controlling the encoder to output generated bidirectionally predicted coding frames promptly, eliminating the need for a waiting period, the encoding latency of the bidirectionally predicted coding frames is reduced, further enhancing the usability of the bidirectionally predicted coding structure in real-time communication scenarios.

[0116] Based on the above technical solution, the first video frame sequence and the current coding frame sequence have the same temporal position relationship; the relative time interval between the forward prediction coding frame and the bidirectional prediction coding frame in the first video frame sequence is the same as the relative time interval between the forward prediction coding frame and the bidirectional prediction coding frame in the second video frame sequence.

[0117] Based on the above technical solutions, the rendering and playing module 430 is specifically used to:

[0118] Based on the video frame acquisition time interval in the current video frame sequence, adjust the rendering time interval between two adjacent target video frames in the second video frame sequence so that the adjusted rendering time interval is the video frame acquisition time interval, wherein the target video frames include video frames corresponding to forward predictive coding frames or video frames corresponding to bidirectional predictive coding frames; render each video frame in the adjusted second video frame sequence, and play the rendered third video frame sequence.

[0119] On the basis of the above technical solutions, the device also includes:

[0120] a smoothing processing module, configured to perform a smoothing process on the decoding start time of two adjacent forward predicted coding frames in the current coding frame sequence before decoding the current coding frame sequence, so that the decoding start time intervals between the two adjacent forward predicted coding frames after smoothing are equal;

[0121] The synchronous updating module is used to synchronously update the decoding start time of the bidirectional predictive coding frame adjacent to the forward predictive coding frame based on the decoding start time after smoothing of the forward predictive coding frame, so that the temporal position relationship between the forward predictive coding frame and the bidirectional predictive coding frame before and after smoothing remains unchanged.

[0122] The video playback device provided in the embodiments of the present disclosure can execute the video playback method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0123] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.

[0124] FIG8 is a schematic diagram of the structure of a video playback system provided by an embodiment of the present disclosure, which is applicable to playing videos in real-time communication scenarios. As shown in FIG8 , the system specifically includes: a transmitting end 510 and a receiving end 520 .

[0125] The transmitting end 510 is used to implement the video playback method provided by the above embodiment applied to the transmitting end; the receiving end 520 is used to implement the video playback method provided by the above embodiment applied to the receiving end.

[0126] In the video playback system of the embodiment of the present disclosure, the sending end determines the target image group coding structure containing at least one bidirectional predictive coding frame by detecting that the bidirectional predictive coding conditions are currently met based on the current delay of the receiving end and the preset allowed delay, and controls the encoder to encode and output the current video frame sequence currently captured based on the target image group coding structure, thereby obtaining the current coded frame sequence output by the encoder, wherein the encoder immediately outputs the bidirectional predictive coding frame when generating the bidirectional predictive coding frame, thereby reducing the coding delay of the bidirectional predictive coding frame. The current coded frame sequence is sent to the receiving end so that the receiving end decodes the current coded frame sequence, adjusts the frame sequence of the first video frame sequence obtained by decoding, and obtains a second video frame sequence with the same frame sequence as the current video frame sequence, so as to ensure that the video is played normally according to the acquisition order. Based on the video frame acquisition time interval in the current video frame sequence, the second video frame sequence is evenly rendered so that the rendered third video frame sequence can be played smoothly to avoid playback jams. By reusing a coding structure containing bidirectionally predicted coding frames when the conditions for bidirectionally predicted coding are detected, the coding structure is dynamically adjusted to avoid impacting the real-time communication experience. Using bidirectionally predicted coding frames for encoding improves video compression, thereby reducing video transmission bandwidth costs. Furthermore, by controlling the encoder to output generated bidirectionally predicted coding frames promptly, eliminating the need for a waiting period, the encoding latency of the bidirectionally predicted coding frames is reduced, further enhancing the usability of the bidirectionally predicted coding structure in real-time communication scenarios.

[0127] FIG9 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Referring to FIG9 , a schematic diagram of the structure of an electronic device (such as a terminal device or server in FIG9 ) 500 suitable for implementing an embodiment of the present disclosure is shown below. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in FIG9 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.

[0128] As shown in FIG9 , the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An edit / output (I / O) interface 505 is also connected to the bus 504.

[0129] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although FIG9 shows the electronic device 500 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may be implemented or present instead.

[0130] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0131] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0132] The electronic device provided by the embodiment of the present disclosure and the video playback method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0133] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon. When the program is executed by a processor, the video playback method provided in the above embodiment is implemented.

[0134] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0135] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0136] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0137] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: if it is detected that the bidirectional predictive coding condition is currently met based on the current delay of the receiving end and the preset allowed delay, then the electronic device determines the target image group coding structure, wherein there is at least one bidirectional predictive coding frame between two adjacent target frames in the target image group coding structure, wherein the target frame includes a key frame or a forward predictive coding frame; controls the encoder to encode and output the current video frame sequence currently captured based on the target image group coding structure, and obtains the current coded frame sequence output by the encoder, wherein the bidirectional predictive coding frame in the current coded frame sequence is output during generation; sends the current coded frame sequence to the receiving end so that the receiving end decodes the current coded frame sequence, adjusts the frame sequence of the first video frame sequence obtained by decoding, obtains a second video frame sequence having the same frame sequence as the current video frame sequence, renders the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence, and plays the rendered third video frame sequence.

[0138] Alternatively, the computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device: receives a current coded frame sequence sent by a transmitter, wherein the current coded frame sequence is obtained by the transmitter controlling the encoder to encode and output the currently captured current video frame sequence based on a target image group coding structure, wherein the target image group coding structure is determined based on the current delay of the receiver and a preset allowed delay, and when it is detected that a bidirectional predictive coding condition is currently met, and there is at least one bidirectional predictive coding frame between two adjacent target frames in the target image group coding structure, wherein the target frame includes a key frame or a forward predictive coding frame; decodes the current coded frame sequence, and adjusts the frame sequence of the first video frame sequence obtained by decoding to obtain a second video frame sequence having the same frame sequence as the current video frame sequence; renders the second video frame sequence based on the video frame capture time interval in the current video frame sequence, and plays the rendered third video frame sequence.

[0139] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0141] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."

[0142] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0143] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0144] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0145] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0146] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A video playback method, comprising: If it is detected that a bidirectional predictive coding condition is currently satisfied based on a current delay of the receiving end and a preset allowed delay, determining a target group of pictures coding structure, wherein at least one bidirectional predictive coding frame exists between two adjacent target frames in the target group of pictures coding structure, wherein the target frames include key frames or forward predictive coding frames; Controlling the encoder to encode and output the currently acquired current video frame sequence based on the target group of pictures coding structure, to obtain a current coded frame sequence output by the encoder, wherein the bidirectionally predicted coded frames in the current coded frame sequence are output during generation; The current encoded frame sequence is sent to a receiving end so that the receiving end decodes the current encoded frame sequence, adjusts the frame sequence of a first video frame sequence obtained by decoding, obtains a second video frame sequence having the same frame sequence as the current video frame sequence, renders the second video frame sequence based on a video frame acquisition time interval in the current video frame sequence, and plays the rendered third video frame sequence.

2. The video playback method according to claim 1, wherein: The detecting that a bidirectional predictive coding condition is currently satisfied based on a current delay of the receiving end and a preset allowed delay includes: Subtract the current delay of the receiving end from the preset allowed delay to obtain a delay difference corresponding to the receiving end; If the delay difference is greater than or equal to a preset delay, it is determined that the bidirectional predictive coding condition is currently met, wherein the preset delay is determined in advance based on the video frame acquisition time interval.

3. The video playback method according to claim 1 or 2, wherein: The determining of the target group of pictures coding structure includes: If the current GOP coding structure is a non-bidirectional prediction coding structure, determining the target GOP coding structure from at least one candidate GOP coding structure; If the current GOP coding structure is a bidirectional prediction coding structure, determining the target GOP coding structure from at least two candidate GOP coding structures based on the current GOP coding structure; There are different numbers of bidirectional prediction coding frames between two adjacent target frames in different candidate image group coding structures.

4. The video playback method according to claim 3, wherein: The determining the target GOP coding structure from at least one candidate GOP coding structure comprises: If there is only one candidate group of pictures coding structure, determining the candidate group of pictures coding structure as the target group of pictures coding structure, wherein there is only one bidirectionally predictive coding frame between two adjacent target frames in the candidate group of pictures coding structure; If there are at least two candidate GOP coding structures, the candidate GOP coding structure with the largest or smallest number of bidirectionally predictive coding frames between two adjacent target frames is determined as the target GOP coding structure.

5. The video playback method according to claim 3, wherein: The determining the target GOP coding structure from at least two candidate GOP coding structures based on the current GOP coding structure includes: determining a target number based on a current number of bidirectionally predictively coded frames existing between two adjacent target frames in a current group of pictures coding structure, wherein the target number is greater than the current number; The target GOP coding structure is determined based on the candidate number of bidirectionally predictive coding frames existing between two adjacent target frames in the candidate GOP coding structure and the target number.

6. The video playback method according to any one of claims 1 to 5, wherein: The controlling encoder encodes and outputs the currently acquired current video frame sequence based on the target image group encoding structure to obtain a current encoded frame sequence output by the encoder, including: Controlling the encoder to determine, based on the target GOP coding structure, a frame type corresponding to each current video frame in a currently acquired current video frame sequence; Encode the current video frame based on the frame type, and output the encoded frame when generating an encoded frame, to obtain a current encoded frame sequence output by the encoder; The encoder continuously outputs a forward predictive coding frame and at least one bidirectional predictive coding frame adjacent to the forward predictive coding frame.

7. A video playback method, comprising: Receiving a current coded frame sequence sent by a transmitting end, wherein the current coded frame sequence is obtained by the transmitting end controlling an encoder to encode and output a currently acquired current video frame sequence based on a target group of pictures coding structure, wherein the target group of pictures coding structure is determined based on a current delay and a preset allowed delay of the receiving end and when it is detected that a bidirectional predictive coding condition is currently satisfied, wherein at least one bidirectional predictive coding frame exists between two adjacent target frames in the target group of pictures coding structure, wherein the target frames include key frames or forward predictive coding frames; Decoding the current coded frame sequence, and adjusting the frame sequence of a first video frame sequence obtained by decoding to obtain a second video frame sequence having the same frame sequence as the current video frame sequence; Based on the video frame capture time interval in the current video frame sequence, the second video frame sequence is rendered, and the rendered third video frame sequence is played.

8. The video playback method according to claim 7, wherein: The first video frame sequence and the current coded frame sequence have the same temporal position relationship; The relative time interval between the forward predictive coded frames and the bidirectional predictive coded frames in the first video frame sequence is the same as the relative time interval between the forward predictive coded frames and the bidirectional predictive coded frames in the second video frame sequence.

9. The video playback method according to claim 7 or 8, wherein: The rendering of the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence and playing the rendered third video frame sequence includes: Adjusting, based on a video frame acquisition time interval in the current video frame sequence, a rendering time interval between two adjacent target video frames in the second video frame sequence so that the adjusted rendering time interval is the video frame acquisition time interval, wherein the target video frames include video frames corresponding to forward predictive coding frames or video frames corresponding to bidirectional predictive coding frames; Render each video frame in the adjusted second video frame sequence, and play the rendered third video frame sequence.

10. The video playback method according to any one of claims 7 to 9, wherein: Before decoding the current coded frame sequence, the method further includes: performing a decoding start time smoothing process on two adjacent forward-predicted coding frames in the current coding frame sequence so that the decoding start time intervals between the two adjacent forward-predicted coding frames after smoothing are equal; Based on the smoothed decoding start time of the forward predictive coding frame, the decoding start time of the bidirectional predictive coding frame adjacent to the forward predictive coding frame is synchronously updated to keep the temporal position relationship between the forward predictive coding frame and the bidirectional predictive coding frame before and after smoothing unchanged.

11. A video playback device, comprising: a coding structure determining module configured to, if it is detected based on a current delay of the receiving end and a preset allowed delay that a bidirectional predictive coding condition is currently satisfied, determine a target group of pictures coding structure, wherein at least one bidirectional predictive coding frame exists between two adjacent target frames in the target group of pictures coding structure, wherein the target frames include key frames or forward predictive coding frames; an encoding control module configured to control an encoder to encode and output a currently acquired current video frame sequence based on the target group of pictures encoding structure, thereby obtaining a current coded frame sequence output by the encoder, wherein bidirectionally predictive coded frames in the current coded frame sequence are output during generation; The coding frame sequence sending module is configured to send the current coding frame sequence to the receiving end so that the receiving end decodes the current coding frame sequence, adjusts the frame sequence of the first video frame sequence obtained by decoding, obtains a second video frame sequence with the same frame sequence as the current video frame sequence, and renders the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence, and plays the rendered third video frame sequence.

12. A video playback device, comprising: a coded frame sequence receiving module configured to receive a current coded frame sequence sent by a transmitting end, wherein the current coded frame sequence is obtained by the transmitting end controlling an encoder to encode and output a currently acquired current video frame sequence based on a target group of pictures coding structure, wherein the target group of pictures coding structure is determined based on a current delay and a preset allowed delay of the receiving end, when it is detected that a bidirectional predictive coding condition is currently satisfied, wherein at least one bidirectional predictive coding frame exists between two adjacent target frames in the target group of pictures coding structure, wherein the target frame includes a key frame or a forward predictive coding frame; a coded frame decoding module configured to decode the current coded frame sequence and adjust the frame sequence of the first video frame sequence obtained by decoding to obtain a second video frame sequence having the same frame sequence as the current video frame sequence; The rendering and playing module is configured to render the second video frame sequence based on the video frame acquisition time interval in the current video frame sequence, and play the rendered third video frame sequence.

13. A video playback system, comprising a transmitting end and a receiving end, wherein: The transmitting end is configured to implement the video playing method according to any one of claims 1 to 6; The receiving end is configured to implement the video playback method according to any one of claims 7 to 10.

14. An electronic device comprising: one or more processors; a storage device configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the video playback method according to any one of claims 1 to 10.

15. A storage medium containing computer-executable instructions, wherein: When executed by a computer processor, the computer executable instructions are used to perform the video playback method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Video encoding method and device, video decoding method and device, electronic equipment and storage medium

    CN112351284A

  • Video coding method and device, electronic equipment and storage medium

    CN116264622A

  • Video playing method, device, system and equipment and storage medium

    CN118138781A

  • Method and system of hardware accelerated video coding with per-frame parameter control

    US20180098083A1