Video rendering method and device and terminal equipment

By determining the second video frame based on the maximum number of consecutive B-frames and the number of the first video frame during the video frame decoding process, and rendering it when the PTS is less than other frames, the problem of low video frame rendering efficiency is solved, and more efficient and accurate video frame rendering is achieved.

CN121967780APending Publication Date: 2026-05-01BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-10-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, the problem of low video frame rendering efficiency, especially when multiple video frames to be decoded include B-frames, requires setting a large threshold for the cache to ensure the rendering order, resulting in low accuracy and flexibility of video frame rendering.

Method used

The terminal device obtains the maximum number of consecutive B-frames in the video and determines the number of the first video frames during the decoding process. Based on the maximum number of consecutive B-frames and the number of the first video frames, the second video frame is determined from the undecoded video frames. If the PTS of the third video frame in the first video frame is less than the PTS of the second video frame, the third video frame is rendered to avoid waiting for enough video frames to be buffered before rendering a frame.

Benefits of technology

It improves the rendering efficiency and accuracy of video frames, allows for flexible rendering of video frames, avoids rendering delays caused by cache threshold limitations, and enhances the flexibility and efficiency of video rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967780A_ABST
    Figure CN121967780A_ABST
Patent Text Reader

Abstract

The invention provides a video rendering method and apparatus, and a terminal device. The video rendering method comprises the steps of obtaining the number of maximum continuous B frames in a video; in the video frame decoding process of the video, the number of first video frames is obtained, and the first video frames are decoded but not rendered video frames in the video; determining a second video frame in undecoded video frames of the video according to the number of the maximum continuous B frames and the number of the first video frames; and if the video rendering timestamp PTS corresponding to a third video frame in the first video frame is smaller than the PTS of the second video frame, rendering the third video frame. And the video rendering efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Video rendering methods, devices and terminal equipment Technical Field

[0001] This disclosure relates to the field of multimedia encoding and decoding technology, and in particular to a video rendering method, apparatus, and terminal device. Background Technology

[0002] When decoding bidirectional predicted frames (B frames), decoding can be performed based on the previous and next video frames.

[0003] Currently, when multiple video frames to be decoded include B-frames, the Decoding Time Stamp (DTS) of these frames is incremental, while the Presentation Time Stamp (PTS) is out of order. To ensure proper rendering order, a fixed threshold needs to be pre-set for the buffer. Only when the number of buffered video frames exceeds this threshold can the decoder output a frame, allowing the terminal device to render it. This results in low efficiency for video frame rendering. Summary of the Invention

[0004] This disclosure provides a video rendering method, apparatus, and terminal device to address the technical problem of low efficiency in video frame rendering in the prior art.

[0005] In a first aspect, this disclosure provides a video rendering method, which includes:

[0006] The maximum number of consecutive B-frames in the video is obtained. During the video frame decoding process of the video, the number of first video frames is obtained. The first video frame is a video frame in the video that has been decoded but not rendered.

[0007] Based on the number of the maximum consecutive B-frames and the number of the first video frames, a second video frame is determined from the undecoded video frames of the video.

[0008] If the video rendering timestamp (PTS) of the third video frame in the first video frame is less than the PTS of the second video frame, then the third video frame is rendered.

[0009] Secondly, this disclosure provides a video rendering apparatus, which includes a first acquisition module, a second acquisition module, a determination module, and a rendering module, wherein:

[0010] The first acquisition module is used to acquire the maximum number of consecutive B-frames in the video;

[0011] The second acquisition module is used to acquire the number of first video frames during the video frame decoding process of the video, wherein the first video frame is a video frame in the video that has been decoded but not rendered.

[0012] The determining module is used to determine a second video frame from the undecoded video frames of the video based on the number of the maximum consecutive B-frames and the number of the first video frames.

[0013] The rendering module is used to render the third video frame if the video rendering timestamp (PTS) corresponding to the third video frame in the first video frame is less than the PTS of the second video frame.

[0014] Thirdly, this disclosure provides a terminal device including: a processor and a memory;

[0015] The memory stores computer-executed instructions;

[0016] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the video rendering methods described in the first aspect above and various possible aspects of the first aspect.

[0017] Fourthly, this disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the video rendering methods described in the first aspect above and various possible aspects of the first aspect.

[0018] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the video rendering methods described in the first aspect above and various possible aspects of the first aspect.

[0019] This disclosure provides a video rendering method, apparatus, and terminal device. The terminal device can obtain the maximum number of consecutive B-frames in a video. During video frame decoding, the terminal device can obtain the number of first video frames, which are decoded but not yet rendered video frames. Based on the maximum number of consecutive B-frames and the number of first video frames, the terminal device can determine a second video frame from the undecoded video frames. If the PTS (Power Time Segmentation) of a third video frame in the first video frame is less than the PTS of the second video frame, then the third video frame is rendered. In this method, since the PTS of the third video frame is less than the PTS of all second video frames, the terminal device's decoder can output frames, and the terminal device can render the third video frame. This eliminates the need to wait until the number of cached video frames exceeds a cache threshold before rendering, improving the efficiency and flexibility of video frame rendering. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of this disclosure;

[0022] Figure 2 is a flowchart illustrating a video rendering method provided in an embodiment of this disclosure;

[0023] Figure 3 is a schematic diagram of a video frame provided in an embodiment of this disclosure;

[0024] Figure 4 is a schematic diagram of a process of inputting video frames to a decoder according to an embodiment of the present disclosure;

[0025] Figure 5 is a schematic diagram of a process for determining a second video frame according to an embodiment of this disclosure;

[0026] Figure 6 is a schematic diagram of rendering a third video frame according to an embodiment of this disclosure;

[0027] Figure 7 is a schematic diagram of a method for rendering a third video frame provided in an embodiment of this disclosure;

[0028] Figure 8 is a structural schematic diagram of a video rendering device provided in an embodiment of this disclosure;

[0029] Figure 9 is a schematic diagram of the structure of a terminal device provided in an embodiment of this disclosure. Detailed Implementation

[0030] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0031] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0032] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0033] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0034] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0035] For ease of understanding, the concepts involved in the embodiments of this disclosure will be explained below.

[0036] Terminal device: A device with wireless transceiver capabilities. Terminal devices can be deployed on land, including indoors or outdoors, handheld, wearable, or vehicle-mounted. These terminal devices can be mobile phones, tablet computers, computers with wireless transceiver capabilities, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, vehicle-mounted terminal devices, wireless terminals in self-driving vehicles, wireless terminal devices in remote medical care, wireless terminal devices in smart grids, wireless terminal devices in transportation safety, wireless terminal devices in smart cities, wireless terminal devices in smart homes, wearable terminal devices, etc. The terminal equipment involved in the embodiments of this disclosure may also be referred to as a terminal, user equipment (UE), access terminal equipment, vehicle-mounted terminal, industrial control terminal, UE unit, UE station, mobile station, mobile station, remote station, remote terminal equipment, mobile device, UE terminal equipment, wireless communication equipment, UE agent, or UE device, etc. The terminal equipment may also be fixed or mobile.

[0037] In related technologies, when multiple video frames to be decoded include B-frames, the DTS (Direct Time Series) of the multiple video frames is incremental, while the PTS (Presentation Time Series) is out of order. For example, since B-frames need to be decoded based on the preceding and following video frames, the DTS of video frames whose PTS is after B-frame can be before the DTS of B-frame. This ensures that the decoder can guarantee the decoding of B-frames. For example, if video frame 1 is a B-frame, video frame 2 is the preceding video frame of video frame 1 (PTS order), and video frame 3 is the following video frame of video frame 1 (PTS order), the PTS order of video frames 1, 2, and 3 can be video frame 2-video frame 1-video frame 3. However, the DTS order of video frames 1, 2, and 3 is video frame 2-video frame 3-video frame 1. This ensures that when decoding video frame 1, video frames 2 and 3 have already been decoded.

[0038] However, when the video frame to be decoded can include multiple consecutive B-frames, a large threshold needs to be set in the buffer to ensure the correct PTS (Presentation Time Schedule) order of the rendered video frames. This is to prevent the decoder from failing to render frames in PTS order. For example, if the maximum number of consecutive B-frames in the video frame to be decoded is 3, then at least 4 video frames need to be buffered before the decoder can render frames. Otherwise, it may result in video frames with later PTS being rendered first, and video frames with earlier PTS being rendered later. This leads to lower accuracy and flexibility in video frame rendering.

[0039] To address the technical problems in related technologies, this disclosure provides a video rendering method. A terminal device can obtain the maximum number of consecutive B-frames in a video. During video frame decoding, it obtains the number of first video frames. The terminal device can determine a target number based on the maximum number of consecutive B-frames and the number of first video frames, and obtain the DTS of the undecoded video frames and the DTS of the target video frames. Based on the DTS order of the undecoded video frames and the DTS of the target video frames, the terminal device determines the target number of second video frames among the undecoded video frames. If the PTS of a third video frame in the first video frames is less than the PTS of the second video frames, then the third video frame is rendered. Thus, since the third video frame is the video frame with the smallest PTS among the first video frames, and the PTS of the third video frame is less than the PTS of each second video frame, the terminal device can flexibly render the third video frame, avoiding the need for video frames after B-frames to be rendered first, eliminating the need to wait for enough first video frames to be buffered before rendering. Furthermore, video frames can be rendered according to the PTS order, improving the accuracy, flexibility, and efficiency of video frame rendering.

[0040] The application scenarios of the embodiments of this disclosure will now be described with reference to Figure 1.

[0041] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of this disclosure. Referring to Figure 1, it includes: a video frame to be decoded and a decoded video frame. The maximum number of consecutive B-frames associated with the video frame to be decoded is 3 (Figure 1 shows a portion of the video frames to be decoded, including 1 B-frame). The video frames to be decoded include video frame 3, video frame 4, and video frame 5, and the decoded video frames include video frame 1 and video frame 2. The PTS of video frame 1 is pts1, and the DTS is dts1; the PTS of video frame 2 is pts3, and the DTS is dts2; the PTS of video frame 3 is pts2, and the DTS is dts3; the PTS of video frame 4 is pts4, and the DTS is dts4; the PTS of video frame 5 is pts5, and the DTS is dts5.

[0042] Please refer to Figure 1. Since the DTS of video frame 3 is DTS3 and the PTS is PTS2, and the PTS of video frame 2 is PTS3 and the DTS is DTS2, the PTS of video frame 2 is after the PTS of video frame 3, and the DTS of video frame 2 is before the DTS of video frame 3. Therefore, video frame 3 can be a B-frame. Since the PTS of video frame 1 is less than the PTS of any video frame to be decoded, the terminal device (not shown in Figure 1) can render video frame 1. In this way, the terminal device can render the video frame corresponding to the current smallest PTS, and the terminal device does not need to wait for enough video frames to be cached before rendering the video frame (in the prior art, four video frames need to be decoded before rendering can be performed), thus improving the efficiency and accuracy of rendering video frames.

[0043] It should be noted that Figure 1 is an example of an application scenario of the embodiments of this disclosure, and is not intended to limit the application scenarios of the embodiments of this disclosure.

[0044] The technical solutions of this disclosure and how they solve the aforementioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this disclosure will now be described with reference to the accompanying drawings.

[0045] Figure 2 is a flowchart illustrating a video rendering method provided in an embodiment of this disclosure. Referring to Figure 2, the method may include:

[0046] S201. Obtain the maximum number of consecutive B-frames in the video.

[0047] The execution entity of this disclosure embodiment can be a terminal device or a video rendering device installed in the terminal device. The video rendering device can be implemented in software, or it can be implemented using a combination of software and hardware.

[0048] The video can be the video to be decoded. For example, the video can include multiple video frames to be decoded. For example, if the video can include 100 video frames, the video frames to be decoded can be all 100 video frames, or it can be 50 of them; this embodiment of the application does not limit this. For example, if the video can include 10 groups of video frames, and each group of video frames includes 5 video frames, then the video frames to be decoded can be each group of video frames, or it can be all 10 groups of video frames.

[0049] Optionally, the terminal device can acquire video in any feasible way (e.g., receive video sent by other devices, or video recorded and saved by the terminal device), and this disclosure does not limit this.

[0050] The video may include keyframes (I-frames), forward prediction frames (P-frames), and bidirectional prediction frames (B-frames). I-frames can be decoded independently, P-frames are decoded based on previous I-frames or P-frames, and B-frames are decoded based on previous I-frames or P-frames and subsequent I-frames or P-frames. For example, if a video contains three consecutive B-frames, the maximum number of consecutive B-frames in the video is three. Similarly, if a video contains two consecutive B-frames and three consecutive B-frames, the maximum number of consecutive B-frames in the video is also three.

[0051] It should be noted that if a video includes B-frames but does not include consecutive B-frames, the maximum number of consecutive B-frames in the video is 1.

[0052] Optionally, the terminal device can determine the maximum number of consecutive B-frames based on the encoding information corresponding to the video. For example, the encoding information corresponding to the video may include any video-related encoding information such as the number of I-frames, the number of B-frames, the maximum number of consecutive B-frames, the position of the I-frames, and the position of the B-frames. Therefore, the terminal device can determine the maximum number of consecutive B-frames from the encoding information corresponding to the video.

[0053] The following section, in conjunction with Figure 3, provides a detailed explanation of the maximum number of consecutive B-frames.

[0054] Figure 3 is a schematic diagram of a video frame provided in an embodiment of this disclosure. Referring to Figure 3, it includes: multiple video frames. These multiple video frames can be video frames from a video to be decoded. The multiple video frames can include I-frames, P-frames, and B-frames. The multiple video frames can be ordered according to the PTS (Programmatical Time Sequence). The multiple video frames can include three consecutive B-frames and two consecutive B-frames. The terminal device can determine that the maximum number of consecutive B-frames in the multiple video frames is three.

[0055] S202. During the video frame decoding process, obtain the number of the first video frames.

[0056] The terminal device can obtain the DTS order of multiple video frames in the video and input multiple video frames into the decoder according to the DTS order. The decoder can then decode the multiple video frames. For example, the video may include video frame 1, video frame 2, and video frame 3. If the DTS of video frame 1 is less than the DTS of video frame 2, and the DTS of video frame 2 is less than the DTS of video frame 3, then the terminal device can input video frame 1 into the decoder. After video frame 1 is decoded, video frame 2 is input into the decoder, and after video frame 2 is decoded, video frame 3 is input into the decoder.

[0057] The terminal device can perform video frame decoding in the following feasible manner: identify the target video frame in the video, and decode the video frame according to the DTS corresponding to the target video frame. In this way, the terminal device decodes the video frames in the correct decoding order, improving decoding efficiency. Furthermore, users can accurately and quickly access the target video frame, improving the user experience.

[0058] The target video frame can be the video frame to be rendered. For example, a video may include video frame 1, video frame 2, and video frame 3. If the video frame to be rendered is video frame 3, then the terminal device can determine video frame 3 as the target video frame. Alternatively, the video frame to be rendered can be a video frame selected or viewed by the user. For example, the video editing track on a video editing page may include video footage, which may include 10 video frames. If the user slides the video editing track and selects the 5th video frame, then the terminal device can determine the 5th video frame as the target video frame.

[0059] It should be noted that the terminal device can determine the target video frame according to any feasible implementation method, and the embodiments disclosed herein do not limit this.

[0060] Specifically, the terminal device performs video frame decoding based on the DTS corresponding to the target video frame. This can be achieved by identifying the fourth video frame in the video and decoding both the fourth and target video frames according to the DTS order. This ensures correct decoding and improves decoding efficiency.

[0061] Optionally, the terminal device may obtain the DTS and PTS corresponding to each video frame in the video according to any feasible implementation method, and this embodiment of the disclosure does not limit this.

[0062] In this context, the DTS of the fourth video frame is less than that of the target video frame. For example, a video may include video frame 1, video frame 2, and video frame 3. The DTS of video frame 1 is less than that of video frame 2, and the DTS of video frame 2 is less than that of video frame 3. If the target video frame is video frame 2, then the fourth video frame can be video frame 1. If the target video frame is video frame 3, then the fourth video frame can be either video frame 2 or video frame 3.

[0063] Specifically, the terminal device can input the fourth video frame and the target video frame into the decoder sequentially, based on the DTS order of the fourth video frame and the target video frame. For example, if the DTS of video frame 1 is less than the DTS of video frame 2, and the DTS of video frame 2 is less than the DTS of the target video frame, then the terminal device can first input video frame 1 into the decoder, then input video frame 2 into the decoder, and finally input the target video frame into the decoder.

[0064] The process of inputting video frames to the decoder will be explained below with reference to Figure 4.

[0065] Figure 4 is a schematic diagram illustrating the process of inputting video frames to a decoder according to an embodiment of this disclosure. Referring to Figure 4, the video to be decoded includes video frame 1, video frame 2, video frame 3, and video frame 4. The PTS of video frame 1 is pts1, and the DTS is dts1; the PTS of video frame 2 is pts3, and the DTS is dts2; the PTS of video frame 3 is pts2, and the DTS is dts3; and the PTS of video frame 4 is pts4, and the DTS is dts4.

[0066] Referring to Figure 4, the target video frame is video frame 3. Since the DTS of video frame 1 and video frame 2 is shorter than that of video frame 3, the terminal device (not shown in Figure 4) can identify video frames 1 and 2 as the third video frame. The terminal device can then input video frames 1, 2, and 3 sequentially to the decoder. This allows the terminal device to decode the target video frame, as well as the video frames preceding it (in DTS order), improving the decoding efficiency of the target video frame.

[0067] In this context, the first video frame refers to a decoded but unrendered video frame. For example, the first video frame can be a Decoded Picture Buffer (DPB) frame. Optionally, during the video decoding process, the terminal device can store the decoded but unrendered first video frames in a buffer and determine the number of first video frames in the buffer. For example, if the buffer contains 2 video frames, the terminal device can determine that the number of first video frames is 2; if the buffer contains 4 video frames, the terminal device can determine that the number of first video frames is 4.

[0068] It should be noted that the number of the first video frames can be the number of video frames that the decoder has decoded but not yet output. For example, if the decoder has decoded 10 video frames and output 8 video frames (the terminal device renders 8 video frames), then there are 2 video frames remaining in the buffer, that is, the number of the first video frames is 2.

[0069] It should be noted that the terminal device can acquire the first video frame in any feasible way, and this embodiment of the present disclosure does not limit this.

[0070] S203. Based on the number of the maximum consecutive B-frames and the number of the first video frames, determine the second video frame from the undecoded video frames of the video.

[0071] The second video frame can be an undecoded video frame from the video. For example, the terminal device can determine L (an integer greater than 0) based on the maximum number of consecutive B-frames and the number of the first video frames, and then select L consecutive video frames from the undecoded video frames to obtain the second video frame. For example, if the undecoded video frames are arranged in DTS order as video frame 1, video frame 2, and video frame 3, and L is 1, and the target video frame is video frame 2, then the terminal device can select one video frame after the target video frame to obtain the second video frame; that is, the second video frame can be video frame 3.

[0072] The terminal device can determine the second video frame according to the following feasible implementation method: determine the target number based on the maximum number of consecutive B-frames and the number of first video frames, obtain the DTS of the undecoded video frame and the DTS of the target video frame, and determine the target number of second video frames in the undecoded video frames based on the DTS order of the undecoded video frames and the DTS of the target video frame.

[0073] The target quantity can satisfy the following formula:

[0074] A = M + 1 - N

[0075] Where A can be the target number, M can be the maximum number of consecutive B frames, and N can be the number of the first video frame.

[0076] Optionally, the terminal device can determine the DTS of each undecoded video frame to obtain the DTS order corresponding to multiple undecoded video frames. The terminal device can also determine the DTS order corresponding to multiple undecoded video frames according to any feasible implementation method. This disclosure does not limit this.

[0077] Optionally, the terminal device can select a target number of undecoded video frames after the target video frame (in DTS order) based on the DTS order of the undecoded video frames and the DTS order of the target video frame to obtain the second video frame. For example, if there are 10 undecoded video frames after the target video frame, arranged in DTS order, and the target number is 2, the terminal device can determine the first and second undecoded video frames as the second video frame.

[0078] The process of determining the second video frame will be explained below with reference to Figure 5.

[0079] Figure 5 is a schematic diagram of a process for determining a second video frame according to an embodiment of this disclosure. Referring to Figure 5, it includes: undecoded video frames. The undecoded video frames include video frame 1, video frame 2, video frame 3, and video frame 4. The PTS of video frame 1 is pts1, and the DTS is dts1; the PTS of video frame 2 is pts2, and the DTS is dts3; the PTS of video frame 3 is pts3, and the DTS is dts2; and the PTS of video frame 4 is pts4.

[0080] Please refer to Figure 5. If the target quantity is 1 and the target video frame is video frame 3, the terminal device can select one video frame after video frame 3 to obtain the second video frame. Since the DTS of video frame 2 is the same as that of video frame 3, the terminal device can determine video frame 2 as the second video frame. This ensures that the terminal device renders video frames according to the PTS order, improving the accuracy of video frame rendering.

[0081] S204. If the PTS of the third video frame in the first video frame is less than the PTS of the second video frame, then render the third video frame.

[0082] In this case, the PTS corresponding to the third video frame is less than the PTS corresponding to any of the second video frames. For example, the first video frame includes video frame 1, video frame 2, and video frame 3, and the second video frame includes video frame 4 and video frame 5. If the PTS corresponding to video frame 1 is less than the PTS corresponding to video frame 5 and the PTS corresponding to video frame 5, then the terminal device can determine that video frame 1 is the third video frame. If the PTS of video frame 1, video frame 2, and video frame 3 are all greater than the PTS of video frame 2, then the terminal device can determine that there is no third video frame in the first video frame.

[0083] Optionally, the third video frame can be the video frame with the smallest PTS among the first video frames. For example, the first video frame includes video frame 1 and video frame 2, and the second video frame includes video frame 3. If the PTS corresponding to video frame 1 is less than the PTS corresponding to video frame 2, the terminal device can determine the size between the PTS corresponding to video frame 1 and the PTS corresponding to video frame 3. If the PTS corresponding to video frame 1 is less than the PTS corresponding to video frame 3, the terminal device can determine video frame 1 as the third video frame. That is, the third video frame is the video frame with the smallest PTS among the first video frames, and the PTS corresponding to the third video frame is less than the PTS corresponding to any of the second video frames.

[0084] If the third video frame is the video frame with the smallest PTS (Power Segment Size) among the first video frames, then the terminal device can render the third video frame. For example, if the first video frame includes video frame 1 and video frame 2, and video frame 1 is the third video frame, then the decoder can output video frame 1, and the terminal device can render video frame 1. If video frame 2 is the third video frame, then the decoder can output video frame 2, and the terminal device can render video frame 2.

[0085] Optionally, the terminal device may also render the third video frame according to the following feasible implementation: obtain the number of the third video frames; if the number of the third video frames is 1, then render the third video frame; if the number of the third video frames is greater than 1, then render the third video frames according to the order of PTS.

[0086] The number of third video frames can be 1 or greater than 1. For example, the first video frame includes video frame 1 and video frame 2, and the second video frame includes video frame 3. If the PTS corresponding to video frame 1 is less than the PTS corresponding to video frame 3, and the PTS corresponding to video frame 2 is greater than the PTS corresponding to video frame 3, then video frame 1 can be the third video frame, that is, the number of third video frames is 1. If the PTS corresponding to video frame 1 is less than the PTS corresponding to video frame 3, and the PTS corresponding to video frame 2 is less than the PTS corresponding to video frame 3, then video frames 1 and 2 can be the third video frames, that is, the number of third video frames is 2.

[0087] Optionally, if the number of third video frames is 1, it means that there is 1 video frame to be rendered (the third video frame) in the first video frame. The PTS corresponding to the video frame to be rendered is the smallest PTS in the first video frame, and the PTS corresponding to the video frame to be rendered is less than the PTS corresponding to any second video frame. Therefore, the terminal device can render the video frame to be rendered to ensure the correct rendering order of the video frames.

[0088] Optionally, if the number of third video frames is greater than 1, it means that there are multiple video frames to be rendered in the first video frame. Therefore, the terminal device can determine the video frame to be rendered with the smallest PTS among the multiple video frames to be rendered and render the video frame to be rendered with the smallest PTS. Alternatively, it can render the third video frame in order of PTS from smallest to largest, thereby ensuring the correct rendering order of video frames.

[0089] The process of rendering the third video frame will now be explained with reference to Figure 6.

[0090] Figure 6 is a schematic diagram of rendering a third video frame according to an embodiment of this disclosure. Referring to Figure 6, it includes: a first video frame and a second video frame. The first video frame includes video frame 1 and video frame 2, and the second video frame includes video frame 3 and video frame 4. The PTS of video frame 1 is PTS3, and the DTS is DTS4; the PTS of video frame 2 is PTS4, and the DTS is DTS3; the PTS of video frame 3 is PTS5, and the DTS is DTS5; the PTS of video frame 6 is PTS6, and the DTS is DTS6.

[0091] Referring to Figure 6, since the PTS3 corresponding to video frame 1 is less than the PTS4 corresponding to video frame 2, and also less than the PTS5 corresponding to video frame 3, and less than the PTS6 corresponding to video frame 4, the terminal device (not shown in Figure 6) can determine video frame 1 as the third video frame and render video frame 1. In this way, the terminal device can flexibly render the video frame with the smallest PTS, ensuring the accuracy of video frame rendering. Furthermore, the terminal device does not need to wait for the number of first video frames to exceed a preset threshold before rendering video frames, thus improving the efficiency of video frame rendering.

[0092] It should be noted that the terminal device can control the output of one third video frame from the buffer and render the third video frame based on any feasible implementation method. This application embodiment does not limit this.

[0093] This disclosure provides a video rendering method. A terminal device can obtain the maximum number of consecutive B-frames in a video and determine a target video frame. The terminal device can also determine a fourth video frame in the video whose DTS is less than the DTS of the target video frame. Based on the DTS order, the fourth video frame and the target video frame are input into the decoder. The terminal device can determine the target number based on the maximum number of consecutive B-frames and the number of first video frames, and obtain the DTS of the undecoded video frames and the target video frame. Based on the DTS order of the undecoded video frames and the DTS of the target video frames, the terminal device can determine the target number of second video frames in the undecoded video frames. If the video rendering timestamp (PTS) of a third video frame in the first video frames is less than the PTS of any second video frame, then the third video frame is rendered. Thus, since there is a third video frame in the first video frames whose PTS is less than the PTS of any second video frame, the terminal device can render the third video frame, improving the accuracy, flexibility, and efficiency of video frame rendering.

[0094] Based on any of the above embodiments, after the terminal device determines the second video frame, the above video rendering method further includes another method for rendering the third video frame. The method for rendering the third video frame in the above video rendering method will be described below with reference to Figure 7.

[0095] Figure 7 is a schematic diagram of a method for rendering a third video frame according to an embodiment of this disclosure. Referring to Figure 7, the method flow includes:

[0096] S701. Determine whether the first video frame includes the third video frame.

[0097] If so, then execute S702.

[0098] If not, then execute S703.

[0099] S702, render the third video frame.

[0100] Optionally, the third video frame may be the video frame with the smallest PTS among the first video frames.

[0101] Optionally, if the number of third video frames is 1, the terminal device may render the third video frame.

[0102] Optionally, if the number of third video frames is greater than 1, the terminal device can render the third video frames according to the order of the PTS.

[0103] S703. Decode the undecoded video frames sequentially according to the DTS order of the undecoded video frames until the first video frame includes the third video frame, then render the third video frame.

[0104] If the first video frame does not include the third video frame, the terminal device can determine the video frame with the smallest DTS (the undecoded video frame may include the target video frame and the second video frame) from the undecoded video frames, and input the video frame with the smallest DTS to the decoder. The decoder can decode and buffer the video frame with the smallest DTS. The terminal device can then determine whether the first video frame includes the third video frame. If it does, the third video frame is rendered. If it does not, the above steps are repeated until the first video frame includes the third video frame, at which point the terminal device can render the third video frame.

[0105] The video rendering method described above will be explained below with specific examples.

[0106] Step 1: The terminal device can obtain the maximum number of consecutive B-frames in the video, denoted as M, where M is an integer greater than or equal to 1.

[0107] Step 2: The terminal device sequentially inputs the target video frame and the undecoded video frame with a smaller DTS than the target video frame into the decoder according to the DTS increment relationship. The terminal device records the number of the first video frame as N and the first video frame with the smallest PTS as P1, where N is an integer greater than or equal to 0.

[0108] Step 3: The terminal device can continuously acquire undecoded video frames after the target video frame according to the DTS order and store them in the queue of undecoded video frames until the number of undecoded video frames in the queue is M+1-N.

[0109] Step 4: Determine the PTS of the undecoded video frames in the queue of undecoded video frames, and determine the undecoded video frame with the smallest PTS in the queue of undecoded video frames, denoted as P2.

[0110] Step 5: If P1 is greater than P2, input the head of the queue of undecoded video frames (the first frame arranged in DTS order) to the decoder, and repeat steps 2-5 until P1 is less than P2, that is, the first video frame includes the third video frame.

[0111] Step 6: Force the decoder to output a video frame (the third video frame) without disrupting the decoder's subsequent decoding.

[0112] Step 7: Repeat step 5 to ensure the correctness of the decoding.

[0113] This disclosure provides a method for rendering a third video frame. A terminal device determines whether a first video frame includes a third video frame. If so, the third video frame is rendered; otherwise, the undecoded video frames are decoded sequentially according to their DTS order until the first video frame includes the third video frame, at which point the third video frame is rendered. This way, if the first video frame includes the third video frame, the terminal device can render that third video frame; if the first video frame does not include the third video frame, the terminal device can continue to input undecoded video frames into the decoder until the first video frame includes the third video frame, at which point the third video frame is rendered. The terminal device does not need to wait for the number of first video frames to exceed a preset threshold before rendering video frames, thus improving the efficiency of video frame rendering.

[0114] Figure 8 is a schematic diagram of a video rendering apparatus provided in an embodiment of this disclosure. Referring to Figure 8, the video rendering apparatus 800 includes a first acquisition module 801, a second acquisition module 802, a determination module 803, and a rendering module 804, wherein:

[0115] The first acquisition module 801 is used to acquire the maximum number of consecutive B-frames in the video;

[0116] The second acquisition module 802 is used to acquire the number of first video frames during the video frame decoding process of the video, wherein the first video frame is a video frame in the video that has been decoded but not rendered.

[0117] The determining module 803 is used to determine a second video frame from the undecoded video frames of the video based on the number of the maximum consecutive B-frames and the number of the first video frames;

[0118] The rendering module 804 is used to render the third video frame if the video rendering timestamp (PTS) corresponding to the third video frame in the first video frame is less than the PTS of the second video frame.

[0119] According to one or more embodiments of this disclosure, the third video frame is the video frame with the smallest PTS among the first video frames.

[0120] According to one or more embodiments of this disclosure, the rendering module 804 is specifically used for:

[0121] Obtain the number of the third video frames;

[0122] If the number of the third video frames is 1, then render the third video frame;

[0123] If the number of the third video frames is greater than 1, then the third video frames are rendered according to the order of PTS.

[0124] According to one or more embodiments of this disclosure, the second acquisition module 802 is specifically used for:

[0125] Identify the target video frame in the video;

[0126] Based on the Video Encoding Timestamp (DTS) corresponding to the target video frame, the video is decoded.

[0127] According to one or more embodiments of this disclosure, the second acquisition module 802 is specifically used for:

[0128] A fourth video frame is determined in the video, wherein the DTS of the fourth video frame is less than the DTS of the target video frame;

[0129] The fourth video frame and the target video frame are decoded according to the DTS sequence.

[0130] According to one or more embodiments of this disclosure, the determining module 803 is specifically used for:

[0131] The target number is determined based on the number of the maximum consecutive B-frames and the number of the first video frames;

[0132] Obtain the DTS of the undecoded video frame and the DTS of the target video frame;

[0133] Based on the DTS order of the undecoded video frames and the DTS of the target video frames, a target number of second video frames are determined in the undecoded video frames.

[0134] According to one or more embodiments of this disclosure, the rendering module 804 is further configured to:

[0135] If the first video frame does not include the third video frame, then the undecoded video frames are decoded sequentially according to the DTS order of the undecoded video frames until the first video frame includes the third video frame, at which point the third video frame is rendered.

[0136] The video rendering apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0137] Figure 9 is a schematic diagram of the structure of a terminal device provided in an embodiment of this disclosure. Referring to Figure 9, it shows a schematic diagram of the structure of a terminal device 900 suitable for implementing an embodiment of this disclosure. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The terminal device shown in Figure 9 is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this disclosure.

[0138] As shown in Figure 9, the terminal device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 902 or a program loaded from storage device 908 into random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the terminal device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0139] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows terminal device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 shows a terminal device 900 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0140] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.

[0141] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0142] The aforementioned computer-readable medium may be included in the aforementioned terminal device; or it may exist independently and not assembled into the terminal device.

[0143] The aforementioned computer-readable medium carries one or more programs, which, when executed by the terminal device, cause the terminal device to perform the method shown in the above embodiments.

[0144] This disclosure provides a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, it implements various methods that may be involved in the above embodiments.

[0145] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements various methods that may be involved in the above embodiments.

[0146] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0148] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0149] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), etc.

[0150] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0151] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0152] In a first aspect, embodiments of this disclosure provide a video rendering method, the video rendering method comprising:

[0153] Get the maximum number of consecutive B-frames in the first video frame;

[0154] The first video frame is decoded, and the number of first video frames buffered after decoding is obtained.

[0155] Based on the maximum number of consecutive B-frames and the number of first video frames in the decoded buffer, the second video frame is determined from the undecoded first video frames;

[0156] If a target PTS exists in the video rendering timestamp PTS corresponding to the first video frame in the decoded buffer, then the first video frame in the decoded buffer is rendered, wherein the target PTS is less than the PTS corresponding to any second video frame.

[0157] According to one or more embodiments of this disclosure, the target PTS is the smallest PTS among the PTS corresponding to the first video frame in the decoded buffer.

[0158] According to one or more embodiments of this disclosure, rendering the first video frame of the decoded buffer includes:

[0159] Obtain the number of the target PTS;

[0160] If the number of target PTSs is 1, then render the first video frame of the decoded buffer corresponding to the target PTS;

[0161] If the number of target PTSs is greater than 1, then render the first video frame of the decoded buffer corresponding to the smallest target PTS.

[0162] According to one or more embodiments of this disclosure, the decoding process for the first video frame includes:

[0163] Determine the target video frame within the first video frame;

[0164] The first video frame is decoded based on the Video Encoding Timestamp (DTS) corresponding to the target video frame.

[0165] According to one or more embodiments of this disclosure, decoding the first video frame based on the Video Coding Timestamp (DTS) of the target video frame includes:

[0166] In the first video frame, a third video frame with a DTS smaller than the DTS of the target video frame is determined;

[0167] According to the DTS sequence, the third video frame and the target video frame are input into the decoder.

[0168] According to one or more embodiments of this disclosure, determining the second video frame from the undecoded first video frames based on the maximum number of consecutive B-frames and the number of first video frames in the decoded buffer includes:

[0169] The target number is determined based on the maximum number of consecutive B-frames and the number of the first video frames in the decoded buffer.

[0170] Based on the DTS order corresponding to the undecoded first video frame and the DTS corresponding to the target video frame, a target number of second video frames are determined in the undecoded first video frame.

[0171] According to one or more embodiments of this disclosure, the method further includes:

[0172] If the target PTS is not present in the PTS corresponding to the first video frame in the decoded buffer, then the first undecoded video frame with the smallest DTS is input to the decoder until the target PTS is present in the PTS corresponding to the first video frame in the decoded buffer.

[0173] Secondly, embodiments of this disclosure provide a video rendering apparatus, wherein the video rendering apparatus includes a first acquisition module, a decoding module, a second acquisition module, a determination module, and a rendering module, wherein:

[0174] The first acquisition module is used to acquire the maximum number of consecutive B-frames in the first video frame;

[0175] The decoding module is used to decode the first video frame;

[0176] The second acquisition module is used to acquire the number of the first video frames buffered after decoding;

[0177] The determining module is used to determine the second video frame from the undecoded first video frames based on the number of the maximum consecutive B-frames and the number of the first video frames in the decoded buffer.

[0178] The rendering module is used to render the first video frame after decoding if a target PTS exists in the video rendering timestamp PTS corresponding to the first video frame in the decoded buffer, wherein the target PTS is less than the PTS corresponding to any second video frame.

[0179] According to one or more embodiments of this disclosure, the target PTS is the smallest PTS among the PTS corresponding to the first video frame in the decoded buffer.

[0180] According to one or more embodiments of this disclosure, the rendering module specifically uses:

[0181] Obtain the number of the target PTS;

[0182] If the number of target PTSs is 1, then render the first video frame of the decoded buffer corresponding to the target PTS;

[0183] If the number of target PTSs is greater than 1, then render the first video frame of the decoded buffer corresponding to the smallest target PTS.

[0184] According to one or more embodiments of this disclosure, the decoding module is specifically used for:

[0185] Determine the target video frame within the first video frame;

[0186] The first video frame is decoded based on the Video Encoding Timestamp (DTS) corresponding to the target video frame.

[0187] According to one or more embodiments of this disclosure, the decoding module is specifically used for:

[0188] In the first video frame, a third video frame with a DTS smaller than the DTS of the target video frame is determined;

[0189] According to the DTS sequence, the third video frame and the target video frame are input into the decoder.

[0190] According to one or more embodiments of this disclosure, the determining module is specifically used for:

[0191] The target number is determined based on the maximum number of consecutive B-frames and the number of the first video frames in the decoded buffer.

[0192] Based on the DTS order corresponding to the undecoded first video frame and the DTS corresponding to the target video frame, a target number of second video frames are determined in the undecoded first video frame.

[0193] According to one or more embodiments of this disclosure, the video rendering apparatus further includes a sending module, wherein the input module is used for:

[0194] If the target PTS is not present in the PTS corresponding to the first video frame in the decoded buffer, then the first undecoded video frame with the smallest DTS is input to the decoder until the target PTS is present in the PTS corresponding to the first video frame in the decoded buffer.

[0195] Thirdly, this disclosure provides a terminal device including: a processor and a memory;

[0196] The memory stores computer-executed instructions;

[0197] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the video rendering methods described in the first aspect above and various possible aspects of the first aspect.

[0198] Fourthly, this disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the video rendering methods described in the first aspect above and various possible aspects of the first aspect.

[0199] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the video rendering methods described in the first aspect above and various possible aspects of the first aspect.

[0200] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0201] It is understood that the data involved in this technical solution (including but not limited to the data itself, its acquisition, or its use) shall comply with the requirements of relevant laws, regulations, and provisions. Data may include information, parameters, and messages, such as flow control instructions.

[0202] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0203] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely exemplary forms of implementing the claims.

Claims

1. A video rendering method, characterized in that, include: The maximum number of consecutive B-frames in the video is obtained. During the video frame decoding process, the number of first video frames is obtained, where the first video frames are the decoded but unrendered video frames in the video. Based on the maximum number of consecutive B-frames and the number of first video frames, a second video frame is determined from the undecoded video frames in the video. If the video rendering timestamp (PTS) corresponding to the third video frame in the first video frame is less than the PTS of the second video frame, then the third video frame is rendered.

2. The method according to claim 1, characterized in that, The third video frame is the video frame with the smallest PTS among the first video frames.

3. The method according to claim 1, characterized in that, Rendering the third video frame includes: obtaining the number of the third video frames; if the number of the third video frames is 1, then rendering the third video frame; if the number of the third video frames is greater than 1, then rendering the third video frames according to the order of PTS.

4. The method according to any one of claims 1-3, characterized in that, The step of decoding the video frames includes: determining a target video frame in the video; and decoding the video frames according to the Video Encoding Timestamp (DTS) corresponding to the target video frame.

5. The method according to claim 4, characterized in that, Based on the Video Encoded Timestamp (DTS) of the target video frame, video frame decoding is performed, including: determining a fourth video frame in the video, wherein the DTS of the fourth video frame is less than the DTS of the target video frame; and decoding the fourth video frame and the target video frame according to the order of the DTS.

6. The method according to any one of claims 1-3, characterized in that, The step of determining the second video frames from the undecoded video frames of the video based on the maximum number of consecutive B-frames and the number of the first video frames includes: determining a target number based on the maximum number of consecutive B-frames and the number of the first video frames; obtaining the DTS of the undecoded video frames and the DTS of the target video frames; and determining the target number of second video frames from the undecoded video frames based on the DTS order of the undecoded video frames and the DTS of the target video frames.

7. The method according to any one of claims 1-3, characterized in that, The method further includes: if the first video frame does not include the third video frame, then according to the DTS order of the undecoded video frames, the undecoded video frames are decoded sequentially until the first video frame includes the third video frame, and then the third video frame is rendered.

8. A video rendering apparatus, characterized in that, The system includes a first acquisition module, a second acquisition module, a determination module, and a rendering module, wherein: the first acquisition module is used to acquire the maximum number of consecutive B-frames in the video; the second acquisition module is used to acquire the number of first video frames during video frame decoding of the video, wherein the first video frames are decoded but not rendered video frames in the video; the determination module is used to determine a second video frame from the undecoded video frames of the video based on the maximum number of consecutive B-frames and the number of first video frames; and the rendering module is used to render the third video frame if the video rendering timestamp (PTS) corresponding to a third video frame in the first video frame is less than the PTS of the second video frame.

9. A terminal device, characterized in that, include: Processor and memory; The memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, causing the processor to perform the video rendering method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the video rendering method as described in any one of claims 1-7.

11. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the video rendering method as described in any one of claims 1-7.