Video playback method, video recording method, video conversion method, and electronic device

By evaluating and selecting optimized candidate paths, the stuttering problem caused by uneven decoding time during video playback was resolved, improving the smoothness of video playback and decoding efficiency, and optimizing the user experience.

WO2026016544A1PCT designated stage Publication Date: 2026-01-22HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/087504
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-18
Filing Date
2025-04-07
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

In existing video playback technologies, the different frame distances between the target display frame and the I-frame during fixed-speed and variable-speed playback result in uneven decoding time, causing video playback to fluctuate and stutter.

Method used

By evaluating the video playback experience and decoding cost of multiple candidate paths, the candidate path with the better video playback experience is selected as the target path. The process of reconstructing the image of the target display frame is optimized, including the decoding path and memory reading path, thereby reducing the decoding cost.

Benefits of technology

It improves the smoothness of video playback and decoding efficiency in both fixed-speed and variable-speed playback scenarios, reduces playback stuttering, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025087504_22012026_PF_FP_ABST
    Figure CN2025087504_22012026_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a video playback method, a video recording method, a video conversion method, and an electronic device. The video playback method comprises: acquiring a video file, wherein the video file comprises a bitstream, and the bitstream comprises encoded data obtained by encoding original images of a plurality of image frames in video data; determining a target display frame from among the plurality of image frames; on the basis of evaluation information corresponding to each candidate path among a plurality of candidate paths, selecting one candidate path from among the plurality of candidate paths as a target path, wherein the evaluation information is used for describing video playback experience; on the basis of the target path and the bitstream, acquiring a reconstructed image of the target display frame; and playing the reconstructed image of the target display frame. In this way, superior video playback experience can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Video playing, video recording and video converting method and electronic device TECHNICAL FIELD

[0001] The present application relates to the field of multimedia, and in particular to a video playing, video recording and video converting method and electronic device. BACKGROUND

[0002] Video fixed-speed playing can refer to adjusting the playing speed by a fixed multiple, such as 1 times, 1.5 times, 2 times, etc. This playing mode is simple and intuitive, and is often used for quickly browsing content, such as quickly watching tutorials or lectures on a video platform, etc. Usually, the video platform will provide several preset speed options, and the user can select a suitable speed option according to his own needs.

[0003] Video variable-speed playing can refer to adjusting the playing speed by a non-fixed ratio; for example, according to user operations (such as dragging the progress bar), the playing speed (such as 1.2 times, 1.3 times, etc.) is adjusted, instead of being able to only select fixed speeds such as 1 times, 2 times, etc. This playing mode is often used in situations where it is necessary to quickly locate interesting content in a video, such as learning a language or analyzing audio / video content in detail, etc.

[0004] Usually, in the process of video playing (including video fixed-speed playing and video variable-speed playing), a plurality of target display frames are sequentially determined from all image frames contained in a video file, and images of the plurality of target display frames are sequentially displayed. In the prior art, the reconstructed image of each target display frame is obtained by decoding image frames in a Group Of Pictures (GOP) from an I frame (i.e. an Intra-coded Frame) to the target display frame. It can be seen that the number of image frames that need to be decoded is positively correlated with the frame distance between the target display frame and the I frame, i.e. the farther the frame distance between the target display frame and the I frame, the more the number of image frames that need to be decoded, and the longer the decoding time; the closer the frame distance between the target display frame and the I frame, the fewer the number of image frames that need to be decoded, and the shorter the decoding time. There will be a part of the plurality of target display frames, and the frame distance between this part of the target display frames and the I frame in the respective GOPs is different, so the time taken to decode each target display frame in this part of the target display frames is different; in this way, the time of the final display of this part of the target display frames is unevenly distributed in time sequence, resulting in that the user will feel that the video playing is fast and slow. In addition, when the time taken to decode a target display frame is relatively long, the user will also feel obvious lag. SUMMARY

[0005] Therefore, the application provides a video playing method, a video recording method, a video converting method and an electronic device. In a fixed speed playing scene, a variable speed playing scene or an original speed playing scene, the video playing method provided by the application can obtain better video playing experience (such as video fluency).

[0006] In a first aspect, the application provides a video playing method, which comprises the following steps: first, obtaining a video file, wherein the video file comprises a code stream, and the code stream comprises encoded data of original images of a plurality of image frames in encoded video data; then, determining a target display frame from the plurality of image frames; subsequently, selecting a candidate path from a plurality of candidate paths as a target path based on evaluation information corresponding to each candidate path in the plurality of candidate paths, wherein the evaluation information is used to describe video playing experience; then, obtaining a reconstructed image of the target display frame based on the target path and the code stream; and playing the reconstructed image of the target display frame.

[0007] That is, the application takes video playing experience as an index (or basis) for measuring the advantages and disadvantages of the plurality of candidate paths, so as to select a target path from the plurality of candidate paths; wherein a candidate path with better video playing experience can be selected as the target path from the plurality of candidate paths. Video playing experience can refer to the overall feeling of a user when watching video content, including video fluency and the like; thus, in a fixed speed playing scene, a variable speed playing scene or an original speed playing scene, the video playing method provided by the application can obtain better video playing experience (such as video fluency and the like).

[0008] Exemplarily, the image frame (Image Frame) in the application can also be referred to as a frame (Frame).

[0009] Exemplarily, the code stream can also be referred to as a bit stream / bitstream, an encoded bit stream, and the like.

[0010] Exemplarily, a candidate path can be selected as the target path from the first M candidate paths with the best video playing experience based on the evaluation information corresponding to each candidate path. Wherein M is a positive integer, and when M is 1, the candidate path with the best video playing experience is selected as the target path.

[0011] Exemplarily, the evaluation information can be used to measure video playing experience, and can be a quantitative representation of video playing experience.

[0012] Exemplarily, “playing the reconstructed image of the target display frame” can also be understood as “displaying the reconstructed image of the target display frame”.

[0013] Exemplarily, the video playing method of the present application can be executed by a video player.

[0014] Exemplarily, the video playing method can further comprise: determining a plurality of candidate paths for obtaining the reconstructed image of the target display frame; and determining evaluation information corresponding to each candidate path in the plurality of candidate paths. Then, based on the evaluation information corresponding to each candidate path in the plurality of candidate paths, one candidate path is selected from the plurality of candidate paths as the target path.

[0015] In one possible manner, the evaluation information of each candidate path can be determined after the candidate path is determined; then in the process of determining the target path, it can be first judged whether the evaluation information (such as evaluation value) of the candidate path is less than the evaluation threshold; when the evaluation information of the candidate path is less than the evaluation threshold, the next candidate path can be determined; when the evaluation information of the candidate path is greater than or equal to the evaluation threshold, the candidate path can be determined as the target path, and the cycle is executed.

[0016] In one possible manner, a plurality of candidate paths for obtaining the reconstructed image of the target display frame can be determined within a certain time length (which can be set in advance according to requirements); then the target path is determined from the plurality of candidate paths.

[0017] In one possible manner, G (G can be set in advance according to requirements, and G is a positive integer) candidate paths for obtaining the reconstructed image of the target display frame can be determined; then the target path is determined from the plurality of candidate paths.

[0018] According to the first aspect, at least one candidate path in the plurality of candidate paths is different from the decoding cost of other candidate paths.

[0019] For example, the decoding cost of any two candidate paths in the plurality of candidate paths is different, that is, the decoding cost of the plurality of candidate paths is different.

[0020] Exemplarily, the decoding cost can refer to the number of decoded image frames or decoding time required to obtain the reconstructed image of the target display frame.

[0021] In one possible manner, the decoding cost of any two candidate paths in the plurality of candidate paths is the same, that is, the decoding cost of the plurality of candidate paths is the same.

[0022] That is, the application first determines a plurality of candidate paths with the same or different decoding costs, and then selects a target path according to the video playback experience (such as the smoothness of the video) of each candidate path. The smoothness of the video is associated with the decoding cost; when the decoding cost is small, the smoothness of the video is better; when the decoding cost is large, the smoothness of the video is poor. It can be seen that the application selects a candidate path with better video playback experience as the target path, which in essence also selects a candidate path with smaller decoding cost; therefore, the application can also reduce the decoding cost.

[0023] It should be noted that the plurality of decoding costs belonging to the same numerical range may correspond to the same smoothness of the video; because the human eye can only feel the stall when the interval between two adjacent frames exceeds a certain length of time; and when performing a drag operation or a flick operation, the human eye also needs a certain reaction time to judge whether the picture is hand-following. Therefore, when the candidate path with the optimal video playback experience is determined to be multiple based on the evaluation information corresponding to each candidate path, a candidate path with the smallest decoding cost can be selected from the multiple candidate paths with the optimal video playback experience to determine the target path.

[0024] According to the first aspect, or any one of the implementations of the first aspect, the candidate path includes at least one of a first type of path, a second type of path, or a third type of path;

[0025] The first type of path is used to indicate one or more image frames required for decoding to obtain a reconstructed image of the target display frame;

[0026] The second type of path is used to indicate reading the reconstructed image of the target display frame from the memory within the decoder;

[0027] The third type of path is used to indicate reading the reconstructed image of the target display frame from the memory outside the decoder.

[0028] Exemplarily, the first type of path can also be referred to as a decoding path. The first type of path can include the frame identifier of one or more image frames required for decoding to obtain the reconstructed image of the target display frame (that is, the frame identifier of the image frame to be decoded). When there are multiple candidate paths belonging to the first type of path, the image frames to be decoded corresponding to the multiple candidate paths are different.

[0029] Exemplarily, the second type of path can include the frame identifier of a decoded image frame in the memory within the decoder. When there are multiple candidate paths belonging to the second type of path, the decoded image frames corresponding to the multiple candidate paths are different.

[0030] Exemplarily, the third type of path can comprise a frame identifier of a to-be-displayed image frame in a memory out of the decoder. When there are multiple candidate paths belonging to the third type of path, the to-be-displayed image frames corresponding to the multiple candidate paths are different. For example, the memory out of the decoder can comprise a display memory.

[0031] Exemplarily, the decoding cost of the second type of path and the third type of path is smaller than that of the first type of path.

[0032] It should be understood that, in some cases, the second type of path and the third type of path can be classified into the same type of path, which is used to indicate a reconstructed image of the target display frame read from the memory.

[0033] According to the first aspect, or any one of the implementations of the first aspect, the first type of path comprises a first decoding path, and when one of the multiple candidate paths is the first decoding path, the method further comprises: obtaining, based on the video file, decoding dependency information of each image frame in a group of pictures to which the target display frame belongs; and determining the first decoding path based on the decoding dependency information of each image frame in the group of pictures to which the target display frame belongs.

[0034] The decoding dependency information of each image frame can be used to indicate a number of image frames after the image frame in the group of pictures to which the image frame belongs and which are dependent on the image frame for decoding. Alternatively, the decoding dependency information of each image frame can be used to indicate a decoding frame dropping strategy of the image frame, and the decoding frame dropping strategy of each image frame can refer to image frames to be dropped after the image frame. Further, based on the decoding dependency information of each image frame in the group of pictures to which the target display frame belongs, the discardable image frames and the non-discardable image frames can be determined, and then the first decoding path can be generated based on the non-discardable image frames. Compared with the prior art which needs to decode all image frames before the target display frame in the group of pictures to which the target display frame belongs, in most cases, the number of image frames to be decoded to obtain the target display frame according to the first decoding path is smaller.

[0035] Exemplarily, “dependent” can be understood as “directly referenced” or “indirectly referenced”.

[0036] According to the first aspect, or any one of the implementations of the first aspect, the first decoding path is determined based on the decoding dependency information of each image frame in the image group to which the target display frame belongs, including: determining one or more first image frames based on the decoding dependency information of each image frame in the image group to which the target display frame belongs, the first image frames being image frames on which the target display frame depends for decoding; and generating the first decoding path using the frame identifiers of the one or more first image frames and the frame identifier of the target display frame. That is, the first decoding path is generated using the frame identifiers of the image frames on which the target display frame depends for decoding (i.e., the non-discardable image frames) and the frame identifier of the target display frame.

[0037] Exemplarily, the number of the first image frames is less than or equal to the number of the image frames before the target display frame in the image group to which the target display frame belongs.

[0038] According to the first aspect, or any one of the implementations of the first aspect, the first type of path includes a second decoding path, and when one of the plurality of candidate paths is the second decoding path, the method further includes: obtaining a second image frame, the decoding cost of the second image frame being less than the decoding cost of the target display frame; when the second image frame is not an I frame, obtaining one or more third image frames, the third image frames being before the second image frame, and the number of the third image frames being less than or equal to the total number of all the image frames before the second image frame in the image group to which the second image frame belongs; and generating the second decoding path using the frame identifiers of the one or more third image frames and the frame identifier of the second image frame, the reconstructed image of the target display frame being the reconstructed image of the second image frame.

[0039] That is, an image frame (i.e., the second image frame) with a decoding cost less than the decoding cost of the target display frame is selected to replace the target display frame, and the decoding path of the second image frame is used as the decoding path of the target display frame; in this way, the decoding cost of the target display frame can be reduced.

[0040] Exemplarily, the one or more third image frames can be determined according to the decoding dependency information of each image frame in the image group to which the second image frame belongs.

[0041] Exemplarily, a plurality of second image frames can be selected first, and then the second image frame with the minimum global cost is selected from the plurality of second image frames as the image frame finally used to replace the target display frame. Exemplarily, the global cost of each image frame can include the decoding cost of the image frame and the decoding dependency cost. Exemplarily, the decoding dependency cost of any image frame can be determined based on the number of image frames on which the candidate frame depends for decoding; the more the number of image frames on which the candidate frame depends for decoding, the smaller the decoding dependency cost; the fewer the number of image frames on which the candidate frame depends for decoding, the greater the decoding dependency cost.

[0042] According to the first aspect, or any one of the implementations of the above first aspect, determining the second decoding path for obtaining the reconstructed image of the target display frame further includes: when the second image frame is an I frame, generating the second decoding path based on a frame identifier of the second image frame. When the second image frame is an I frame, the decoding cost of obtaining the reconstructed image of the target display frame according to the second decoding path is smaller than that when the second image frame is not an I frame.

[0043] According to the first aspect, or any one of the implementations of the above first aspect, the first type of path includes a third decoding path. When the one of the plurality of candidate paths is the third decoding path, the method further includes: determining one or more predicted display frames from the plurality of image frames based on the user operation and the target display frame; obtaining a set of image frames, the set of image frames including a fourth image frame and one or more fifth image frames, the number of the fifth image frames being the same as the number of the predicted display frames, the decoding cost of the set of image frames being smaller than the sum of the decoding cost of the predicted display frames and the decoding cost of the target display frame; obtaining one or more sixth image frames, the one or more sixth image frames being before the fourth image frame, the number of the one or more sixth image frames being smaller than or equal to the total number of all image frames before the fourth image frame in the image group to which the fourth image frame belongs; and generating the third decoding path based on the frame identifiers of the one or more sixth image frames and the frame identifier of the fourth image frame, the reconstructed image of the target display frame being the reconstructed image of the fourth image frame.

[0044] Exemplarily, the predicted display frame can refer to a predicted subsequent target display frame. The predicted display frame can be the same as or different from the subsequent actual target display frame. That is, the target display frame and the predicted display frame are taken as a whole, and the image frame (i.e., the fourth image frame) used to replace the target display frame is selected according to the sum of the decoding costs of the target display frame and the predicted display frame. In this way, when the predicted display frame is relatively accurate (e.g., the predicted display frame is the same as or close to the actual target display frame), the decoding cost of the subsequent target display frame can be reduced, and thus the decoding cost of the whole (the whole composed of the target display frame and the predicted display frame) can be reduced.

[0045] Exemplarily, in the process of "obtaining a group of image frames", a plurality of groups of candidate image frames can be determined first; then a group of candidate image frames that is optimal in terms of uniformity of image frames and consistency of decoding paths can be selected as the group of image frames. The uniformity of image frames of each group of candidate image frames can be determined according to the difference in frame interval between any two pairs of adjacent frames in the fourth image frame and the plurality of fifth image frames in each group of candidate image frames. For example, the smaller the difference in interval between any two pairs of adjacent frames in the fourth image frame and the plurality of fifth image frames in each group of candidate image frames, the better the uniformity of image frames. The consistency of decoding paths of each group of candidate image frames can be determined according to the decoding paths of each image frame in the fourth image frame and the plurality of fifth image frames in each group of candidate image frames; the higher the coincidence degree of the decoding paths of each image frame in the fourth image frame and the plurality of fifth image frames in each group of candidate image frames, the better the consistency of decoding paths (because in the process of decoding a later image frame, the image frame decoded in the process of decoding a previous image frame can be reused, which can be specifically referred to the description of the determination process of the fourth decoding path).

[0046] Exemplarily, one or more sixth image frames can be determined according to the decoding dependency information of each image frame in the image group to which the fourth image frame belongs.

[0047] According to the first aspect, or any one of the implementations of the first aspect, the first type of path includes a fourth decoding path, when the one candidate path in the plurality of candidate paths is the fourth decoding path, the method further includes: determining a decoded image frame closest to the target display frame in the memory in the decoder; generating the fourth decoding path using the frame identifier of the one or more seventh image frames and the frame identifier of the target display frame, the seventh image frame being an image frame after the decoded image frame closest to the target display frame and before the target display frame.

[0048] It can be seen that the fourth decoding path is generated on the basis of reusing the decoded image frame in the memory in the decoder, that is, the decoding is continued from the decoded image frame closest to the target display frame; compared with decoding frame by frame from the I frame of the image group to which the target display frame belongs, the reconstruction image of the target display frame obtained according to the fourth decoding path can also reduce the decoding cost.

[0049] Exemplarily, one or more seventh image frames can be determined according to the decoding dependency information of each image frame in the image group to which the target display frame belongs.

[0050] According to the first aspect, or any one of the implementations of the first aspect, the obtaining the reconstructed image of the target display frame based on the target path and the code stream comprises: when the target path is the first type of path, determining a to-be-decoded image frame based on the target path; and decoding a part of the code stream corresponding to the to-be-decoded image frame to obtain the reconstructed image of the target display frame.

[0051] For example, when the target path is the second type of path, the reconstructed image of the image frame corresponding to the frame identifier in the second type of path is read from the memory in the decoder as the reconstructed image of the target display frame.

[0052] For example, when the target path is the third type of path, the reconstructed image of the image frame corresponding to the frame identifier in the third type of path is read from the memory outside the decoder as the reconstructed image of the target display frame.

[0053] According to the first aspect, or any one of the implementations of the first aspect, the video file further comprises decoding dependency information of each image frame in the plurality of image frames, and the obtaining the decoding dependency information of each image frame in the image group to which the target display frame belongs based on the video file comprises: decapsulating the decoding dependency information of each image frame in the image group to which the target display frame belongs from the video file.

[0054] In this way, the hierarchical structure (i.e., the decoding dependency information) in the entire code stream or a segment of the code stream can be known early, the decoding frame loss ratio can be controlled as a whole, and the second image frame with the minimum global cost can be accurately selected.

[0055] According to the first aspect, or any one of the implementations of the first aspect, the code stream further comprises decoding dependency information of each image frame in the plurality of image frames, and the obtaining the decoding dependency information of each image frame in the image group to which the target display frame belongs based on the video file comprises: decapsulating the video file to obtain the code stream; and parsing the decoding dependency information of each image frame in the image group to which the target display frame belongs from the code stream.

[0056] According to the first aspect, or any one of the implementations of the first aspect, the obtaining the decoding dependency information of each image frame in the image group to which the target display frame belongs based on the video file comprises: decapsulating the video file to obtain the code stream; parsing the code stream to determine the reference relationship of each image frame in the image group to which the target display frame belongs; and determining the decoding dependency information of each image frame in the image group to which the target display frame belongs according to the reference relationship of all the image frames in the image group to which the target display frame belongs. In this way, even if the video file does not carry the decoding dependency information, the decoding dependency information can also be determined through analysis of the information parsed from the code stream, so that the video playing method of the present application is applied more widely.

[0057] Exemplarily, the long-term reference frame, the short-term reference frame and the rearrangement information of the reference picture list of the first image frame can be parsed from the code stream; then, the reference picture list of the first image frame is determined based on the long-term reference frame, the short-term reference frame and the rearrangement information of the reference picture list of the first image frame; subsequently, the number S of reference frames of the first image frame is parsed from the code stream; then, the first S image frames are selected from the reference picture list of the first image frame as the reference frames of the first image frame; and finally, the reference relationship of the first image frame is generated based on the frame identifier of the first image frame and the frame identifiers of the reference frames of the first image frame.

[0058] According to the first aspect or any possible implementation of the above first aspect, the evaluation information comprises first evaluation information, the first evaluation information being used to describe a video watching experience; and the method further comprises: obtaining a video playing frame rate and / or a freezing information in a preset time length; and determining the first evaluation information corresponding to each candidate path in the plurality of candidate paths based on the video playing frame rate and / or the freezing information in the preset time length.

[0059] Exemplarily, the video watching experience can refer to the smoothness of the video.

[0060] Exemplarily, the preset time length can be set according to requirements, which is not limited in the present application.

[0061] According to the first aspect or any possible implementation of the above first aspect, the freezing information in the preset time length comprises at least one of a freezing frequency, a freezing time length or a speed variation amplitude in the preset time length.

[0062] According to the first aspect or any possible implementation of the above first aspect, the evaluation information further comprises second evaluation information, the second evaluation information being used to describe an interactive experience; and the method further comprises: obtaining an interactive starting time delay and an interactive ending time delay; and determining the second evaluation information corresponding to each candidate path in the plurality of candidate paths based on the interactive starting time delay and the interactive ending time delay.

[0063] It should be noted that the evaluation information comprises the second evaluation information only when the user operation is a variable speed playing operation (the variable speed playing operation can be an operation input by the user in real time). When the user operation is a fixed speed playing operation, the evaluation information comprises only the first evaluation information.

[0064] Exemplarily, the interactive experience can be understood as whether the response is in time with the starting point and the ending point.

[0065] Exemplarily, the interactive starting response time delay can refer to the difference between the time of the starting point detected by the mobile phone (such as the time when the user starts to drag or flick) and the actual display time of the target display frame corresponding to the starting point. The shorter the interactive starting response time delay is, the faster the response of the starting point is.

[0066] Exemplarily, the interaction end response time delay can refer to a difference between a time of the end point detected by the mobile phone (e.g., a time when the user ends the dragging or flicking) and an actual display time of the target display frame corresponding to the end point. The shorter the interaction end response time delay is, the faster the response of the end point is.

[0067] According to the first aspect, or any one of the implementations of the above first aspect, the selecting one candidate path from the plurality of candidate paths as the target path based on the evaluation information corresponding to each candidate path comprises: performing weighted calculation on the first evaluation information and the second evaluation information corresponding to the each candidate path according to the first weight and the second weight to obtain a weighted result corresponding to the each candidate path; and selecting one candidate path with the largest weighted result from the plurality of candidate paths as the target path.

[0068] Exemplarily, the first weight is a weight of the first evaluation information, and the second weight is a weight of the second evaluation information; wherein the first weight and the second weight can be set according to requirements, for example, the user pays more attention to the video watching experience, and the first weight can be set to be greater than the second weight; for another example, the user pays more attention to the interaction experience, and the second weight can be set to be greater than the first weight.

[0069] Exemplarily, the larger the weighted result is, the better the video playing experience is.

[0070] According to the first aspect, or any one of the implementations of the above first aspect, the determining the target display frame from the plurality of image frames comprises: determining the target display frame from the plurality of image frames based on the user operation, and the user operation can comprise one of a fixed multiple speed playing operation or a variable speed playing operation.

[0071] Exemplarily, the fixed multiple speed playing operation can comprise an operation of clicking a multiple option in FIG. 1A(3) (e.g., an operation of clicking a 1.5 times option). When the user operation is the fixed multiple speed operation, the corresponding playing scenario can be referred to as a fixed multiple speed playing scenario.

[0072] Exemplarily, the variable speed playing operation can comprise a dragging operation or a flicking operation on a progress bar (or a thumbnail preview bar or a video playing window). When the user operation is the variable speed playing operation, the corresponding playing scenario can be referred to as a variable speed playing scenario.

[0073] Wherein, the dragging can also be referred to as dragging; dragging: a touch screen control gesture, starting with a press-down action, continuously sliding on the screen, changing the sliding direction in the process, ending with a lift-up action, and the speed of the object (e.g., the progress bar) in the screen being 0 at the end.

[0074] Fling: A touch-screen gesture that starts with a press, followed by a short, fast swipe across the screen, and ends with a quick flick. When the flick is released, the object on the screen (such as a progress bar) has a high initial velocity, and then decelerates until it stops moving at a certain velocity.

[0075] In a second aspect, the present application provides a video recording method, which comprises: first, obtaining video data, the video data comprising original images of a plurality of image frames; then, encoding the original images of the plurality of image frames to obtain a code stream; subsequently, obtaining decoding dependency information corresponding to each image frame in the plurality of image frames, the decoding dependency information of a first image frame being used to indicate the number of second image frames that depend on the decoding of the first image frame, the first image being any image frame in the plurality of image frames, and the second image frame being an image frame after the first image frame in a group of images to which the first image frame belongs; and then, generating a video file based on the decoding dependency information of each image frame in the plurality of image frames and the code stream.

[0076] In this way, the obtained video file can carry the decoding dependency information; in a subsequent video playback process, a first type of path (such as a first decoding path, a second decoding path, a third decoding path, and a fourth decoding path) with a decoding cost smaller than that of the prior art can be determined based on the decoding dependency information.

[0077] In another aspect, the decoding dependency information of the first image frame can be used to indicate a decoding frame loss strategy of the first image frame, and the decoding frame loss strategy of the first image frame can refer to the second image frames to be discarded after the first image frame.

[0078] Exemplarily, all the image frames after the first image frame in the group of images to which the first image frame belongs can be referred to as the second image frames.

[0079] It should be noted that the image frames after a certain image frame in the present application refer to the image frames in decoding (encoding) order after the certain image frame; and the image frames before a certain image frame in the present application refer to the image frames in decoding (encoding) order before the certain image frame.

[0080] According to the second aspect, the decoding dependency information comprises any one of a first dependency level, a second dependency level, or a third dependency level;

[0081] The first dependency level of the first image frame indicates that all the second image frames after the first image frame depend on the decoding of the first image frame;

[0082] The second dependency level of the first image frame indicates that part of the second image frames after the first image frame depend on the decoding of the first image frame;

[0083] The third dependency level of the first image frame indicates that there is no second image frame after the first image frame that depends on the first image frame for decoding.

[0084] Exemplarily, the decoding loss frame strategy of the first image frame indicated by the first dependency level of the first image frame can be that, if the first image frame is discarded in the decoding process, all second image frames after the first image frame in a GOP to which the first image frame belongs are discarded.

[0085] Exemplarily, the decoding loss frame strategy of the first image frame indicated by the second dependency level of the first image frame can be that, if the first image frame is discarded in the decoding process, all second image frames after the first image frame in a GOP to which the first image frame belongs are discarded until a second image frame after the first image frame that has the first dependency level in the decoding dependency information is found.

[0086] Exemplarily, the decoding loss frame strategy of the first image frame indicated by the third dependency level of the first image frame can be that, if the first image frame is discarded in the decoding process, no second image frame after the first image frame in a GOP to which the first image frame belongs is discarded.

[0087] In a possible manner, the decoding dependency information includes any one of the first dependency level, the third dependency level or the frame identifier. The meanings of the first dependency level and the third dependency level can refer to the above description. The frame identifier included in the decoding dependency information can be used to indicate a second image frame that depends on the first image frame for decoding. The decoding loss frame strategy of the first image frame indicated by the frame identifier included in the decoding dependency information can be that, if the first image frame is discarded in the decoding process, the second image frame corresponding to the frame identifier is discarded.

[0088] In a possible manner, the decoding dependency information includes any one of the first dependency level, the third dependency level or (the frame identifier and the second dependency level).

[0089] According to the second aspect, or any one of the implementation manners of the above second aspect, the first dependency level includes any one of a first sub-level or a second sub-level.

[0090] The first sub-level of the first image frame indicates that all second image frames after the first image frame depend on the first image frame for decoding, and the first image frame is an I frame.

[0091] The second sub-level of the first image frame indicates that all second image frames after the first image frame depend on the first image frame for decoding, and the first image frame is a non-I frame.

[0092] According to the second aspect, or any one of the implementation manners of the above second aspect, the second dependency level includes any one of a third sub-level or a fourth sub-level.

[0093] The third sub-level of the first image frame indicates that all the second image frames after the first image frame depend on the first image frame for decoding, and the number of the second image frames depending on the first image frame for decoding is greater than the number threshold;

[0094] The fourth sub-level of the first image frame indicates that all the second image frames after the first image frame depend on the first image frame for decoding, and the number of the second image frames depending on the first image frame for decoding is less than or equal to the number threshold.

[0095] It should be understood that the third sub-level and the fourth sub-level can also be subdivided.

[0096] According to the second aspect, or any one of the implementation forms of the second aspect, the obtaining of the decoding dependency information of each image frame in the plurality of image frames comprises: parsing the code stream to obtain reference relationships of each image frame in the plurality of image frames; and determining the decoding dependency information of each image frame in the plurality of image frames according to the reference relationships of all the image frames in the plurality of image frames.

[0097] In a possible manner, the parsing of the code stream to obtain the reference relationship of each image frame in the plurality of image frames comprises: for the first image frame, parsing, from the code stream, long-term reference frames, short-term reference frames, and rearrangement information of a reference picture list of the first image frame; determining the reference picture list of the first image frame based on the long-term reference frames, the short-term reference frames, and the rearrangement information of the reference picture list of the first image frame; parsing, from the decoding, a reference frame number S of the first image frame; selecting, from the reference picture list of the first image frame, the first S image frames as the reference frames of the first image frame; and generating the reference relationship of the first image frame based on the frame identifier of the first image frame and the frame identifiers of the reference frames of the first image frame.

[0098] In a possible manner, the parsing of the code stream to obtain the reference relationship of each image frame in the plurality of image frames comprises: obtaining encoding information corresponding to original images of the plurality of image frames; and wherein the encoding information comprises the reference relationship of each image frame in the plurality of image frames.

[0099] According to the second aspect, or any one of the implementation forms of the second aspect, the generating of the video file based on the decoding dependency information of each image frame in the plurality of image frames and the code stream comprises: writing the decoding dependency information of each image frame in the plurality of image frames into the code stream; and encapsulating the code stream to obtain the video file.

[0100] According to a second aspect, or any possible implementation mode of the second aspect, the video file is generated based on the decoding dependency information of each image frame in the plurality of image frames and the code stream, and the generating includes: encapsulating the decoding dependency information of each image frame in the plurality of image frames and the code stream to obtain the video file. Compared with carrying the decoding dependency information in the code stream, this way can make the transmission process obtain the hierarchical structure (decoding dependency information) in the whole or a segment of the code stream, and can control the transmission frame loss ratio as a whole; and can make the video playing process know the hierarchical structure (decoding dependency information) in the whole or a segment of the code stream early, can control the decoding frame loss ratio as a whole, and can accurately select the second image frame and the fourth image frame with the minimum global cost.

[0101] According to a third aspect, the application provides a video conversion method, which includes: obtaining a first video file, the first video file including a code stream, the code stream being obtained from original images of a plurality of image frames in encoded video data; decapsulating the first video file to obtain the code stream; parsing the code stream to obtain reference relationships of each image frame in the plurality of image frames; determining decoding dependency information of each image frame in the plurality of image frames according to the reference relationships of all image frames in the plurality of image frames; and generating a second video file based on the decoding dependency information of each image frame in the plurality of image frames and the code stream. In this way, a video file without carrying decoding dependency information can be converted into a video file carrying decoding dependency information, and then in the subsequent video playing process, a first type path (such as a first decoding path, a second decoding path, a third decoding path and a fourth decoding path) with a decoding cost smaller than that of the prior art can be determined based on the decoding dependency information.

[0102] According to the third aspect, the second video file is generated based on the decoding dependency information of each image frame in the plurality of image frames and the code stream, and the generating includes: writing the decoding dependency information of each image frame in the plurality of image frames into the code stream; and encapsulating the code stream to obtain the second video file.

[0103] According to the third aspect, or any possible implementation mode of the third aspect, the second video file is generated based on the decoding dependency information of each image frame in the plurality of image frames and the code stream, and the generating includes: encapsulating the decoding dependency information of each image frame in the plurality of image frames and the code stream to obtain the second video file.

[0104] In a fourth aspect, the present application provides a video file, the video file comprising a bitstream, the bitstream comprising encoded data obtained by encoding original pictures of a plurality of image frames in video data and decoding dependency information of each image frame in the plurality of image frames, the decoding dependency information of a first image frame being used to indicate a number of second image frames that are dependent on decoding of the first image frame, the first image being any one of the plurality of image frames, and the second image frames being image frames after the first image frame in a group of pictures to which the first image frame belongs.

[0105] In a fifth aspect, the present application provides a video file, the video file comprising a bitstream and decoding dependency information of each image frame in a plurality of image frames; the bitstream being obtained by encoding original pictures of the plurality of image frames; and the decoding dependency information of a first image frame being used to indicate a number of second image frames that are dependent on decoding of the first image frame, the first image being any one of the plurality of image frames, and the second image frames being image frames after the first image frame in a group of pictures to which the first image frame belongs.

[0106] In a sixth aspect, the present application provides a bitstream, the bitstream comprising encoded data obtained by encoding original pictures of a plurality of image frames in video data and decoding dependency information of each image frame in the plurality of image frames, the decoding dependency information of a first image frame being used to indicate a number of second image frames that are dependent on decoding of the first image frame, the first image being any one of the plurality of image frames, and the second image frames being image frames after the first image frame in a group of pictures to which the first image frame belongs.

[0107] In a seventh aspect, the present application provides a video playing apparatus, the video playing apparatus comprising:

[0108] a first obtaining module, configured to obtain a video file, the video file comprising a bitstream, the bitstream comprising encoded data obtained by encoding original pictures of a plurality of image frames in video data;

[0109] a display frame determining module, configured to determine a target display frame from the plurality of image frames based on a user operation, the user operation comprising one of a fixed-speed playing operation or a variable-speed playing operation;

[0110] a target path determining module, configured to select one candidate path from a plurality of candidate paths as a target path based on evaluation information corresponding to each candidate path in the plurality of candidate paths, the evaluation information being used to describe a video playing experience;

[0111] a second obtaining module, configured to obtain a reconstructed picture of the target display frame based on the target path and the bitstream;

[0112] a display module, configured to play the reconstructed picture of the target display frame.

[0113] Exemplarily, the video playing device can be used to execute the video playing method in the first aspect or any possible implementation manner of the first aspect.

[0114] In an eighth aspect, the present application provides a video recording device, comprising:

[0115] a video data obtaining module, configured to obtain video data, the video data comprising original images of a plurality of image frames;

[0116] an encoding module, configured to encode the original images of the plurality of image frames to obtain a bitstream;

[0117] a decoding dependency information generating module, configured to obtain decoding dependency information corresponding to each image frame in the plurality of image frames, the decoding dependency information of a first image frame being used to indicate a number of second image frames that depend on decoding of the first image frame, the first image being any one of the plurality of image frames, and the second image frame being an image frame after the first image frame in a group of pictures to which the first image frame belongs;

[0118] a video file generating module, configured to generate a video file based on the decoding dependency information of each image frame in the plurality of image frames and the bitstream.

[0119] Exemplarily, the video recording device can be used to execute the video recording method in the second aspect or any possible implementation manner of the second aspect.

[0120] In a ninth aspect, the present application provides a video conversion device, comprising:

[0121] a video file obtaining module, configured to obtain a first video file, the first video file comprising a bitstream, the bitstream being obtained by encoding original images of a plurality of image frames in video data;

[0122] a video unpacking module, configured to unpack the first video file to obtain the bitstream;

[0123] a decoding dependency information generating module, configured to parse the bitstream to obtain a reference relationship of each image frame in the plurality of image frames, and determine decoding dependency information of each image frame in the plurality of image frames according to the reference relationship of all the image frames in the plurality of image frames;

[0124] a video file generating module, configured to generate a second video file based on the decoding dependency information of each image frame in the plurality of image frames and the bitstream.

[0125] Exemplarily, the video conversion device can be used to execute the video conversion method in the third aspect or any possible implementation manner of the third aspect.

[0126] In a tenth aspect, the present application provides an electronic device, comprising: a memory and a processor, the memory coupled to the processor; the memory storing program instructions that, when executed by the processor, cause the electronic device to perform the method of the first aspect or any possible implementation of the first aspect, or, perform the method of the second aspect or any possible implementation of the second aspect, or, perform the method of the third aspect or any possible implementation of the third aspect.

[0127] In an eleventh aspect, the present application provides a chip, comprising one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits; when the one or more processors execute computer instructions, the electronic device performs the method of the first aspect or any possible implementation of the first aspect, or, performs the method of the second aspect or any possible implementation of the second aspect, or, performs the method of the third aspect or any possible implementation of the third aspect.

[0128] In a twelfth aspect, the present application provides a computer-readable storage medium, the computer-readable storage medium storing a computer program, when the computer program runs on a computer or a processor, the computer or the processor performs the method of the first aspect or any possible implementation of the first aspect, or, performs the method of the second aspect or any possible implementation of the second aspect, or, performs the method of the third aspect or any possible implementation of the third aspect.

[0129] In a thirteenth aspect, the present application provides a computer program product, the computer program product comprising computer instructions, when the computer instructions are executed by a computer or a processor, the computer or the processor performs the method of the first aspect or any possible implementation of the first aspect, or, performs the method of the second aspect or any possible implementation of the second aspect, or, performs the method of the third aspect or any possible implementation of the third aspect.

[0130] In a fourteenth aspect, the present application provides a computer-readable storage medium, the computer-readable storage medium storing a video file, the video file comprising a bitstream, the bitstream comprising encoded data of original pictures of a plurality of picture frames in encoded video data and decoding dependency information of each picture frame in the plurality of picture frames, the decoding dependency information of a first picture frame indicating a number of second picture frames that depend on decoding of the first picture frame, the first picture being any one of the plurality of picture frames, the second picture frames being picture frames after the first picture frame in a group of pictures to which the first picture frame belongs.

[0131] In a fifteenth aspect, the present application provides a computer readable storage medium, which stores a video file, the video file comprising a bitstream and decoding dependency information of each image frame in a plurality of image frames; the bitstream is obtained by encoding original images of the plurality of image frames; the decoding dependency information of a first image frame is used to indicate a number of second image frames that depend on decoding of the first image frame, the first image being any one of the plurality of image frames, and the second image frame being an image frame after the first image frame in a group of pictures to which the first image frame belongs.

[0132] In a sixteenth aspect, the present application provides a device for storing a video file, the device comprising: a receiver configured to receive the video file; and at least one storage medium configured to store the video file; the video file is generated according to the second aspect and any one of the implementation manners of the second aspect, or the video file is a second video file generated according to the third aspect and any one of the implementation manners of the third aspect.

[0133] In a seventeenth aspect, the present application provides a device for transmitting a video file, the device comprising: a transmitter and at least one storage medium, the at least one storage medium being configured to store the video file, the video file being generated according to the second aspect and any one of the implementation manners of the second aspect, or the video file being a second video file generated according to the third aspect and any one of the implementation manners of the third aspect; the transmitter being configured to obtain the video file from the storage medium and transmit the video file to an end-side device through a transmission medium.

[0134] In an eighteenth aspect, the present application provides a system for distributing a video file, the system comprising: at least one storage medium configured to store at least one video file, the at least one video file being generated according to the second aspect and any one of the implementation manners of the second aspect, or the video file being a second video file generated according to the third aspect and any one of the implementation manners of the third aspect; and a streaming device configured to obtain a target video file from the at least one storage medium and transmit the target video file to an end-side device, wherein the streaming device comprises a content server or a content distribution server.

[0135] In a nineteenth aspect, the present application provides a video player, the video player being configured to:

[0136] obtain a video file, the video file comprising a bitstream, the bitstream comprising encoded data obtained by encoding original images of a plurality of image frames;

[0137] determine a target display frame from the plurality of image frames based on a user operation, the user operation comprising one of a fixed-speed playback operation or a variable-speed playback operation;

[0138] determine a plurality of candidate paths for obtaining a reconstructed image of the target display frame;

[0139] determine evaluation information corresponding to each candidate path in the plurality of candidate paths, the evaluation information being used to describe a video playback experience;

[0140] select a candidate path from the plurality of candidate paths as a target path based on the evaluation information corresponding to each candidate path;

[0141] obtain the reconstructed image of the target display frame based on the target path and the bitstream;

[0142] play the reconstructed image of the target display frame.

[0143] In a twentieth aspect, a video recorder is provided, and the video recorder is used to:

[0144] obtain video data, the video data comprising original images of a plurality of image frames;

[0145] encode the original images of the plurality of image frames to obtain a bitstream;

[0146] obtain decoding dependency information corresponding to each image frame in the plurality of image frames, the decoding dependency information of a first image frame being used to indicate a number of second image frames that depend on decoding of the first image frame, the first image frame being any one of the plurality of image frames, and the second image frame being an image frame after the first image frame in a group of pictures to which the first image frame belongs;

[0147] generate a video file based on the decoding dependency information of each image frame in the plurality of image frames and the bitstream.

[0148] The electronic device, the computer-readable storage medium, the computer program product, the chip or codec, the system, and the like provided in the embodiments of the present application are used to execute the corresponding method provided above, and thus the beneficial effects achieved thereby can refer to the beneficial effects of the corresponding method provided above. BRIEF DESCRIPTION OF DRAWINGS

[0149] FIG. 1A is a schematic diagram of an application scenario of an embodiment of the present application;

[0150] FIG. 1B is a schematic diagram of another application scenario of an embodiment of the present application;

[0151] FIG. 2 is a schematic diagram of still another application scenario of an embodiment of the present application;

[0152] FIG. 3A is a schematic diagram of a structure of a video recorder and a video player according to an embodiment of the present application;

[0153] FIG. 3B is a schematic diagram of a structure of another video player according to an embodiment of the present application;

[0154] FIG. 4 is a structural schematic diagram of a video converter according to an embodiment of the present application;

[0155] FIG. 5 is a software structural block diagram of an electronic device according to an embodiment of the present application;

[0156] FIG. 6 is a schematic diagram of a video playing process 600 according to an embodiment of the present application;

[0157] FIG. 7A is a schematic diagram of another video playing process 700 according to an embodiment of the present application;

[0158] FIG. 7B is a schematic diagram of decoding dependency information of each image frame in a group of pictures according to an embodiment of the present application;

[0159] FIG. 7C is a schematic diagram of decoding dependency information of each image frame in a group of pictures according to another embodiment of the present application;

[0160] FIG. 7D is a schematic diagram of a first decoding path according to an embodiment of the present application;

[0161] FIG. 7E is a schematic diagram of adjusting a target display frame according to an embodiment of the present application;

[0162] FIG. 7F is a schematic diagram of adjusting a target display frame according to another embodiment of the present application;

[0163] FIG. 7G is a schematic diagram of adjusting a target display frame according to another embodiment of the present application;

[0164] FIG. 7H is a schematic diagram of a group of pictures according to an embodiment of the present application;

[0165] FIG. 7I is a schematic diagram of a target display frame in a reverse playing process according to an embodiment of the present application;

[0166] FIG. 8A is a schematic diagram of a decoding dependency information determination process 800 according to an embodiment of the present application;

[0167] FIG. 8B is a schematic diagram of a reference relationship determination process according to an embodiment of the present application;

[0168] FIG. 8C is a schematic diagram of a bitstream structure according to an embodiment of the present application;

[0169] FIG. 8D is a schematic diagram of a reference relationship according to an embodiment of the present application;

[0170] FIG. 9 is a schematic diagram of a video recording process 900 according to an embodiment of the present application;

[0171] FIG. 10 is a schematic diagram of a video conversion process 1000 according to an embodiment of the present application;

[0172] FIG. 11 is a structural schematic diagram of a video playing apparatus 1100 according to an embodiment of the present application;

[0173] FIG. 12 is a structural schematic diagram of a video recording device 1200 according to an embodiment of the present application;

[0174] FIG. 13 is a structural schematic diagram of a video conversion device 1300 according to an embodiment of the present application;

[0175] FIG. 14 is a structural schematic diagram of a device according to an embodiment of the present application. DETAILED DESCRIPTION

[0176] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0177] The term “and / or” in the present application is merely used to describe an association relationship of associated objects, and means that there can be three relationships, for example, A and / or B can mean that there are three cases of A alone, A and B together, and B alone.

[0178] The terms “first” and “second” and the like in the specification and claims of the embodiments of the present application are used to distinguish different objects, and are not used to describe a specific order of the objects. For example, the first target object and the second target object are used to distinguish different target objects, and are not used to describe a specific order of the target objects.

[0179] In the embodiments of the present application, the words “exemplarily” or “for example” and the like are used to mean as an example, illustration or description. Any embodiment or design scheme described as “exemplarily” or “for example” in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words “exemplarily” or “for example” and the like are intended to present the relevant concept in a specific manner.

[0180] In the description of the embodiments of the present application, unless otherwise specified, the meaning of “a plurality of” is two or more. For example, a plurality of processing units means two or more processing units; a plurality of systems means two or more systems.

[0181] In the embodiments of the present application, the modules / components shown in the framework diagram (or structure diagram or system diagram) are only one example of the present application, and the actual framework (or structure or system) can include more or fewer modules / components than those shown in the diagram, or can have a different configuration of components. In addition, the various components / modules shown in the diagram can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application specific integrated circuits.

[0182] FIG. 1A is a schematic diagram of an application scenario of an embodiment of the present application. FIG. 1A shows a scenario in which a video is played at a fixed speed in a video application.

[0183] For example, after a user clicks on a video (such as a TV series, a movie, etc.) in a video application, the phone can enter a video playing interface 101 in response to the user's operation, as shown in FIG. 1A(1). The video playing interface 101 can include, but is not limited to, the following controls: a video playing window, a media title, a media selection, a comment editing box, a more function option 102, and the like.

[0184] When the user expects the video to be played at a fixed speed, the more function option 102 can be clicked, and the phone can display a screen mirroring option, a speed option 103, a sharing option, and the like in response to the user's operation, as shown in FIG. 1A(2). Then, the user can click the speed option 103, and the phone can display a plurality of multiple options 104 in response to the user's operation, such as a 0.5 times option (x0.5), a 0.75 times option (x0.75), a 1.0 times option (x1.0), a 1.5 times option (x1.5), a 2.0 times option (x2.0), a 3.0 times option (x3.0), and the like, as shown in FIG. 1A(3). Then, if the user expects the video to be played at 1.5 times speed, the user can click the 1.5 times option, and the phone can adjust the playing speed of the video in the video playing window to 1.5 times the original playing speed of the video in response to the user's operation. It should be understood that if the user expects the video to be played at other speeds, the user can also click other multiple options, and the phone can adjust the playing speed of the video in the video playing window to other multiples (i.e., speeds corresponding to the other multiple options) of the original playing speed of the video in response to the user's operation.

[0185] FIG. 1B is another schematic diagram of an application scenario of an embodiment of the present application. FIG. 1B shows a scenario in which a video is played at variable speed in a video application.

[0186] For example, after a user clicks on a video (such as a TV series or movie) in a video application, the mobile phone can respond to the user's operation and enter the video playback interface 101, as shown in Figure 1B. The video playback interface 101 may include, but is not limited to, the following controls: video playback window, progress bar 105, pause / start button, next episode button, previous episode button, media title, media selection, comment editing box, more function options, etc.

[0187] In one possible approach, the user can drag the progress bar (e.g., the user can move the progress bar left (or right) to rewind (or fast forward) the video, as shown in Figure 1B) to adjust the video playback speed. For example, while the user is dragging the progress bar, the phone can periodically obtain the progress bar position and display the corresponding screen after each position is obtained. It should be understood that the user can also perform drag operations within the video playback window to adjust the video playback speed.

[0188] Among them, dragging can also be called dragging; dragging: a touch screen control gesture that starts with a pressing action, slides continuously on the screen, and can change the direction of the slide during the process, and ends with a lifting action, at which point the movement speed of objects on the screen (such as progress bars) is 0.

[0189] In one possible approach, the user can perform a flicking motion on the video playback window or progress bar (no corresponding illustration provided) to adjust the video playback speed. Fling: A touchscreen gesture that begins with a press, followed by a brief, rapid swipe across the screen and then a quick release. Upon release, the object on the screen (such as a progress bar) has a high initial velocity, then decelerates until the velocity falls below a certain value. For example, during the flicking motion (before the user releases their finger), the phone can periodically acquire the progress bar position and display the corresponding screen after each position is acquired. Upon completion of the flicking motion (when the user releases their finger), the initial velocity of the progress bar is determined; then, multiple progress bar positions are sequentially determined based on the initial velocity; subsequently, screens corresponding to each progress bar position are displayed sequentially.

[0190] Figure 2 is a schematic diagram of another application scenario of the present application embodiment. Figure 2 shows a scenario of previewing a video in an album / gallery.

[0191] Figure 2(1) shows the main interface 201 of the gallery, which includes image options, video options 202, discovery options, etc. When a user needs to view a video, they can click on the video option 202, and the mobile phone can respond to the user's operation and display the video browsing interface 203, as shown in Figure 2(2).

[0192] Referring to FIG. 2(2), the video browsing interface 203 can include, but is not limited to, a video playing window 205, a thumbnail preview bar 204, a sharing option, a deleting option, and the like. In one possible implementation, a user can perform a drag operation (as shown in FIG. 2) on the thumbnail preview bar 204 to adjust the playing speed of the video. During the user performing the drag operation, the mobile phone can periodically acquire the position of the thumbnail preview bar, and can display the frame corresponding to the position of the thumbnail preview bar in the video playing window 205 after acquiring each position of the thumbnail preview bar.

[0193] In one possible implementation, a user can perform a flick operation (not shown in the corresponding diagram) on the thumbnail preview bar 204 to adjust the playing speed of the video. For example, during the user performing the flick operation, the mobile phone can periodically acquire the position of the thumbnail preview bar, and can display the frame corresponding to the position of the thumbnail preview bar. When the user finishes the flick operation, the initial speed is determined; then, a plurality of positions of the thumbnail preview bar are determined according to the initial speed; and subsequently, the frames corresponding to the plurality of positions of the thumbnail preview bar are displayed in sequence.

[0194] It should be understood that the present application can also include other application scenarios, such as the scenario of previewing the video before (or after) editing in a video editing application, and the like, which are not limited in the present application.

[0195] It should be understood that the mobile phone in FIGS. 1A, 1B, and 2 can also be replaced by other terminal devices, such as a tablet computer, a personal computer, a smart wearable device, and the like, which are not limited in the present application.

[0196] FIG. 3A is a structural schematic diagram of a video recorder and a video player according to an embodiment of the present application.

[0197] The video recorder 300 in FIG. 3A can include a video generation module 301, an encoder 302, a decoding dependency information generation module 303, and a video packaging module 304.

[0198] For example, the video generation module 301 can be a camera, a screen recording tool, or the like. The video generation module 301 can generate video data 11, which can include original images of a plurality of image frames.

[0199] For example, the encoder 302 can be a video encoder (such as an H264 encoder, an H265 encoder, or the like). The encoder 302 can encode the input video data 11 to obtain a code stream 12. The code stream can also be referred to as a bit stream, a bit stream, an encoded code stream, or the like, which are not limited in the present application.

[0200] Exemplarily, the encoder 302 outputs 13 to the decoding dependency information generation module 303 can include at least one of the code stream 12 or the encoding information. Wherein, the encoding information can refer to information related to the encoding of the video data 11.

[0201] Exemplarily, the decoding dependency information generation module 303 can determine the decoding dependency information of each image frame in the plurality of image frames of the video data based on at least one of the code stream 12 or the encoding information. Wherein, the decoding dependency information of each image frame can be used to indicate the number of image frames that depend on the decoding (or encoding) of the image frame; “depend” can be understood as “directly reference” or “indirectly reference”.

[0202] Exemplarily, the decoding dependency information generation module 303 can write the decoding dependency information of each image frame in the plurality of image frames of the video data into the code stream 12.

[0203] Exemplarily, the video encapsulation module 304 can be used for video encapsulation.

[0204] In one possible way, the decoding dependency information generation module 303 outputs 14 to the video encapsulation module 304 is the decoding dependency information of each image frame in the plurality of image frames of the video data; at this time, the encoder 302 can output the code stream 12 to the video encapsulation module 304. After that, the video encapsulation module 304 can encapsulate the decoding dependency information 14 of each image frame in the plurality of image frames of the video data and the code stream 12 to obtain the video file 15 (such as mp4 file).

[0205] In one possible way, the decoding dependency information generation module 303 outputs 14 to the video encapsulation module 304 is the code stream carrying the decoding dependency information of each image frame in the plurality of image frames of the video data, at this time, the encoder 302 does not need to output the code stream 12 to the video encapsulation module 304. After that, the video encapsulation module 304 can encapsulate the code stream 14 to obtain the video file 15.

[0206] The video player 310 in FIG. 3A can include a video decapsulation module 311, a video playback experience evaluation module 312, a decoder 313, a display control module 314 and a display module 315.

[0207] Exemplarily, the video decapsulation module 311 can decapsulate the video file 15.

[0208] In one possible way, the video decapsulation module 311 decapsulates the video file 15 to obtain the code stream 12.

[0209] In a possible implementation, the video decapsulation module 311 decapsulates the video file 15 to obtain the code stream 12 and the decoding dependency information 14 of each image frame in the plurality of image frames of the video data.

[0210] Exemplarily, the video playback experience evaluation module 312 can determine a plurality of candidate paths (the candidate paths include at least one of a first type of path, a second type of path or a third type of path, which will be described later) of the reconstructed image of the target display frame based on the decoding dependency information 14 of each image frame in the plurality of image frames of the video data, determine evaluation information of each candidate path, and select a candidate path as a target path according to the evaluation information of each candidate path.

[0211] In a possible implementation, the video playback experience evaluation module 312 can output 17 (which can include any one of the first type of path or the second type of path) to the decoder 313. The first type of path can include the frame identifier of one or more to-be-decoded image frames, and the second type of path can include the frame identifier of a decoded image frame in the memory in the decoder.

[0212] In a possible implementation, the video playback experience evaluation module 312 can output the third type of path 18 to the display control module 314. The third type of path can include the frame identifier of a to-be-displayed image frame in the memory (such as a graphics card memory (i.e., a display memory) for storing an image to be displayed) outside the decoder.

[0213] Exemplarily, when 17 is the first type of path, the decoder 313 can decode the code stream according to the first type of path to obtain the reconstructed image 19 of the target display frame and output the reconstructed image 19 to the memory outside the decoder.

[0214] Exemplarily, when 17 is the second type of path, the decoder 313 can select the reconstructed image 19 of the image frame corresponding to the frame identifier in the second type of path from the cache of the decoder 313 (at this time, the reconstructed image of the image frame corresponding to the frame identifier in the second type of path is determined as the reconstructed image of the target display frame) and output the reconstructed image 19 to the memory outside the decoder.

[0215] Exemplarily, the display control module 314 is configured to send the reconstructed image 20 of the target display frame in the memory outside the decoder to the display module 315 at a predicted display time corresponding to the target display frame (that is, control which frame is sent to the display at what time).

[0216] Exemplarily, the display module 315 can be configured to display an image (such as the reconstructed image 20 of the target display frame).

[0217] FIG. 3B is a structural schematic diagram of another video player according to an embodiment of the present application. The video file 21 in FIG. 3B is different from the video file 15 in FIG. 3A, the video file 21 in FIG. 3B is not generated by the video recorder according to an embodiment of the present application, but is generated by a video recorder of the prior art, and the video file 21 does not carry the decoding dependency information.

[0218] The video player 310 in FIG. 3B can include a video un-encapsulation module 311, a video playback experience evaluation module 312, a decoder 313, a display control module 314, a display module 315, and a decoding dependency information generation module 316. The decoding dependency information generation module 316 in FIG. 3B is similar to the decoding dependency information generation module 303 in FIG. 3A, and will not be described here. The decoding dependency information generation module 316 in FIG. 3B generates the decoding dependency information 22 of each of the plurality of image frames of the video data based on the code stream; the specific process will be described later.

[0219] In this way, for the video file generated by the video recorder of the prior art, the video player in FIG. 3B according to an embodiment of the present application can also perform video playback.

[0220] FIG. 4 is a structural schematic diagram of a video converter according to an embodiment of the present application. The first video file 31 in FIG. 4 is different from the video file 15 in FIG. 3A, the first video file 31 in FIG. 4 is not generated by the video recorder according to an embodiment of the present application, but is generated by a video recorder of the prior art, and the video file 31 does not carry the decoding dependency information.

[0221] The video converter 400 in FIG. 4 can include a video un-encapsulation module 401, a decoding dependency information generation module 402, and a video encapsulation module 403. Exemplarily, after the video un-encapsulation module 401 un-encapsulates the code stream 32 from the first video file 31, the video un-encapsulation module 401 can output the code stream 32 to the decoding dependency information generation module 402 and the video encapsulation module 403.

[0222] In one possible manner, the decoding dependency information generation module 402 can generate the decoding dependency information 33 of each of the plurality of image frames of the video data based on the code stream 32 and output to the video encapsulation module 403. Then, the video encapsulation module 403 can encapsulate the decoding dependency information 33 of each of the plurality of image frames of the video data and the code stream 32 to obtain the second video file 34.

[0223] In a possible implementation, the decoding dependency information generation module 402 can generate decoding dependency information of each of the plurality of image frames of the video data based on the code stream 32, and then write the decoding dependency information of each of the plurality of image frames of the video data into the code stream 32 to obtain a code stream 33 (the code stream 33 carries the decoding dependency information of each of the plurality of image frames of the video data) and output to the video encapsulation module 403. Subsequently, the video encapsulation module 403 can encapsulate the code stream 33 to obtain the second video file 34.

[0224] Thus, after the first video file 31 is converted into the second video file 34, the video player 310 in FIG. 3A can perform video playing on the second video file 34.

[0225] It should be noted that the video recorder 300 and the video player 310 in FIG. 3A can be arranged in the same terminal device, or can be arranged in different terminal devices, which are not limited in the present application. The video converter 400 in FIG. 4 can be arranged in a server or a terminal device. The server can include but is not limited to a cloud server, a physical (independent) server, a station cluster server, etc., which are not limited in the present application. The terminal device can include but is not limited to a personal computer, a computer workstation, a smart phone, a tablet computer, a smart camera, a smart car or other types of cellular phones, a media consumption device, a wearable device, a set-top box, a game console, etc.

[0226] It should be noted that the modules in the video recorder 300, the video player 310 and the video converter 400 can be realized by a combination of software and hardware, or can be realized by software, which are not limited in the present application.

[0227] FIG. 5 is a software structure block diagram of an electronic device according to an embodiment of the present application.

[0228] The layered architecture of the electronic device divides the software into several layers, each of which has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the mobile operating system is divided into four layers, from top to bottom, the application layer, the application framework layer, the system library, and the kernel layer.

[0229] The application layer can include a series of application packages.

[0230] As shown in FIG. 5, the application package can include camera, gallery, calendar, call, Bluetooth, music, video, wireless projection, etc.

[0231] The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes some pre-defined functions.

[0232] As shown in FIG. 5, the application framework layer can include a window manager, a content provider, a view system, a telephony manager, a resource manager, a notification manager, etc.

[0233] The window manager is used to manage window programs. The window manager can acquire the size of a display screen, determine whether there is a status bar, lock a screen, and take a screenshot, etc.

[0234] The content provider is used to store and acquire data, and make the data accessible to applications. The data can include videos, images, audios, dialed and received calls, browsing history and bookmarks, a phonebook, etc.

[0235] The view system includes visual controls, such as a control for displaying text, a control for displaying pictures, etc. The view system can be used to build an application. A display interface can be composed of one or more views. For example, a display interface including a short message notification control can include a view for displaying text and a view for displaying pictures.

[0236] The telephony manager is used to provide communication functions of the electronic device. For example, management of a call state (including call connection, call hang-up, etc.).

[0237] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc.

[0238] The notification manager enables an application to display notification information in a status bar. The notification manager can be used to convey a message of the notification type, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to notify a download completion, a message reminder, etc. The notification manager can also be a notification appearing in a system top status bar in a form of a graph or a scrolling text, a notification of an application running in the background, or a notification appearing on a screen in a form of a dialog window. For example, a text information is prompted in a status bar, a prompt sound is emitted, the electronic device is vibrated, a light flashes, etc.

[0239] The application layer and the application framework layer run in a virtual machine. The virtual machine executes java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions of management of an object life cycle, stack management, thread management, security and exception management, and garbage collection, etc.

[0240] The system library can include a plurality of functional modules. For example, a surface manager, a three-dimensional graphics processing library (e.g., OpenGL ES), a 2D graphics engine (e.g., SGL), a media framework (also referred to as media libraries in some cases), and the like.

[0241] The surface manager is used to manage a display subsystem and provides fusion of 2D and 3D layers for a plurality of applications.

[0242] The media framework supports playback and recording of a plurality of commonly used audio, video formats, and still image files. The media libraries can support a plurality of audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, and the like.

[0243] In particular, the media framework can include software portions of the video recorder 300, the video player 310, and the video converter 400. That is, to implement the video recording method, the video playback method, and the video conversion method of the present application, software programs (or software codes) of the decoding dependency information generation module 303 (or 316 or 402) and the video playback experience evaluation module 312 can be added to the media framework.

[0244] The three-dimensional graphics processing library is used to implement three-dimensional graphics drawing, image rendering, composition, and layer processing.

[0245] The 2D graphics engine is a drawing engine for 2D drawing.

[0246] The kernel layer is a layer between hardware and software. The kernel layer at least includes display drivers, camera drivers, audio drivers, sensor drivers, encoders, decoders, and the like.

[0247] It should be noted that the software portions of the video recorder 300, the video player 310, and the video converter 400 can also be implemented in the application framework layer. That is, to implement the video recording method, the video playback method, and the video conversion method of the present application, software programs (or software codes) of the decoding dependency information generation module 303 (or 316 or 402) and the video playback experience evaluation module 312 can be added to the application framework layer.

[0248] The video playback process, the video recording process, and the video conversion process are described below, respectively.

[0249] FIG. 6 is a schematic diagram of a video playback process 600 according to an embodiment of the present application. The video playback process 600 can be performed by the video player 310 in FIG. 3A or FIG. 3B.

[0250] S601, obtaining a video file, the video file comprising a bitstream, the bitstream comprising encoded data of original images of a plurality of image frames in encoded video data.

[0251] Exemplarily, the image frame in the present application can also be referred to as a frame.

[0252] In a possible manner, after a user selects a video (such as a TV series, a movie, etc.) to play in a video application, the video player can obtain a video file of the video.

[0253] In a possible manner, after a user selects a shot video to preview (or play) in a gallery, the video player can obtain a video file of the video.

[0254] In a possible manner, after a user selects an edited (or to-be-edited) video to preview (or play) in a video editing software, the video player can obtain a video file of the video.

[0255] Exemplarily, the video file obtained by the video player can be any one of the video file 15 in FIG. 3A, the video file 21 in FIG. 3B, or the second video file 34 in FIG. 4.

[0256] S602, determining a target display frame from the plurality of image frames.

[0257] In a possible manner, the video player can determine the target display frame from the plurality of image frames when the video player starts to play the video or in the process of playing the video. This case corresponds to an original speed playing scenario.

[0258] In a possible manner, in the process of watching the played video, when a user needs to adjust the playing speed of the video, any one of a fixed multiple speed playing operation or a variable speed playing operation can be performed as required. Then, the video player can determine the target display frame from the plurality of image frames based on the user operation, the user operation comprising one of the fixed multiple speed playing operation or the variable speed playing operation. This case corresponds to a fixed multiple speed scenario or a variable speed playing scenario. The present application takes the fixed multiple speed scenario or the variable speed playing scenario as an example for description.

[0259] For example, the fixed multiple speed playing operation can comprise an operation of clicking a multiple option in FIG. 1A(3) (such as an operation of clicking a 1.5 multiple option).

[0260] For example, the variable speed playing operation can comprise a drag operation or a flick operation on a progress bar (or a thumbnail preview bar or a video playing window).

[0261] After the mobile phone receives the user operation, the target display frame can be determined from the multiple image frames based on the user operation. The target display frame can refer to an image frame to be displayed.

[0262] When the user operation is a fixed speed play operation, the mobile phone can determine a next play position (which can be referred to as a target play position) according to a multiple corresponding to a multiple option selected by the user and a current play position (which can be understood as a current position of a progress bar). An image frame corresponding to a time of the target play position is determined as a target display frame. Then, the next target play position can be determined according to the multiple corresponding to the multiple option selected by the user and the last determined target play position. Then, an image frame corresponding to a time of the target play position is determined as the next target display frame. The above process is repeated.

[0263] When the user operation is a drag operation, the mobile phone can obtain the position of the progress bar according to a preset period (such as a screen acquisition period) because the position of the progress bar moves with the movement of the user's finger. After the mobile phone obtains each progress bar position, the progress bar position is determined as a target play position. Then, an image frame corresponding to a time of the target play position is determined as a target display frame.

[0264] When the user operation is a flick operation, the mobile phone can obtain the position of the progress bar according to a preset period (such as a screen acquisition period) because the position of the progress bar moves with the movement of the user's finger during the execution of the flick operation. After the mobile phone obtains each progress bar position, the progress bar position is determined as a target play position. Then, an image frame corresponding to a time of the target play position is determined as a target display frame.

[0265] When the user operation is a flick operation, the mobile phone can determine the initial speed of the progress bar when the user's finger is lifted within a period of time after the completion of the flick operation. Then, the next play position (i.e., a target play position) and the corresponding moving speed of the progress bar can be determined according to the play position when the finger is lifted, the initial speed, and a speed function (or referred to as a speed model) corresponding to the flick operation. An image frame corresponding to a time of the target play position is determined as a target display frame. Then, the next target play position and the corresponding moving speed of the progress bar can be determined according to the last determined target play position, the last determined moving speed of the progress bar corresponding to the last determined target play position, and the speed function corresponding to the flick operation. An image frame corresponding to a time of the target play position is determined as the next target display frame. The above process is repeated until the moving speed of the progress bar is 0.

[0266] It should be noted that when the user operation is a fixed speed play operation, or when the user operation is a flick operation and a time period after the user completes the flick operation, the mobile phone can continuously determine multiple target display frames; or the next target display frame can be determined after the display of the target display frame determined this time. When the user operation is a drag operation, or when the user operation is a flick operation and the user performs the flick operation, the mobile phone can determine a target display frame after obtaining the position of the progress bar each time.

[0267] S602 can refer to determining one or more target display frames, and S603-S605 are described by taking one target display frame as an example.

[0268] In S603, a candidate path is selected from the multiple candidate paths to determine a target path based on the evaluation information corresponding to each candidate path in the multiple candidate paths.

[0269] First, the multiple candidate paths for obtaining the reconstructed image of the target display frame can be determined.

[0270] In a possible manner, the decoding costs of any two candidate paths in the multiple candidate paths are the same, that is, the decoding costs of the multiple candidate paths are the same. The decoding cost can refer to the number of decoded image frames or decoding time required to obtain the reconstructed image of the target display frame.

[0271] In a possible manner, the decoding cost of at least one candidate path in the multiple candidate paths is different from that of the other candidate paths.

[0272] For example, the decoding costs of any two candidate paths in the multiple candidate paths are different, that is, the decoding costs of the multiple candidate paths are different.

[0273] Exemplarily, the candidate paths include at least one of a first type of path, a second type of path, or a third type of path; wherein:

[0274] The first type of path is used to indicate one or more image frames required to be decoded to obtain the reconstructed image of the target display frame;

[0275] The second type of path is used to indicate reading the reconstructed image of the target display frame from the memory in the decoder;

[0276] The third type of path is used to indicate reading the reconstructed image of the target display frame from the memory outside the decoder.

[0277] Exemplarily, the first type of path can also be referred to as a decoding path. The first type of path can include the frame identifier of one or more image frames required to be decoded to obtain the reconstructed image of the target display frame (that is, the frame identifier of the image frame to be decoded). When there are multiple candidate paths belonging to the first type of path, the image frames to be decoded corresponding to the multiple candidate paths are different.

[0278] Exemplarily, the second type of path can include a frame identifier of a decoded image frame in the memory within the decoder. When there are multiple candidate paths belonging to the second type of path, the decoded image frames corresponding to the multiple candidate paths are different.

[0279] Exemplarily, the third type of path can include a frame identifier of a to-be-displayed image frame in the memory outside the decoder. When there are multiple candidate paths belonging to the third type of path, the to-be-displayed image frames corresponding to the multiple candidate paths are different.

[0280] Then, evaluation information corresponding to each candidate path in the multiple candidate paths is determined, and the evaluation information is used to describe the video playback experience.

[0281] Exemplarily, for each candidate path in the multiple candidate paths, the video playback experience of the reconstructed image of the target display frame obtained according to the candidate path can be evaluated to obtain the evaluation information corresponding to the candidate path. The video playback experience can refer to the overall feeling of a user when watching the video content, including the smoothness of the video, the interactivity (in the case where the user performs a drag operation or a flick operation), and the like. The smoothness of the video, the interactivity, and the like are associated with the decoding cost; when the decoding cost is small, the smoothness of the video and the interactivity are good; when the decoding cost is large, the smoothness of the video and the interactivity are poor.

[0282] It should be noted that the multiple decoding costs belonging to the same numerical range can correspond to the same smoothness of the video and the same interactivity; because the human eye can only feel a stall when the interval between two adjacent frames exceeds a certain time length; and when performing a drag operation or a flick operation, the human eye also needs a certain reaction time to determine whether the picture is in hand.

[0283] Subsequently, based on the evaluation information corresponding to each candidate path, one candidate path is selected from the multiple candidate paths to be determined as the target path.

[0284] Exemplarily, based on the evaluation information corresponding to each candidate path, one candidate path can be selected from the first M candidate paths with the optimal video playback experience to be determined as the target path. M is a positive integer, and when M is 1, the candidate path with the optimal video playback experience is selected to be determined as the target path.

[0285] Exemplarily, when the candidate paths with the optimal video playback experience are multiple based on the evaluation information corresponding to each candidate path, one candidate path with the minimum decoding cost can be selected from the multiple candidate paths with the optimal video playback experience to be determined as the target path.

[0286] S604, based on the target path and the code stream, a reconstructed image of the target display frame is obtained.

[0287] Exemplarily, no matter which kind of path the target path is, the reconstructed image of the target display frame is obtained by decoding the code stream; only when the target path is the first kind of path, the code stream is decoded according to the target path to obtain the reconstructed image of the target display frame. When the target path is the second kind of path or the third kind of path, since in the last process of obtaining the reconstructed image of the target display frame, the decoding of the code stream can obtain the reconstructed image of the target display frame needed to be obtained this time, and the reconstructed image of the target display frame needed to be obtained this time is stored in the memory in the decoder or the memory outside the decoder; therefore, this time, the decoding of the code stream is not needed, but the reconstructed image of the target display frame needed to be obtained this time is obtained from the memory in the decoder or the memory outside the decoder.

[0288] S605, playing the reconstructed image of the target display frame.

[0289] Exemplarily, the display control module can also estimate the predicted display time of the target display frame. Then, the display control module can send the reconstructed image of the target display frame to the display module for display at the predicted display time of the target display frame.

[0290] In summary, in the video playing process, the present application first determines a plurality of candidate paths for obtaining the reconstructed image of the target display frame; then, determines the video playing experience corresponding to each candidate path in the plurality of candidate paths; and then, selects the candidate path with better video playing experience as the target path. The video playing experience can refer to the overall feeling of the user when watching the video content, including the smoothness of the video; in this way, after the reconstructed image of the target display frame is obtained according to the target path and played, the video playing experience is also better, that is, the smoothness of the video played at a fixed speed or variable speed is better.

[0291] FIG. 7A is a schematic diagram of another video playing process 700 of an embodiment of the present application. The video playing process 700 can be performed by the video player 310 in FIG. 3A. The video file received in the video playing process 700 includes any one of the video file 15 in FIG. 3A or the second video file 34 in FIG. 4. The candidate paths in the video playing process 700 include the first kind of path, the second kind of path and the third kind of path. The video playing process 700 describes the process of obtaining a plurality of candidate paths, and describes the process of determining the evaluation information.

[0292] S701, obtaining a video file, the video file including a code stream, the code stream including encoded data obtained by encoding a plurality of image frames in video data.

[0293] Exemplarily, the video file in S701 includes any one of the video file 15 in FIG. 3A or the second video file 34 in FIG. 4.

[0294] S702, determining a target display frame from the plurality of image frames based on the user operation, the user operation comprising one of a fixed speed play operation or a variable speed play operation.

[0295] Exemplarily, S702 can refer to the description of S602 described above, and will not be described here again.

[0296] S703, determining one or more predicted display frames from the plurality of image frames based on the user operation and the target display frame.

[0297] Exemplarily, after determining each target display frame, one or more predicted display frames (that is, predicted subsequent target display frame(s)) can be determined; the one or more predicted display frames can be used to assist in determining the candidate path.

[0298] It should be noted that when the user operation is a fixed speed play operation, or when the user operation is a flick operation and the predicted display frame is the same as the subsequently determined target display frame (both are determined in the same way) within a period of time after the user completes the flick operation. When the user operation is a drag operation, or when the user operation is a flick operation and the predicted display frame can be the same as the subsequently determined target display frame, or can be different (the determination methods are different).

[0299] Exemplarily, in each play scenario (such as a fixed speed play scenario (corresponding to the user operation of clicking the speed option), a drag variable speed play scenario (corresponding to the user operation of dragging), and a flick variable speed play scenario (corresponding to the user operation of flicking)), the moving speed of the progress bar is modelable; specifically, one or more predicted display frames can be determined according to the speed function (or speed model) corresponding to the scenario type, the target play position, and the moving speed of the progress bar corresponding to the target play position.

[0300] The determination process of the first predicted display frame is as follows: the moving speed of the progress bar corresponding to the first predicted play position can be determined according to the moving speed of the progress bar corresponding to the target play position, the speed function corresponding to the scenario type, and the time interval. Then, the first predicted play position is determined according to the target play position, the moving speed of the progress bar corresponding to the target play position, and the moving speed of the progress bar corresponding to the first predicted play position. Then, the time corresponding to the first predicted play position is determined according to the correspondence between the first predicted play position and the total length of the progress bar and the total duration of the video data; the image frame corresponding to the time of the first predicted play position is determined as the first predicted display frame. The time interval can be the difference between the predicted display time of the target display frame determined this time and the predicted display time (or the actual display time) of the target display frame determined last time.

[0301] The determination process of the second predicted display frame is as follows: the moving speed of the progress bar corresponding to the second predicted play position can be determined according to the moving speed of the progress bar corresponding to the first predicted play position, the speed function corresponding to the scene type, and the time interval. Then, the second predicted play position is determined according to the first predicted play position, the moving speed of the progress bar corresponding to the first predicted play position, the moving speed of the progress bar corresponding to the second predicted play position, and the like. Then, the time corresponding to the second predicted play position is determined according to the second predicted play position, the corresponding relationship between the total length of the progress bar and the total duration of the video data; and the image frame corresponding to the time corresponding to the second predicted play position is determined as the second predicted display frame. The time interval can refer to the difference between the predicted display time of the target display frame determined this time and the predicted display time of the first predicted display frame.

[0302] By analogy, details are not described herein.

[0303] For example, the play scene is a drag play scene. The first predicted play position can be obtained by referring to the following formula (1) to formula (2): n+1 = V(Velocity n , Δt) (1)

[0304] In formula (1), V() is a speed function / speed model, Velocity n and Δt (time interval) are inputs of the speed function, and Velocity n+1 is an output of the speed function.

[0305] In formula (2), Velocity represents the average speed between two play positions. It should be understood that the average speed between two play positions can also be determined by other methods, such as being determined according to a speed function / speed model.

[0306] When Velocity n is the moving speed of the progress bar corresponding to the target play position, Velocity n+1 is the moving speed of the progress bar corresponding to the first predicted play position, postion n is the target play position, postion n+1 is the first predicted play position.

[0307] When Velocity n is the moving speed of the progress bar corresponding to the first predicted play position, Velocity n+1Velocity n postion n+1 postion

[0308] For example, the playing scene is a swinging playing scene, the first predicted playing position can be obtained by referring to the following formula (3) and formula (4): n+1 Velocity n Δt-1000*f (3)

[0309] In formula (3) and formula (4), f is an adjustable friction coefficient, for example, f=-4.2.

[0310] S704, determine a plurality of candidate paths for obtaining a reconstructed image of the target display frame, the candidate paths including a first type of path, a second type of path and a third type of path.

[0311] Exemplarily, the first type of path can include a plurality of decoding paths: a first decoding path, a second decoding path, a third decoding path, a fourth decoding path and a fifth decoding path.

[0312] Exemplarily, the first decoding path can be determined according to a first decoding manner. The first decoding manner can refer to discarding the decoding of part of the image frames in the process of decoding the code stream. The determination process of the first decoding path can be as follows S11-S12:

[0313] S11, based on the video file, obtaining the decoding dependency information of each image frame in the image group to which the target display frame belongs.

[0314] In a possible manner, the video file can further include the decoding dependency information; the decoding dependency information of each image frame in the image group to which the target display frame belongs can be obtained by decapsulating the video file.

[0315] In a possible manner, the code stream can further include the decoding dependency information; the code stream can be obtained by decapsulating the video file; the decoding dependency information of each image frame in the image group to which the target display frame belongs can be obtained by parsing the code stream.

[0316] In a possible manner, the decoding dependency information of any image frame can include any one of a first dependency level, a second dependency level or a third dependency level; the meanings (semantics) of the first dependency level, the second dependency level or the third dependency level can be as shown in Table 1:

[0317] Table 1

[0318] In Table 1, the image frame with decoding dependency information of the first dependency level can be referred to as a key reference frame, and the corresponding level identifier can be 1. The image frame with decoding dependency information of the second dependency level can be referred to as a general reference frame, and the corresponding level identifier can be 2. The image frame with decoding dependency information of the third dependency level can be referred to as a non-reference frame, and the corresponding level identifier can be 3. The decoding frame loss strategies corresponding to the first dependency level, the second dependency level, and the third dependency level are different.

[0319] FIG. 7B is a schematic diagram of decoding dependency information of each image frame in a group of images according to an embodiment of the present application. In the figure, the video data includes a plurality of image frames, each of which has a corresponding frame sequence number (Frame Poc (Picture Order Count) (numbered according to the display order of the image frames), hereinafter referred to as F). The frame sequence number of the first image frame is 0, and the first image frame can also be referred to as F0. The frame sequence number of the second image frame is 1, and the second image frame can also be referred to as F1. Similarly, the frame sequence number of the third image frame is 2, and the third image frame can also be referred to as F2. Similarly, the frame sequence number of the fourth image frame is 3, and the fourth image frame can also be referred to as F3. Similarly, the frame sequence number of the fifth image frame is 4, and the fifth image frame can also be referred to as F4. Similarly, the frame sequence number of the sixth image frame is 5, and the sixth image frame can also be referred to as F5. Similarly, the frame sequence number of the seventh image frame is 6, and the seventh image frame can also be referred to as F6. Similarly, the frame sequence number of the eighth image frame is 7, and the eighth image frame can also be referred to as F7. Similarly, the frame sequence number of the ninth image frame is 8, and the ninth image frame can also be referred to as F8. Similarly, the frame sequence number of the tenth image frame is 9, and the tenth image frame can also be referred to as F9. Similarly, the frame sequence number of the eleventh image frame is 10, and the eleventh image frame can also be referred to as F10. Similarly, the frame sequence number of the twelfth image frame is 11, and the twelfth image frame can also be referred to as F11. Similarly, the frame sequence number of the thirteenth image frame is 12, and the thirteenth image frame can also be referred to as F12. Similarly, the frame sequence number of the fourteenth image frame is 13, and the fourteenth image frame can also be referred to as F13. Similarly, the frame sequence number of the fifteenth image frame is 14, and the fifteenth image frame can also be referred to as F14. Similarly, the frame sequence number of the sixteenth image frame is 15, and the sixteenth image frame can also be referred to as F15. Similarly, the frame sequence number of the seventeenth image frame is 16, and the seventeenth image frame can also be referred to as F16.

[0320] In FIG. 7B(1), F0, F8, and F17 are key reference frames, F2, F4, F6, F10, F12, and F14 are general reference frames, and F1, F3, F5, F7, F9, F11, F13, F15, and F16 are non-reference frames.

[0321] In FIG. 7B(2), F0, F4, F8, F12, and F17 are key reference frames, F2, F6, F10, and F14 are general reference frames, and F1, F3, F5, F7, F9, F11, F13, F15, and F16 are non-reference frames. It should be noted that the decoding order in FIG. 7B(2) is: F0, F4, F2, F1, F3, F8, F6, F5, F7, F12, F10, F9, F11, F16, F14, F13, F15.

[0322] In FIG. 7B(3), F0, F3, F6, F9, F12, F15, and F17 are key reference frames, F1, F4, F7, F10, and F13 are general reference frames, and F2, F5, F8, F11, F14, and F16 are non-reference frames.

[0323] In FIG. 7B(3), F0 to F15, and F17 are key reference frames, and F16 is a non-reference frame.

[0324] For example, the first dependency level can be further divided into a first sub-level or a second sub-level, and the second dependency level can be further divided into a third sub-level or a fourth sub-level. Table 2 shows the example as follows:

[0325] Table 2 For example, the first dependency level can be further divided into a first sub-level or a second sub-level, and the second dependency level can be further divided into a third sub-level or a fourth sub-level. Table 2 shows the example as follows:

[0326] In Table 2, the image frame whose decoding dependency information is the first sub-decoding dependency information can be referred to as an I frame (which is a kind of key reference frame); the image frame whose decoding dependency information is the second sub-decoding dependency information can be referred to as another key reference frame.

[0327] FIG. 7C is a schematic diagram of decoding dependency information of each image frame in another group of pictures according to an embodiment of the present application.

[0328] In FIG. 7C, F0 and F17 are I frames; F4, F8 and F12 are another key reference frames; F1, F5, F9 and F13, and F2, F6, F10 and F14 are general reference frames, wherein the decoding dependency information of F1, F5, F9 and F13 is the third sub-level, and the decoding dependency information of F2, F6, F10 and F14 is the fourth sub-level; F3, F7, F11, F15 and F16 are non-reference frames.

[0329] It should be understood that the third sub-level and the fourth sub-level can also be subdivided, which is not limited in the present application.

[0330] In a possible manner, the decoding dependency information of any image frame can include any one of a first dependency level, a third dependency level or a frame identifier of a second image frame on which the first image frame is decoded; wherein the meaning (semantics) of the first dependency level, the second dependency level or the frame identifier can be as shown in Table 3:

[0331] Table 3

[0332] It should be understood that the decoding dependency information of any image frame can include both the second dependency level and the frame identifier of the second image frame on which the first image frame is decoded.

[0333] S12, determining a first decoding path based on the decoding dependency information of each image frame in the group of pictures to which the target display frame belongs.

[0334] Exemplarily, one or more first image frames can be determined based on the decoding dependency information of each image frame in the group of pictures to which the target display frame belongs and the corresponding decoding frame loss strategy, the first image frame being an image frame on which the target display frame is dependent; a frame identifier of the one or more first image frames and a frame identifier of the target display frame are used to generate the first decoding path. The number of the first image frames is less than or equal to the total number of all image frames before the target display frame in the group of pictures to which the target display frame belongs.

[0335] Embodiments of the present application are described taking the decoding dependency information in Table 1 as an example.

[0336] FIG. 7D is a schematic diagram of a first decoding path according to an embodiment of the present application.

[0337] In FIG. 7D, the target display frame is the 9th image frame (F8), and the GOP to which the target display frame belongs includes 12 image frames (F0 to F11).

[0338] In FIG. 7D, F0, F3, F6 and F9 are key reference frames (or the decoding dependency information is the first dependency level), and the corresponding level identifier is "1"; F1, F4, F7 and F10 are general reference frames (or the decoding dependency information is the second dependency level), and the corresponding level identifier is "2"; F2, F5, F8 and F11 are non-reference frames (or the decoding dependency information is the third dependency level), and the corresponding level identifier is "3".

[0339] According to the decoding dependency information of each image frame in FIG. 7D and the corresponding decoding frame loss strategy, it can be determined that the reconstructed image of F8 can discard F1, F2, F4 and F5, and cannot discard F0, F3, F6, F7 and F8; that is, the first image frame can include F0, F3, F6, F7; and correspondingly, the first decoding path can be {0, 3, 6, 7, 8}.

[0340] Exemplarily, the second decoding path can be determined according to a second decoding manner. The second decoding manner can refer to adjusting (or replacing) the target display frame to a second image frame (the decoding cost of the second image frame is less than the decoding cost of the target display frame). The determination process of the second decoding path can be as follows S21-S23:

[0341] S21, obtaining a second image frame, the decoding cost of the second image frame being less than the decoding cost of the target display frame.

[0342] Exemplarily, one or more candidate image frames can be determined from a plurality of image frames before or after the target display frame according to a first preset condition. When the candidate image frames are multiple, one candidate image frame is selected from the multiple candidate image frames to determine the second image frame according to a second preset condition. When the candidate image frames are one, the candidate image frame can be determined as the second image frame.

[0343] Exemplarily, the first preset condition can be that the frame distance between the image frame used to replace the target display frame and the target display frame is less than a preset value (in this way, the content of the reconstructed image of the image frame used to replace the target display frame can be similar to the content of the reconstructed image of the target display frame), and the decoding cost of the image frame used to replace the target display frame is less than the decoding cost of the target display frame. The preset value can be set according to requirements, such as 1, 2, etc.

[0344] In a possible manner, the second preset condition can be to select the candidate image frame with the minimum decoding cost to determine the second image frame.

[0345] FIG. 7E is a schematic diagram of adjusting a target display frame according to an embodiment of the present application.

[0346] In FIG. 7E, each GOP includes 12 image frames, the first GOP includes F0 to F11, and the second GOP includes F12 to F23.

[0347] Suppose the target display frame is F7, the 8th image frame in the first GOP. If the prior art method is used, F7 needs to be decoded from F0 frame by frame, i.e., 8 image frames need to be decoded, to obtain the reconstructed image of F7.

[0348] For example, the preset value is 1, and the candidate image frame is determined as 1, i.e., F6. At this time, F6 is determined as the second image frame.

[0349] For example, the preset value is 3, and the candidate image frame is determined as 3, i.e., F4, F5 and F6. The decoding cost of F4 is 5 image frames (which can be understood as 5 image frames need to be decoded), the decoding cost of F5 is 6 image frames (which can be understood as 6 image frames need to be decoded), and the decoding cost of F6 is 7 image frames (which can be understood as 7 image frames need to be decoded). It can be seen that the decoding cost of F4 is the smallest. Therefore, F4 is determined as the second image frame.

[0350] Suppose the target display frame is F11, the 12th image frame in the first GOP. If the prior art method is used, F11 needs to be decoded from F0 frame by frame, i.e., 12 image frames need to be decoded, to obtain the reconstructed image of F11.

[0351] For example, the preset value is 1, and the candidate image frame is determined as F10 or F12. The decoding cost of F12 (only 1 image frame needs to be decoded) is smaller than the decoding cost of F10 (11 image frames need to be decoded). Therefore, F12 is determined as the second image frame.

[0352] In one possible manner, the second preset condition can be that the candidate image frame with the smallest global cost is selected as the second image frame. Illustratively, the global cost of each image frame can include the decoding cost of the image frame and the decoding dependency cost. Illustratively, the decoding dependency cost of any image frame can be determined based on the number of image frames that depend on the decoding of the candidate frame. The more image frames that depend on the decoding of the candidate frame, the smaller the decoding dependency cost. The fewer image frames that depend on the decoding of the candidate frame, the larger the decoding dependency cost.

[0353] Suppose the target display frame is F7, the 8th image frame in the first GOP in FIG. 7E. If the first decoding method described above is used, F0, F3, F6 and F7 need to be decoded, i.e., 4 image frames need to be decoded, to obtain the reconstructed image of F7.

[0354] For example, the preset value is 3, and two candidate image frames, F4 and F6, can be determined. Since the decoding cost of F6 and F4 is the same (both need to decode 3 image frames), but 5 image frames F7 to F11 are dependent on F6 decoding, and only one image frame F5 is dependent on F4 decoding, it can be seen that the decoding dependency cost of F6 is less than that of F4. Therefore, F6 can be determined as the second image frame.

[0355] Suppose that the target display frame in FIG. 7E is F11, the 12th image frame in the first GOP. If the first decoding method is used, 6 image frames F0, F3, F6, F9, F10 and F11 need to be decoded, and the reconstructed image of F11 can be obtained.

[0356] For example, the preset value is 1, and F10 or F12 can be determined as the candidate image frame. The decoding cost of F12 (only one image frame needs to be decoded) is less than that of F10 (5 image frames need to be decoded), and 11 image frames F13 to F23 are dependent on F12 decoding, and only one image frame F11 is dependent on F10 decoding. It can be seen that the global cost of F12 is less than that of F10. Therefore, F12 can be determined as the second image frame.

[0357] In one possible manner, the second preset condition can be that the candidate image frame with the minimum decoding cost is selected as the second image frame under the condition that the video is played in time sequence.

[0358] FIG. 7F is another schematic diagram of adjusting the target display frame according to an embodiment of the present application.

[0359] In FIG. 7F, suppose that the target display frame is F3, F6 and F9. If the prior art method is used, 4 image frames need to be decoded from F0 to F3, and the reconstructed image of F3 can be obtained. If the prior art method is used, 7 image frames need to be decoded from F0 to F6, and the reconstructed image of F6 can be obtained. If the prior art method is used, 10 image frames need to be decoded from F0 to F9, and the reconstructed image of F9 can be obtained.

[0360] For example, the preset value is equal to 12, F3 can be adjusted to F0, F6 can be adjusted to F1, and F9 can be adjusted to F2. In this way, the reconstructed image of F0 obtained by decoding is taken as the reconstructed image of the target display frame, and only one image frame needs to be decoded. The reconstructed image of F1 obtained by decoding is taken as the reconstructed image of the target display frame, and only two image frames need to be decoded. The reconstructed image of F2 obtained by decoding is taken as the reconstructed image of the target display frame, and only three image frames need to be decoded. In this way, the decoding cost can be greatly reduced.

[0361] S22, when the second image frame is not an I frame, obtaining one or more third image frames, the third image frames being before the second image frame, the number of the third image frames being less than or equal to the total number of all image frames before the second image frame in the image group to which the second image frame belongs; generating the second decoding path by using the frame identifiers of the one or more third image frames and the frame identifier of the second image frame.

[0362] In a possible manner, when the second image frame is not an I frame, all image frames before the second image frame in the image group to which the second image frame belongs can be determined as the third image frames. Then, the second decoding path is generated by using the frame identifiers of the one or more third image frames and the frame identifier of the second image frame.

[0363] For example, when the target display frame is F7 and the second image frame is F6 (F6 is a non-I frame) in FIG. 7E, F0 to F5 can be determined as the third image frames. In this case, the second decoding path is {0, 1, 2, 3, 4, 5}.

[0364] In a possible manner, the one or more third image frames can be determined in the manner of determining the first decoding path (i.e., the first decoding manner). Specifically, the one or more third image frames can be determined according to the decoding dependency information of each image frame in the plurality of image frames. Then, the second decoding path is generated by using the frame identifiers of the one or more third image frames and the frame identifier of the second image frame.

[0365] For example, when the target display frame is F7 and the second image frame is F6 (F6 is a non-I frame) in FIG. 7E, F0 and F3 can be determined as the third image frames. In this case, the second decoding path is {0, 3, 6}.

[0366] S23, when the second image frame is an I frame, generating the second decoding path by using the frame identifier of the second image frame.

[0367] For example, when the target display frame is F11 and the second image frame is F12 (F12 is an I frame) in FIG. 7E, the second decoding path can be generated by using the frame identifier of F12; in this case, the second decoding path is {12}.

[0368] In summary, the reconstructed image of the target display frame obtained according to the second decoding path is the reconstructed image of the second image frame.

[0369] It should be noted that the prediction display frame does not need to be determined before the first decoding path and the second decoding path are determined.

[0370] Exemplarily, the third decoding path can be determined according to a third decoding manner. The third decoding manner can refer to adjusting (or replacing) the target display frame into a fourth image frame in a group of image frames in combination with one or more predicted display frames, where the decoding cost of the group of image frames is less than the sum of the decoding cost of the target display frame and the one or more predicted display frames. The determination process of the third decoding path can be as follows S31-S34:

[0371] S31, obtaining a group of image frames, the group of image frames including a fourth image frame and one or more fifth image frames, the number of the fifth image frames being the same as the number of the predicted display frames, and the decoding cost of the group of image frames being less than the sum of the decoding cost of the predicted display frames and the target display frame.

[0372] Exemplarily, one or more fourth image frames can be selected from a plurality of image frames to replace the target display frame, and one or more fifth image frames can be selected from the image frames to replace one predicted display frame; where the fourth image frames and the fifth image frames can be selected according to the first preset condition described above, which will not be repeated here. Wherein one fourth image frame and one fifth image frame used to replace each predicted display frame can constitute a group of candidate image frames; one fifth image frame in the group of candidate image frames is used to replace one predicted display frame, and the number of the fifth image frames in the group of candidate image frames is the same as the number of the preset display frames determined in S703. Then, a group of candidate image frames can be selected from a plurality of groups of candidate image frames according to a third preset condition to determine a group of image frames in S31.

[0373] Exemplarily, the third preset condition can be a group of candidate image frames with the uniformity of the frame interval of the image frames, the consistency of the frame interval of the image frames and the original frame interval, the consistency of the decoding path of the image frames, and the overall decoding frame number of the image frames being the most optimal, and the group of image frames in S31 is determined. Wherein, the uniformity of the frame interval of the image frames in each group of candidate image frames can be determined according to the difference of the frame interval of any two pairs of adjacent frames in the fourth image frame and the plurality of fifth image frames in each group of candidate image frames. For example, the smaller the difference of the interval of any two pairs of adjacent frames in the fourth image frame and the plurality of fifth image frames in each group of candidate image frames, the better the uniformity of the image frames. The consistency of the frame interval of the image frames in each group of candidate image frames and the original frame interval can refer to the consistency of the frame interval of the group of candidate image frames and the target display frame and one or more predicted display frames. For example, the predicted display frame is 2, the frame intervals of the target display frame and the 2 predicted display frames are 1 and 3 respectively, the first group of image frames includes 3 image frames, and the frame intervals of the 3 image frames are 2 and 4 respectively, the second group of image frames includes 3 image frames, and the frame intervals of the 3 image frames are 1 and 3 respectively, it can be determined that the second group of image frames is more consistent with the target display frame and one or more predicted display frames. The consistency of the decoding path of each group of candidate image frames can be determined according to the decoding path of each image frame in the fourth image frame and the plurality of fifth image frames in each group of candidate image frames; the higher the coincidence degree of the decoding path of each image frame in the fourth image frame and the plurality of fifth image frames in each group of candidate image frames, the better the consistency of the decoding path (because in the process of decoding the image frames behind, the image frames decoded in the process of decoding the image frames in front can be reused, which can be referred to the description of the determination process of the fourth decoding path).

[0374] FIG. 7G is another schematic diagram of adjusting the target display frame according to an embodiment of the present application.

[0375] In FIG. 7G, F1 is the target display frame, F5 and F10 are the predicted display frames. Wherein, the fourth image frame F0 is selected from the plurality of image frames to replace F1, the fifth image frames F3, F4 and F6 are selected from the plurality of image frames to replace F5, and the fifth image frames F9 and F12 are selected from the plurality of image frames to replace F10. The fourth image frame and the fifth image frame can form the following groups of candidate image frames:

[0376] The first group of candidate image frames = {F0, F3, F9}, the second group of candidate image frames = {F0, F4, F9}, the third group of candidate image frames = {F0, F6, F9}, the fourth group of candidate image frames = {F0, F3, F12}, the fifth group of candidate image frames = {F0, F4, F12}, and the sixth group of candidate image frames = {F0, F6, F12}.

[0377] The first group of candidate image frames F0 and F3 are spaced apart by 2 image frames, and F3 and F9 are spaced apart by 5 image frames; F9, F3 and F0 in the decoding path of F0 coincide twice, F3 coincides once, and F6 and F9 do not coincide.

[0378] The second group of candidate image frames F0 and F4 are spaced apart by 3 image frames, and F4 and F9 are spaced apart by 4 image frames. F9, F4 and F0 in the decoding path of F0 coincide twice, F3 coincides once, and F4, F6 and F9 do not coincide.

[0379] The third group of candidate image frames F0 and F6 are spaced apart by 5 image frames, and F6 and F9 are spaced apart by 2 image frames. F9, F6 and F0 in the decoding path of F0 coincide twice, F3 and F6 coincide once, and F9 does not coincide.

[0380] The fourth group of candidate image frames F0 and F3 are spaced apart by 2 image frames, and F3 and F12 are spaced apart by 8 image frames. F12, F3 and F0 in the decoding path of F0 coincide once, and F3 and F12 do not coincide.

[0381] The fifth group of candidate image frames F0 and F4 are spaced apart by 3 image frames, and F4 and F12 are spaced apart by 7 image frames. F12, F4 and F0 in the decoding path of F0 coincide once, and F4, F3 and F12 do not coincide.

[0382] The sixth group of candidate image frames F0 and F6 are spaced apart by 5 image frames, and F6 and F9 are spaced apart by 2 image frames. F12, F6 and F0 in the decoding path of F0 coincide once, and F6, F3 and F12 do not coincide.

[0383] Exemplarily, the adjustment score of each group of candidate image frames can be calculated according to the frame interval, the number of coinciding image frames (and / or the number of non-coinciding image frames) or the total number of decoding frames; the group of candidate image frames with the highest (or the lowest, depending on the calculation method of the adjustment score) adjustment score is determined as the group of image frames in S31.

[0384] Suppose the adjustment score of the first group of candidate image frames is the highest, then F0 can be determined as the fourth image frame, and F3 and F9 can be determined as the fifth image frame.

[0385] S32, when the fourth image frame is not an I frame, one or more sixth image frames are obtained, the sixth image frames are before the fourth image frame, and the number of the sixth image frames is less than or equal to the total number of all image frames before the fourth image frame in the image group to which the fourth image frame belongs.

[0386] S33, the frame identifiers of the one or more sixth image frames and the frame identifier of the fourth image frame are used to generate a third decoding path.

[0387] S34, when the fourth image frame is not an I frame, generating the third decoding path according to the frame identifier of the fourth image frame.

[0388] For example, S32-S34 can refer to the description of S22-S23 above, and will not be repeated here.

[0389] For example, the fourth decoding path can be determined according to a fourth decoding manner. The fourth decoding manner can be to reuse the decoded image frames in the memory within the decoder. For example, the process of determining the fourth decoding path can be as follows S41-S42:

[0390] S41, determining the decoded image frame closest to the target display frame in the memory within the decoder.

[0391] FIG. 7H is a schematic diagram of an image group according to an embodiment of the present application.

[0392] In FIG. 7H, the target display frame determined this time is F8, and the last target display frame determined (i.e., the last target display frame) is F4.

[0393] In the embodiment of the present application, after the application layer determines the target playback position, it can issue a seek task to the media framework, and the media framework can execute the video playback method of the present application to determine the reconstructed image of the target display frame according to the target path. When the media framework decodes the code stream according to the target path to obtain the reconstructed image of the target display frame, if the media framework determines that there is no end identifier (used to indicate whether the drag operation or the flick operation is ended, or whether the fixed speed operation is ended) in the seek task, after the media framework obtains the reconstructed image of the target display frame output by the decoder and returns the reconstructed image of the target display frame to the application layer, the media framework does not instruct the decoder to clear the cache (in the prior art, the media framework will instruct the decoder to clear the cache after executing each seek task). In this way, the memory within the decoder can include the reconstructed images of R (R is a positive integer, which can be set according to requirements, such as R=4) decoded image frames. Assuming that R=4, if the last target display frame determined is F4, the memory within the decoder must contain F1, F2, F3 and F4. Therefore, the decoded image frame closest to the target display frame in the memory within the decoder is F4.

[0394] It should be understood that before S41, it can also be determined according to the decoding dependency information of each image frame whether the decoded image frame in the memory within the decoder is the image frame on which the reconstructed image of the target display frame depends. If it is determined that the decoded image frame in the memory within the decoder is the image frame on which the reconstructed image of the target display frame depends, S41 can be executed. Otherwise, at least one of the first decoding path, the second decoding path and the third decoding path can be determined.

[0395] S42, generating a fourth decoding path by using the frame identifier of the one or more seventh image frames and the frame identifier of the target display frame, the seventh image frames being all or part of the image frames after the nearest decoded image frame to the target display frame and before the target display frame.

[0396] Exemplarily, in the forward playing process, the fourth decoding path can be generated by using the frame identifier of the one or more seventh image frames and the frame identifier of the target display frame, the seventh image frames being all image frames after the nearest decoded image frame to the target display frame and before the target display frame. For example, in FIG. 7H, the seventh image frames are F5, F6 and F7, and the fourth decoding path is {5, 6, 7, 8}.

[0397] Exemplarily, in the forward playing process, the one or more seventh image frames can be determined based on the decoding dependency information of each image frame, the seventh image frames being the image frames required to be depended on for decoding the target display frame, and then the fourth decoding path is generated by using the frame identifier of the one or more seventh image frames and the frame identifier of the target display frame. In this case, the seventh image frames can be part of the image frames after the nearest decoded image frame to the target display frame and before the target display frame. For example, in FIG. 7H, the seventh image frames are F6 and F7, and the fourth decoding path is {6, 7, 8}.

[0398] It should be understood that the determination method (i.e., the first decoding manner) for determining the first decoding path can be used in the determination of the second decoding path, the third decoding path and the fourth decoding path. The determination method (i.e., the fourth decoding manner) for determining the fourth decoding path can be used in the determination of the second decoding path and the third decoding path. That is, the determination methods of the first decoding path, the second decoding path (or the third decoding path) and the fourth decoding path can be combined with each other.

[0399] Exemplarily, the fifth decoding path can be determined in the fifth decoding manner, the fifth decoding manner being to decode frame by frame starting from an I frame until the target display frame is decoded. Thus, the determined fifth decoding path can include the frame identifier of the eighth image frame and the frame identifier of the target display frame, the eighth image frame being all image frames before the target display frame in the image group to which the target display frame belongs.

[0400] Exemplarily, in the case that the maximum number of frames allowed to be decoded is insufficient to meet the decoding requirement (e.g., insufficient decoding capability, excessive power consumption, too long time consumption, etc.), the target display frame can be adjusted to a decoded image frame, so that the decoding cost can be greatly reduced in the case of equivalent frame rate. Further, in this case, the second type of path can be determined in this way. It should be understood that the second type of path can also be determined in this way in the case of small decoding pressure, which is not limited in the present application.

[0401] It should be noted that the second type of path can include one or more reading paths.

[0402] It should be noted that the prediction display frame can not be determined before the second type of path.

[0403] Exemplarily, in the scenario of reverse playback, in the process of decoding the code stream to obtain the reconstructed image of the target display frame, the reconstructed image of the subsequent target display frame can be decoded. Therefore, in order to reduce the decoding cost of the subsequent multiple target display frames, when the reconstructed image of the prediction display frame is determined to be decoded, the reconstructed image of the prediction display frame can be stored in the memory outside the decoder. In this way, when the subsequent target display frame is the same as the prediction display frame, the third type of path for indicating that the reconstructed image of the target display frame is read from the memory outside the decoder can be determined. It should be understood that when the subsequent target display frame is not the same as the prediction display frame but the frame distance between the subsequent target display frame and the prediction display frame is less than a distance threshold (which can be set as required, such as 1, 2, etc.), the third type of path can also be determined in this way, which is not limited in the present application.

[0404] It should be noted that the third type of path can include one or more reading paths.

[0405] FIG. 7I is a schematic diagram of a target display frame in a reverse playback process according to an embodiment of the present application.

[0406] In FIG. 7I, the last target display frame is F8, and in the process of obtaining the reconstructed image of F8, if the decoding mode of the prior art is followed, F0-F8, i.e., 9 image frames, need to be decoded. If F4 is determined to be the prediction display frame in the process of determining the last target display frame, the reconstructed image of F4 can be output to the video RAM. If the distance threshold is 1, when the current determined target display frame is any one of F3, F4 or F5, the third type of path can be determined as = {Video RAM, 4}, wherein “4” is the frame identifier; it should be understood that the third type of path here is only an example.

[0407] If F4 and F2 are determined to be the prediction display frame in the process of determining the last target display frame, the reconstructed images of F4 and F2 can be output to the video RAM. If the distance threshold is 1, when the current determined target display frame is F3, reading path 1 can be determined as = {Video RAM, 4}, wherein “4” is the frame identifier; or reading path 2 can be determined as = {Video RAM, 2}; it should be understood that the second reading path here is only an example. In this case, two reading paths can be determined.

[0408] In summary, the decoding cost of any one of the first decoding path, the second decoding path, the third decoding path, the fourth decoding path, the second type path, and the third type path is less than the decoding cost of the fifth decoding path.

[0409] S705, obtain the video playback frame rate and / or the freezing information in the preset time length.

[0410] It should be noted that the video playback frame rate in S705 is not the original playback frame rate of the video data, but the playback frame rate of the target display frame.

[0411] Exemplarily, the predicted display time of the target display frame can be determined first; then, in one possible manner, the predicted display time of the target display frame determined this time and the actual display time (or the predicted display time) of the plurality of target display frames determined historically can be used to determine the number of target image frames (target image frames in the time range [the predicted display time of the target display frame determined this time - the preset time length, the predicted display time of the target display frame determined this time]) in the preset time length; and then, the number of target display frames in the preset time length is divided by the preset time length, and the video playback frame rate in the preset time length can be obtained.

[0412] Exemplarily, the unpackaging module, the decoder, and the display module in the video playback process can be processed in parallel, which can improve the efficiency. In the fixed speed or variable speed playback scene, the unpackaging time consumption is much lower than the decoding time consumption, and the display time consumption is also less than the decoding time consumption (because the number of display frames is less than the number of decoding frames); therefore, the predicted display time of the target display frame when the reconstructed image of the target display frame is obtained according to the first type path can be calculated according to the decoding time consumption.

[0413] Exemplarily, the predicted display time of the target display frame when the reconstructed image of the target display frame is obtained according to the first type path can refer to the following formula (5): ToDispTime=DecFinishTime=lastDecFinishTime+m*AvgDec Cos t (5)

[0414] Wherein, ToDispTime is the predicted display time of the target display frame when the reconstructed image of the target display frame is obtained according to the first type path, DecFinishTime is the decoding completion time of the reconstructed image of the target display frame when the reconstructed image of the target display frame is obtained according to the first type path, lastDecFinishTime is the time of completing decoding of the target display frame determined last time (or the actual display time of the target display frame determined last time), and AvgDec Cos t is the average decoding time consumption, and m is the number of to-be-decoded image frames.

[0415] For example, AvgDec Cos t can refer to the number of image frames that can be decoded per unit time (i.e., decoding capability).

[0416] One possible approach is to determine AvgDec Cos t based on the decoder's decoding throughput specifications. For example, if the decoder's decoding throughput specifications are 1080p@200fps (meaning 200 frames can be decoded per second of 1080p video), then the average decoding time can be determined to be AvgDec Cos t = 5ms.

[0417] One possible approach is to statistically analyze the start time decIn1 and end time decOut of the decoded n frames. n Then, calculate AvgDec Cos t using the following formula (6):

[0418] One possible approach is to divide the decoding time into one first-frame decoding delay Dec Cos t1 and n-1 average decoded frame intervals DecInterval. i Where DecCos t1 refers to the difference between the time from the input of the first frame to the decoder and the time from the output of the first frame by the decoder; i is a positive integer greater than or equal to 2, i is less than or equal to n, and DecInterval i Let AvgDec Cos t be the difference between the time when the decoder outputs the i-th image frame and the time when it outputs the (i-1)-th image frame. Then, calculate AvgDec Cos t using the following formula (7):

[0419] For example, for target display frames in the second and third types of paths, the predicted display time of the target display frame can be the average rendering time. For instance, the rendering time of each displayed target display frame can be counted; the average rendering time can be obtained by calculating the average rendering time of multiple displayed target display frames.

[0420] For example, when multiple predicted display frames are determined, the predicted display times of these multiple predicted display frames can also be determined. In one possible approach, the number of image frames (including the target display frame and the predicted display frames) within a preset duration can be determined based on the predicted display time of the target display frame and the predicted display times of the multiple predicted display frames. Then, dividing the number of image frames within the preset duration by the preset duration yields the video playback frame rate within the preset duration. It should be understood that the predicted display time of each predicted display frame is calculated in the same way as the predicted display time of the target display frame, and will not be elaborated upon here.

[0421] Exemplarily, the stalling information comprises at least one of a stalling frequency in the preset time length, a stalling duration in the preset time length, or a variable speed range in the preset time length.

[0422] The stalling frequency in the preset time length can refer to a ratio of a number of stallings in the preset time length to the preset time length. The stalling can refer to a display interval of adjacent two image frames exceeding a threshold (e.g., 100 ms). The stalling duration in the preset time length can refer to a sum of durations of all stallings. The variable speed range in the preset time length can refer to a difference between a maximum playing speed and a minimum playing speed.

[0423] S706, determine first evaluation information corresponding to each candidate path in the plurality of candidate paths based on the video playing frame rate in the preset time length and / or the stalling information.

[0424] Exemplarily, the first evaluation information is used to describe a video watching experience, for example, to describe a smoothness of the video.

[0425] In a possible manner, the first evaluation information can be determined based on only the video playing frame rate in the prediction time length. For example, the first evaluation information can be determined with reference to the following formula (8): View=f(frameRate) (8)

[0426] Wherein, f() is a function of evaluating a subjective score according to the video playing frame rate, and frameRate is the video playing frame rate in the preset time length.

[0427] In a possible manner, the first evaluation information can be determined based on only the stalling information in the prediction time length. For example, the first evaluation information can be determined with reference to the following formula (9): View=g(stallingFrequent,stallingDuration,variableSpeed) (9)

[0428] Wherein, g() is a function of evaluating a subjective score according to the stalling information, stallingFrequent is the stalling frequency in the preset time length, stallingDuration is the stalling duration in the preset time length, and variableSpeed is the variable speed range in the preset time length.

[0429] In a possible manner, the first evaluation information can be determined based on the video playing frame rate in the prediction time length and the stalling information. For example, the first evaluation information can be determined with reference to the following formula (10): View=f(frameRate)-g(stallingFrequent,stallingDuration,variableSpeed) (10)

[0430] Exemplarily, when the video playing scene is the drag playing or the swing playing scene, the evaluation information can further include second evaluation information for describing the interactive experience; wherein the second evaluation information can be determined according to S707-S708 as follows:

[0431] S707, obtaining an interactive start response delay and an interactive end response delay.

[0432] Exemplarily, the interactive experience can be understood as whether the interaction is in time with the hand, which can mainly consider whether the response of the start point and the end point is in time.

[0433] Exemplarily, the interactive start response delay can refer to the difference between the time of the start point detected by the mobile phone (such as the time when the user starts to drag or swing) and the actual display time of the target display frame corresponding to the start point. The shorter the interactive start response delay is, the faster the response of the start point is.

[0434] Exemplarily, the interactive end response delay can refer to the difference between the time of the end point detected by the mobile phone (such as the time when the user ends to drag or swing) and the actual display time of the target display frame corresponding to the end point. The shorter the interactive end response delay is, the faster the response of the end point is.

[0435] S708, determining the second evaluation information corresponding to each candidate path in the plurality of candidate paths based on the interactive start delay and the interactive end delay.

[0436] For example, the second evaluation information can refer to the following formula (11): interaction = h(startDelay, endDelay) (11)

[0437] Wherein, h() is a function for evaluating the subjective score according to the interactive response delay, interaction is the second evaluation information, startDelay is the interactive start response delay, and endDelay is the interactive end response delay.

[0438] Exemplarily, whether the end point position is accurate and the average moving delay will also affect the interactive experience. Therefore, the application embodiment can further obtain at least one of the accuracy value of the end point or the average moving delay; then, the second evaluation information can be determined in combination with at least one of the accuracy value of the end point or the average moving delay, and the interactive start delay and the interactive end delay.

[0439] The accuracy value of the end point can be represented by the frame distance between the target display frame corresponding to the end point and the actually displayed image frame. The average moving delay can be determined as follows: calculating the difference between the time of the detected user finger position and the actual display time of the target display frame corresponding to the user finger position to obtain a moving delay; and calculating the average of all moving delays in the interaction process to obtain the average moving delay.

[0440] For example, the second evaluation information can refer to the following formula (12):

[0441] S709, the first evaluation information and the second evaluation information corresponding to each candidate path are weighted and calculated according to the first weight and the second weight to obtain the weighted result corresponding to each candidate path.

[0442] For example, the first weight is the weight of the first evaluation information, and the second weight is the weight of the second evaluation information; wherein the first weight and the second weight can be set according to the demand, for example, the user pays more attention to the video watching experience, and the first weight can be set to be greater than the second weight; for example, the user pays more attention to the interaction experience, and the second weight can be set to be greater than the first weight.

[0443] For example, the weighted calculation can be performed with reference to the following formula (13):

[0444] In formula (13), a is the first weight, and (1-a) is the second weight.

[0445] S710, selecting a candidate path with the largest weighted result from the plurality of candidate paths to determine the target path.

[0446] For example, the larger the weighted result obtained in S709, the better the video playing experience; and then, a candidate path with the largest weighted result can be selected from the plurality of candidate paths to determine the target path.

[0447] It should be understood that S707-S710 are optional steps; when S707-S710 are not executed, the evaluation information, i.e. the first evaluation information (View), can select the candidate path with the largest View to determine the target path.

[0448] In one possible implementation, the implementation of S704-S710 can be that, after determining each candidate path, S704 can execute S705-S709 to determine the evaluation information of the candidate path, and then execute S710; the execution of S710 can include: judging whether the evaluation information of the candidate path is less than the evaluation threshold; when the evaluation information of the candidate path is less than the evaluation threshold, S704 can be executed to continue to determine the next candidate path; when the evaluation information of the candidate path is greater than or equal to the evaluation threshold, the candidate path can be determined as the target path; and the above process can be repeated.

[0449] For example, one possible implementation of S704 can be that, within a certain time length (which can be set in advance according to requirements), a plurality of candidate paths for obtaining the reconstructed image of the target display frame are determined; and then S705-S710 are executed to determine the target path.

[0450] For example, one possible implementation of S704 can be that G (G can be set in advance according to requirements, and G is a positive integer) candidate paths for obtaining the reconstructed image of the target display frame are determined; and then S705-S710 are executed to determine the target path.

[0451] For example, one possible implementation of S704 can be that one candidate path for obtaining the reconstructed image of the target display frame is determined; then S705-S709 are executed to determine the weighting result corresponding to the candidate path; and then it is judged whether the weighting result corresponding to the candidate path is greater than the weighting threshold; if the weighting result corresponding to the candidate path is greater than the weighting threshold, the candidate path is determined as the target path; if the weighting result corresponding to the candidate path is less than the weighting threshold, S704 can be executed to obtain another candidate path for obtaining the reconstructed image of the target display frame; and then the above process can be repeated; in this way, the target path can also be obtained.

[0452] S711, when the target path is the first type of path, determining the to-be-decoded image frame based on the target path.

[0453] For example, when the target path is the first type of path, the target path can include one or more frame identifiers of the to-be-decoded image frame; and then the to-be-decoded image frame can be determined according to the frame identifier in the target path; and one frame identifier in the target path can be used to determine one to-be-decoded image frame.

[0454] For example, when the target path is the first decoding path {0, 3, 6, 7, 8} determined in FIG. 7D, the to-be-decoded image frame can be F0, F3, F6, F7 and F8.

[0455] For example, the target path is the second decoding path {0, 1, 2, 3, 4, 5} determined in FIG. 7E, and the image frames to be decoded can be F0, F1, F2, F3, F4 and F5.

[0456] S712, decoding the part of the code stream corresponding to the image frame to be decoded to obtain the reconstructed image of the target display frame.

[0457] S713, when the target path is the second type of path, reading the reconstructed image of the image frame corresponding to the frame identifier in the second type of path from the memory in the decoder as the reconstructed image of the target display frame.

[0458] S714, when the target path is the third type of path, reading the reconstructed image of the image frame corresponding to the frame identifier in the third type of path from the memory outside the decoder as the reconstructed image of the target display frame.

[0459] S715, playing the reconstructed image of the target display frame.

[0460] In this way, in the drag play scene or the flick play scene, the video is played according to the video playing process 700, which not only ensures that the video is relatively smooth, but also ensures that the video is relatively interactive.

[0461] For example, for the video file 15 containing decoding dependency information (as shown in FIG. 3A), the above S701-S714 can be executed by the video player 310 in FIG. 3A to play the reconstructed image of the target display frame. For the video file 21 not containing decoding dependency information (as shown in FIG. 3B), S701-S714 can be executed by the video player 310 in FIG. 3B to play the reconstructed image of the target display frame. However, the way of determining decoding dependency information in the process of S704 executed by the video player 310 in FIG. 3B is different from the way of determining decoding dependency information shown in the above video playing process 700 (i.e., the way of determining decoding dependency information in the process of S704 executed by the video player 310 in FIG. 3A). The way of determining decoding dependency information by the video player 310 in FIG. 3B can refer to the following decoding dependency information determination process 800.

[0462] FIG. 8A is a schematic diagram of a decoding dependency information determination process 800 provided by an embodiment of the present application.

[0463] S801, unpacking the video file to obtain a code stream.

[0464] S802, parsing the code stream to determine the reference relationship of each image frame in the plurality of image frames.

[0465] The following takes the first image frame (any one of the plurality of image frames) and the code stream as an example of H265 code stream to illustrate the process of determining the reference relationship of the first image frame.

[0466] FIG. 8B is a schematic diagram of a reference relationship determination process according to an embodiment of the present application.

[0467] S802 can include the following steps S8021-S8025.

[0468] S8021, parse the long-term reference frame, the short-term reference frame and the rearrangement information of the reference picture list of the first image frame from the code stream.

[0469] FIG. 8C is a schematic diagram of a code stream structure according to an embodiment of the present application.

[0470] For example, the code stream of H265 can include a plurality of network abstraction layer (Network Abstraction Layer, NAL) units, each NAL unit can include a NAL header and a NAL body. The NAL body in different NAL units can be used to carry different information.

[0471] For example, referring to FIG. 8C, the NAL body is used to carry the sequence parameter set (Sequence Parameter Set, SPS), at this time, the NAL body includes the raw byte sequence payload (Raw Byte Sequence Payload, RBSP) of the SPS (i.e. SPS_RBSP). For example, the NAL body is used to carry the picture parameter set (Picture Paramater Set, PPS), at this time, the NAL body includes PPS_RBSP. For example, the NAL body is used to carry the video parameter set (Video Parameter Set, VPS), at this time, the NAL body includes VPS_RBSP. For example, the NAL body is used to carry the original image in the video data, at this time, the NAL body includes the Slice_header and Slice_RBSP of a slice (one original image can be divided into at least one slice). If the original image in the video data includes N (N is a positive integer) slices, the original image in the video data can be carried by the NAL body of N NAL units. The NAL body in the first NAL unit of the N NAL units can include the Slice 1_header and Slice1_RBSP of a slice (i.e. Slice 1); …; the NAL body in the Nth NAL unit can include the SliceN_header and Slice N_RBSP of a slice (Slice N).

[0472] For example, the st_ref_pic_set (short-term reference picture set) and other related fields in the SPS_RBSP and Slice Header (corresponding to the first image frame) can be parsed to obtain the short-term reference picture set (including one or more short-term reference frames) of the first image frame.

[0473] For example, the lt_ref_pic_poc_lsb_sps (used to indicate the least significant bits of the picture order count (POC) of the long-term reference picture), used_by_curr_pic_lt_sps_flag (used to indicate whether the long-term reference picture defined in the SPS is used in the reference picture list of the current picture), and other related fields can be parsed to obtain the long-term reference picture set (including one or more long-term reference frames) of the first image frame.

[0474] For example, the ref_pic_list_modification (a process in H.264 / AVC and H.265 / HEVC video coding standards, which involves modification of the reference picture list during decoding), list_entry (an entry in the reference picture list, which can include the reference picture index, reference picture flag, and POC of the reference picture), and other related fields in the Slice Header (corresponding to the first image frame) can be parsed to obtain the rearrangement information of the reference picture list of the first image frame.

[0475] S8022, based on the long-term reference frame, short-term reference frame, and rearrangement information of the reference picture list of the first image frame, determine the reference picture list of the first image frame.

[0476] For example, the initial reference picture list of the first image frame can be constructed based on the short-term reference picture set of the first image frame and the long-term reference picture set of the first image frame, following the principle of short-term reference frames first and long-term reference frames last. Then, the initial reference picture list of the first image frame can be reordered based on the rearrangement information of the reference picture list of the first image frame to obtain the reference picture list of the first image frame.

[0477] S8023, parse the number S of reference frames of the first image frame from the code stream.

[0478] Then, the num_ref_idx_active_minus1 (used to indicate the number of active reference picture indexes in the current slice) and other related fields in the PPS and Slice Header (corresponding to the first image frame) can be parsed to obtain the number S of reference pictures of the first image frame (S is a positive integer).

[0479] S8024, selecting the first S image frames from the reference image list of the first image frame as the reference frames of the first image frame.

[0480] Subsequently, selecting the first S image frames from the reference image list of the first image frame as the reference frames of the first image frame.

[0481] S8025, generating the reference relationship of the first image frame based on the frame identifier of the first image frame and the frame identifier of the reference frame of the first image frame.

[0482] For example, the frame identifier of the first image frame is 1, the reference frame of the first image frame is 1 image frame (the frame identifier of the image frame is 0), and the reference relationship of the first image frame is {(0), 1} (the representation of the reference relationship is only an example, and does not represent the actual reference relationship in this way).

[0483] For example, the frame identifier of the first image frame is 2, the reference frame of the first image frame is 2 image frames (the frame identifiers of the 2 image frames are 0 and 1 respectively), and the reference relationship of the first image frame is {(0, 1), 2}.

[0484] For example, the frame identifier of the first image frame is 4, the reference frame of the first image frame is 2 image frames (the frame identifiers of the 2 image frames are 4 and 5 respectively), and the reference relationship of the first image frame is {(4, 5), 6}.

[0485] FIG. 8D is a schematic diagram of a reference relationship according to an embodiment of the present application.

[0486] In FIG. 8D, the reference frame of F1 is F0, the reference frame of F2 is F0 and F1, the reference frame of F3 is F0 and F1. The reference frame of F4 is F0, the reference frame of F5 is F4 and F0, the reference frame of F6 is F4 and F5, the reference frame of F7 is F4 and F5. The reference frame of F8 is F4, the reference frame of F9 is F8, the reference frame of F10 is F8 and F9, and the reference frame of F11 is F8 and F9.

[0487] S803, determining the decoding dependency information of each image frame in the plurality of image frames according to the reference relationship of the plurality of image frames.

[0488] Assuming that the definition of the decoding dependency information is as shown in Table 1, the decoding dependency information of the 12 image frames determined based on the reference relationship in FIG. 8D can be as follows:

[0489] F0, F4, and F8 are key reference frames, and the corresponding decoding dependency information is the first dependency level; F1, F5, and F9 are general reference frames, and the corresponding decoding dependency information is the second dependency level; F2, F3, F6, F7, F10, and F11 are non-reference frames, and the corresponding decoding dependency information is the third dependency level.

[0490] FIG. 9 is a schematic diagram of a video recording process 900 according to an embodiment of the present application. The video recording process 900 can be performed by the video recorder 300 in FIG. 3A.

[0491] S901, obtaining video data, the video data comprising original images of a plurality of image frames.

[0492] For example, the video data can be captured by a camera; for another example, the video data can be obtained by a screen recording tool; and the like.

[0493] S902, encoding the original images of the plurality of image frames to obtain a bitstream.

[0494] For example, the video data can be input into an encoder, and the original images of the plurality of image frames in the video data can be encoded by the encoder to output the bitstream.

[0495] S903, obtaining decoding dependency information corresponding to each image frame in the plurality of image frames, the decoding dependency information of a first image frame indicating a number of second image frames that depend on decoding of the first image frame, the first image frame being any image frame in the plurality of image frames, and the second image frames being image frames after the first image frame in a group of pictures to which the first image frame belongs.

[0496] In one possible implementation, the decoding dependency information generation module 303 can obtain encoding information of the video data encoded by the encoder 302, wherein the encoding information can comprise reference relationships of each image frame. Then, the decoding dependency information generation module 303 can determine the decoding dependency information corresponding to each image frame based on the reference relationships of each image frame in the encoding information.

[0497] For example, for a first image frame (any image frame in the plurality of image frames), the number of second image frames that depend on decoding (or encoding) of the first image frame can be determined according to the reference relationships of each image frame in a group of pictures to which the first image frame belongs. Then, the decoding dependency information of the first image frame can be determined based on the number of the second image frames that depend on decoding (or encoding) of the first image frame. The image frames after the first image frame in the group of pictures to which the first image frame belongs can be referred to as the second image frames.

[0498] For example, referring to the explanation of the decoding dependency information in Table 1, when all the second image frames after the first image frame are decoded in dependence on the first image frame, the decoding dependency information of the first image frame can be determined as the first dependency level. When part of the second image frames after the first image frame are decoded in dependence on the first image frame, the decoding dependency information of the first image frame can be determined as the second dependency level. When no second image frame after the first image frame is decoded in dependence on the first image frame, the decoding dependency information of the first image frame can be determined as the third dependency level.

[0499] In other words, the first dependency level of the first image frame indicates that all the second image frames after the first image frame are decoded in dependence on the first image frame; the second dependency level of the first image frame indicates that part of the second image frames after the first image frame are decoded in dependence on the first image frame; and the third dependency level of the first image frame indicates that there is no second image frame after the first image frame that is decoded in dependence on the first image frame.

[0500] For example, referring to the explanation of the decoding dependency information in Table 2, when all the second image frames after the first image frame are decoded in dependence on the first image frame, and the first image frame is an I frame, the decoding dependency information of the first image frame can be determined as the first sub-level. When all the second image frames after the first image frame are decoded in dependence on the first image frame, and the first image frame is a non-I frame, the decoding dependency information of the first image frame can be determined as the second sub-level.

[0501] For example, referring to the explanation of the decoding dependency information in Table 2, when part of the second image frames after the first image frame are decoded in dependence on the first image frame, and the number of the second image frames decoded in dependence on the first image frame is greater than a number threshold, the decoding dependency information of the first image frame can be determined as the third sub-level. When part of the second image frames after the first image frame are decoded in dependence on the first image frame, and the number of the second image frames decoded in dependence on the first image frame is less than or equal to the number threshold, the decoding dependency information of the first image frame can be determined as the fourth sub-level.

[0502] It should be understood that the third sub-level and the fourth sub-level can be further subdivided, which can be determined according to application scenarios, and the present application does not limit this.

[0503] In other words, the first sub-level of the first image frame indicates that all the second image frames after the first image frame are decoded depending on the first image frame, and the first image frame is an I frame; the second sub-level of the first image frame indicates that all the second image frames after the first image frame are decoded depending on the first image frame, and the first image frame is a non-I frame. The third sub-level of the first image frame indicates that part of the second image frames after the first image frame are decoded depending on the first image frame, and the number of the second image frames decoded depending on the first image frame is greater than the number threshold; the fourth sub-level of the first image frame indicates that part of the second image frames after the first image frame are decoded depending on the first image frame, and the number of the second image frames decoded depending on the first image frame is less than or equal to the number threshold.

[0504] In a possible manner, the decoding dependency information generation module 303 can parse the code stream to determine the reference relationship of each image frame in the plurality of image frames; then, based on the reference relationship of all the image frames in the plurality of image frames, determine the decoding dependency information of each image frame in the plurality of image frames. Wherein, the process of determining the reference relationship of each image frame can refer to the description of S8021-S8025 above, which will not be repeated here.

[0505] S904, generating a video file based on the decoding dependency information of each image frame in the plurality of image frames and the code stream.

[0506] In a possible manner, the decoding dependency information generation module 303 can write the decoding dependency information of each image frame in the plurality of image frames into the code stream; then, the video encapsulation module 304 encapsulates the code stream to obtain a video file. For example, the level identifier corresponding to the decoding dependency information of each image frame can be written into the code stream. For example, when the decoding dependency information of the image frame is the first dependency level, "1" can be written into the code stream; when the decoding dependency information of the image frame is the second dependency level, "2" can be written into the code stream; when the decoding dependency information of the image frame is the third dependency level, "3" can be written into the code stream.

[0507] For example, taking the code stream of H265 as an example. The level identifier corresponding to the decoding dependency information of each image frame can be written into the Slice_header of each image frame. When one image frame corresponds to one Slice, the level identifier corresponding to the decoding dependency information of the corresponding image frame can be included in each Slice_header. When one image frame corresponds to multiple Slices, the level identifier corresponding to the decoding dependency information of the image frame can be written into the Slice_header of the first Slice corresponding to the image frame; of course, the level identifier corresponding to the decoding dependency information of the image frame can also be written into the Slice_header of each Slice corresponding to the image frame.

[0508] It should be understood that when the decoding dependency information of the image frame is the frame identifier, the frame identifier of the image frame on which the decoding of the image frame depends can be written into the bitstream.

[0509] In a possible manner, the decoding dependency information of each of the plurality of image frames and the bitstream can be encapsulated by the video encapsulation module 304 to obtain a video file.

[0510] For example, taking an MP4 file as an example. The MP4 file can include a plurality of types of BOX (sample container / box), such as a moov BOX and an mdat BOX.

[0511] The moov BOX, that is, a Movie Box, contains metadata (such as encoding parameters of video and audio, timestamps, sample descriptions, sample-to-time mapping, and the like) of all media data in the MP4 file.

[0512] The mdat BOX, that is, a Media Data Box, contains encoded video data (that is, the bitstream described above).

[0513] In a possible manner, the existing syntax element in the moov BOX of the MP4 file can be reused, and the existing syntax element can be used to indicate the decoding dependency information of the image frame.

[0514] For example, the syntax element sample_is_depended_on in the moov BOX is reused. Exemplarily, the semantics of different values of sample_is_depended_on can be extended according to the meaning of the decoding dependency information in the present application, and the semantics after extension (the extended semantics include original semantics and extended semantics) is obtained. The semantics of sample_is_depended_on can be as shown in Table 4:

[0515] Table 4

[0516] In Table 4, the sample can refer to the image frame.

[0517] In a possible manner, a new syntax element can be added in the moov BOX, and the new syntax element can be used to indicate the decoding dependency information of the image frame. The moov BOX can be added in a new BOX (such as a decoding_priority box), and then the new syntax element can be added to the new BOX.

[0518] It should be understood that when the decoding dependency information of the image frame is the frame identifier, the frame identifier of the image frame on which the decoding of the image frame depends and the code stream can be encapsulated to obtain the video file.

[0519] Compared with carrying the decoding dependency information in the code stream, carrying the decoding dependency information in the BOX of the MP4 can enable the hierarchical structure (decoding dependency information) in the whole or a section of the code stream to be obtained in the transmission process, so that the transmission frame loss ratio can be controlled as a whole. In addition, the hierarchical structure (decoding dependency information) in the whole or a section of the code stream can be known earlier in the video playback process, so that the decoding frame loss ratio can be controlled as a whole, and the second image frame and the fourth image frame with the minimum global cost can be accurately selected.

[0520] FIG. 10 is a schematic diagram of a video conversion process 1000 according to an embodiment of the present application. A server or the like can convert a video file not carrying decoding dependency information into a video file carrying decoding dependency information, so that an electronic device without the decoding dependency information generation module 316 can perform the video playback process 700 to realize video playback. The video conversion process 1000 can be performed by the video converter 400 in FIG. 4.

[0521] S1001, a first video file is obtained, the first video file including a code stream, the code stream being obtained from original images of a plurality of image frames in encoded video data.

[0522] For example, the first video file does not include decoding dependency information of any image frame. For example, the first video file is the first video file 31 in FIG. 4.

[0523] S1002, the first video file is unpacked to obtain the code stream.

[0524] S1003, the code stream is parsed to determine a reference relationship of each image frame in the plurality of image frames.

[0525] S1004, decoding dependency information of each image frame in the plurality of image frames is determined according to the reference relationship of all image frames in the plurality of image frames.

[0526] For example, S1003-S1004 can refer to the description of S802-S803 above, and will not be described here again.

[0527] S1005, a second video file is generated based on the decoding dependency information of each image frame in the plurality of image frames and the code stream.

[0528] In a possible implementation, the decoding dependency information generation module 402 can write the decoding dependency information of each of the plurality of image frames into the code stream; and then the video encapsulation module 403 encapsulates the code stream to obtain the second video file (e.g., the second video file 34 in FIG. 4).

[0529] In a possible implementation, the video encapsulation module 403 can encapsulate the decoding dependency information of each of the plurality of image frames and the code stream to obtain the second video file (e.g., the second video file 34 in FIG. 4).

[0530] FIG. 11 is a structural schematic diagram of a video playing apparatus 1100 according to an embodiment of the present application. The video playing apparatus 1100 can be used to execute the method according to the foregoing embodiments, and therefore, the beneficial effects that can be achieved by the video playing apparatus 1100 can refer to the beneficial effects provided in the corresponding method, which will not be described herein again.

[0531] For example, the video playing apparatus 1100 includes:

[0532] The first obtaining module 1101 is configured to obtain a video file, where the video file includes a code stream, and the code stream includes encoded data obtained by encoding a plurality of image frames in video data.

[0533] The display frame determination module 1102 is configured to determine a target display frame from the plurality of image frames.

[0534] The target path determination module 1103 is configured to select one candidate path from the plurality of candidate paths as a target path based on evaluation information corresponding to each of the plurality of candidate paths, where the evaluation information is used to describe a video playing experience.

[0535] The second obtaining module 1104 is configured to obtain a reconstructed image of the target display frame based on the target path and the code stream.

[0536] The display module 1105 is configured to play the reconstructed image of the target display frame.

[0537] It should be noted that the target path determination module 1103 belongs to the video playing experience evaluation module 312 described above. The second obtaining module 1104 can include the decoder 313 described above.

[0538] For example, at least one candidate path in the plurality of candidate paths is different from the decoding cost of the other candidate paths.

[0539] For example, the candidate path includes at least one of a first type of path, a second type of path, or a third type of path.

[0540] The first type of path is used to indicate one or more image frames that need to be decoded to obtain the reconstructed image of the target display frame.

[0541] The second type of path is used to indicate that the reconstructed image of the target display frame is read from the memory within the decoder;

[0542] The third type of path is used to indicate that the reconstructed image of the target display frame is read from the memory outside the decoder.

[0543] Exemplarily, the first type of path includes a first decoding path, one of the plurality of candidate paths is the first decoding path, and the video playing apparatus 1100 further includes:

[0544] The candidate path determination module is configured to: acquire, based on the video file, decoding dependency information of each image frame in the image group to which the target display frame belongs; and determine the first decoding path based on the decoding dependency information of each image frame in the image group to which the target display frame belongs.

[0545] Exemplarily, the candidate path determination module is specifically configured to: determine one or more first image frames based on the decoding dependency information of each image frame in the image group to which the target display frame belongs, the first image frame being an image frame on which the target display frame depends; and generate the first decoding path by using a frame identifier of the one or more first image frames and a frame identifier of the target display frame.

[0546] Exemplarily, the first type of path includes a second decoding path, one of the plurality of candidate paths is the second decoding path, and the candidate path determination module is configured to: acquire a second image frame, the decoding cost of the second image frame being less than the decoding cost of the target display frame; acquire one or more third image frames when the second image frame is not an I frame, the third image frame being before the second image frame, and the number of the third image frames being less than or equal to the total number of all image frames before the second image frame in the image group to which the second image frame belongs; generate the second decoding path by using a frame identifier of the one or more third image frames and a frame identifier of the second image frame; and the reconstructed image of the target display frame being a reconstructed image of the second image frame.

[0547] Exemplarily, the candidate path determination module is further configured to: when the second image frame is an I frame, generate the second decoding path by using a frame identifier of the second image frame.

[0548] Exemplarily, the video playing apparatus 1110 further includes:

[0549] The third acquisition module is configured to determine one or more predicted display frames from the plurality of image frames based on a user operation and the target display frame.

[0550] The first type of path includes a third decoding path, one of the plurality of candidate paths is the third decoding path, the candidate path determination module is used for obtaining a group of image frames, the group of image frames includes a fourth image frame and one or more fifth image frames, the number of the fifth image frames is the same as the number of the predicted display frames, and the decoding cost of the group of image frames is less than the sum of the decoding costs of the predicted display frames and the target display frame; one or more sixth image frames are obtained, the sixth image frames are before the fourth image frame, and the number of the sixth image frames is less than or equal to the total number of all image frames before the fourth image frame in the image group to which the fourth image frame belongs; the third decoding path is generated by using the frame identifiers of the one or more sixth image frames and the frame identifier of the fourth image frame; and the reconstructed image of the target display frame is the reconstructed image of the fourth image frame.

[0551] Exemplarily, the first type of path includes a fourth decoding path, one of the plurality of candidate paths is the fourth decoding path, and the candidate path determination module is specifically used for determining the decoded image frame closest to the target display frame in the memory in the decoder; the fourth decoding path is generated by using the frame identifiers of one or more seventh image frames and the frame identifier of the target display frame, the seventh image frame being an image frame after the decoded image frame closest to the target display frame and before the target display frame.

[0552] Exemplarily, the second obtaining module 1104 is specifically used for, when the target path is the first type of path, determining the to-be-decoded image frame based on the target path; and decoding the part corresponding to the to-be-decoded image frame in the code stream to obtain the reconstructed image of the target display frame.

[0553] Exemplarily, the video file further includes decoding dependency information of each image frame in the plurality of image frames, and the candidate path determination module is used for obtaining, from the video file, the decoding dependency information of each image frame in the image group to which the target display frame belongs.

[0554] Exemplarily, the code stream further includes decoding dependency information of each image frame in the plurality of image frames, and the candidate path determination module is used for obtaining the code stream by decapsulating the video file; and obtaining, from the code stream, the decoding dependency information of each image frame in the image group to which the target display frame belongs.

[0555] Exemplarily, the candidate path determination module is used for obtaining the code stream by decapsulating the video file; parsing the code stream to determine the reference relationship of each image frame in the image group to which the target display frame belongs; and determining the decoding dependency information of each image frame in the image group to which the target display frame belongs according to the reference relationship of all the image frames in the image group to which the target display frame belongs.

[0556] Exemplarily, the evaluation information includes first evaluation information, the first evaluation information is used for describing the video watching experience; and the video playing device 1100 further includes:

[0557] The evaluation information determination module is configured to acquire a video playback frame rate and / or a freezing information within a preset time length; and determine first evaluation information corresponding to each candidate path in the plurality of candidate paths based on the video playback frame rate and / or the freezing information within the preset time length.

[0558] For example, the freezing information within the preset time length includes at least one of a freezing frequency, a freezing time length, or a variable speed amplitude within the preset time length.

[0559] For example, the evaluation information further includes second evaluation information, the second evaluation information is used to describe an interactive experience; the evaluation information determination module is further configured to acquire an interactive start delay and an interactive end delay; and determine the second evaluation information corresponding to each candidate path in the plurality of candidate paths based on the interactive start delay and the interactive end delay.

[0560] For example, the evaluation information determination module is specifically configured to perform a weighted calculation on the first evaluation information and the second evaluation information corresponding to each candidate path according to a first weight and a second weight, to obtain a weighted result corresponding to each candidate path; and select a candidate path with a maximum weighted result from the plurality of candidate paths as the target path.

[0561] For example, the display frame determination module 1102 is specifically configured to determine a target display frame from the plurality of image frames based on a user operation, the user operation including one of a fixed speed playback operation or a variable speed playback operation.

[0562] FIG. 12 is a structural schematic diagram of a video recording device 1200 according to an embodiment of the present application. The video recording device 1200 can be used to execute the method of the foregoing embodiments, and thus the beneficial effects that can be achieved thereby can refer to the beneficial effects provided in the corresponding method, which will not be described herein again.

[0563] For example, the video recording device 1200 includes:

[0564] The video data acquisition module 1201 is configured to acquire video data, the video data including original images of a plurality of image frames;

[0565] The encoding module 1202 is configured to encode the original images of the plurality of image frames to obtain a bitstream;

[0566] The decoding dependency information generation module 1203 is configured to acquire decoding dependency information corresponding to each image frame in the plurality of image frames, the decoding dependency information of a first image frame being used to indicate a number of second image frames that depend on decoding of the first image frame, the first image being any one of the plurality of image frames, and the second image frame being an image frame after the first image frame in a group of pictures to which the first image frame belongs;

[0567] The video file generation module 1204 is configured to generate a video file based on the decoding dependency information of each image frame in the plurality of image frames and the bitstream.

[0568] For example, the decoding dependency information includes any one of a first dependency level, a second dependency level, or a third dependency level.

[0569] The first dependency level of the first image frame indicates that all the second image frames after the first image frame are decoded in dependence on the first image frame.

[0570] The second dependency level of the first image frame indicates that part of the second image frames after the first image frame are decoded in dependence on the first image frame.

[0571] The third dependency level of the first image frame indicates that there is no second image frame after the first image frame that is decoded in dependence on the first image frame.

[0572] For example, the first dependency level includes any one of a first sub-level or a second sub-level.

[0573] The first sub-level of the first image frame indicates that all the second image frames after the first image frame are decoded in dependence on the first image frame, and the first image frame is an I frame.

[0574] The second sub-level of the first image frame indicates that all the second image frames after the first image frame are decoded in dependence on the first image frame, and the first image frame is a non-I frame.

[0575] For example, the second dependency level includes any one of a third sub-level or a fourth sub-level.

[0576] The third sub-level of the first image frame indicates that part of the second image frames after the first image frame are decoded in dependence on the first image frame, and the number of the second image frames decoded in dependence on the first image frame is greater than a number threshold.

[0577] The fourth sub-level of the first image frame indicates that part of the second image frames after the first image frame are decoded in dependence on the first image frame, and the number of the second image frames decoded in dependence on the first image frame is less than or equal to the number threshold.

[0578] For example, the decoding dependency information generation module 1203 is specifically configured to parse the bitstream to obtain a reference relationship of each image frame in the plurality of image frames; and determine the decoding dependency information of each image frame in the plurality of image frames according to the reference relationship of all the image frames in the plurality of image frames.

[0579] For example, the video file generation module 1204 is specifically configured to write the decoding dependency information of each image frame in the plurality of image frames into the bitstream; and encapsulate the bitstream to obtain the video file.

[0580] The video file generation module 1204 encapsulates the decoding dependency information and the bitstream of each of the plurality of image frames to obtain a video file.

[0581] FIG. 13 is a structural schematic diagram of a video conversion apparatus 1300 according to an embodiment of the present application. The video conversion apparatus 1300 can be configured to execute the method according to the above-described embodiments, and thus can achieve the beneficial effects as described above with respect to the corresponding method. Therefore, the detailed description is omitted here.

[0582] The video conversion apparatus 1300 includes, for example:

[0583] The video file acquisition module 1301 is configured to acquire a first video file, and the first video file includes a bitstream, and the bitstream is obtained from original images of a plurality of image frames in encoded video data.

[0584] The video unencapsulation module 1302 is configured to unencapsulate the first video file to obtain the bitstream.

[0585] The decoding dependency information generation module 1303 is configured to parse the bitstream to obtain a reference relationship of each of the plurality of image frames, and determine decoding dependency information of each of the plurality of image frames according to the reference relationship of all the plurality of image frames.

[0586] The video file generation module 1304 is configured to generate a second video file based on the decoding dependency information of each of the plurality of image frames and the bitstream.

[0587] The video file generation module 1304 is configured to, for example, write the decoding dependency information of each of the plurality of image frames into the bitstream, and encapsulate the bitstream to obtain the second video file.

[0588] The video file generation module 1304 is configured to, for example, encapsulate the decoding dependency information and the bitstream of each of the plurality of image frames to obtain the second video file.

[0589] In one example, FIG. 14 shows a schematic block diagram of an apparatus 1400 according to an embodiment of the present application. The apparatus 1400 can include a processor 1401 and a transceiver 1402, and optionally further include a memory 1403.

[0590] The various components of the apparatus 1400 are coupled by a bus 1404, which can include a data bus, a power bus, a control bus, and a state signal bus. However, for the sake of clarity, the various buses are shown as the bus 1404.

[0591] Optionally, the memory 1403 can be used to store instructions in the foregoing method embodiments. The processor 1401 can be used to execute the instructions in the memory 1403, and control the transceiver 1402 to receive signals, and control the transceiver 1402 to send signals.

[0592] The apparatus 1400 can be an electronic device or a chip of an electronic device in the foregoing method embodiments. The electronic device can be a terminal device or a server.

[0593] All relevant contents of each step involved in the foregoing method embodiments can be cited to the function description of the corresponding function module, and will not be repeated here.

[0594] The embodiment of the present application further provides a chip, comprising one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits; when the one or more processors execute computer instructions, the above-mentioned related method steps are executed to realize the steps of the method in the foregoing embodiment. The interface circuit is the transceiver 1402.

[0595] The embodiment further provides a computer readable storage medium, which stores computer instructions, and when the computer instructions run on an electronic device, the electronic device executes the above-mentioned related method steps to realize the method in the foregoing embodiment. Exemplarily, the computer readable storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various storage program codes.

[0596] The embodiment further provides a computer program product, which contains computer instructions, and when the computer instructions are executed by a computer or a processor, the computer executes the above-mentioned related steps to realize the method in the foregoing embodiment. Exemplarily, the computer program product can be stored in a random access memory (RAM), a flash memory, a ROM, an erasable programmable ROM (EPROM), an electrically EPROM (EEPROM), a register, a hard disk, a mobile hard disk, a read-only compact disc (CD-ROM), or any other form of storage medium well known in the art.

[0597] The electronic device, the computer readable storage medium, the computer program product or the chip provided in the embodiment are used for executing the corresponding method provided above, and thus the beneficial effects achieved thereby can refer to the beneficial effects of the corresponding method provided above, which will not be described herein again.

[0598] Through the description of the above embodiments, those skilled in the art can understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0599] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed to multiple different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0600] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0601] Any content of each embodiment of the present application, and any content of the same embodiment, can be freely combined. Any combination of the above is within the scope of the present application.

[0602] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the above specific embodiments, and the above specific embodiments are only illustrative and not restrictive. Those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope of protection of the claims, and all of them belong to the protection of the present application.

Claims

A video playing method, characterized in that, The method comprises: obtaining a video file, wherein the video file comprises a code stream, and the code stream comprises encoded data of original images of a plurality of image frames in encoded video data; determining a target display frame from the plurality of image frames; selecting a candidate path from a plurality of candidate paths as a target path based on evaluation information corresponding to each of the plurality of candidate paths, wherein the evaluation information is used to describe video playback experience; obtaining a reconstructed image of the target display frame based on the target path and the code stream; playing the reconstructed image of the target display frame. The method of claim 1, wherein At least one of the plurality of candidate paths has a different decoding cost from other candidate paths. The method according to claim 1 or 2, characterized in that The candidate path comprises at least one of a first type of path, a second type of path, or a third type of path. The first type of path is used to indicate one or more image frames required for decoding to obtain the reconstructed image of the target display frame. The second type of path is used to indicate reading the reconstructed image of the target display frame from memory within a decoder. The third type of path is used to indicate reading the reconstructed image of the target display frame from memory outside the decoder. The method according to claim 3, characterized in that The first type of path comprises a first decoding path, and when one of the plurality of candidate paths is the first decoding path, the method further comprises: obtaining decoding dependency information of each image frame in a group of pictures to which the target display frame belongs based on the video file; determining the first decoding path based on the decoding dependency information of each image frame in the group of pictures to which the target display frame belongs. The method according to claim 4, characterized in that The determining of the first decoding path based on the decoding dependency information of each image frame in the group of pictures to which the target display frame belongs comprises: determining one or more first image frames based on the decoding dependency information of each image frame in the group of pictures to which the target display frame belongs, wherein the first image frames are image frames on which decoding of the target display frame depends; generating the first decoding path by using frame identifiers of the one or more first image frames and a frame identifier of the target display frame. The method according to claim 3, characterized in that The first type of path comprises a second decoding path, and when one of the plurality of candidate paths is the second decoding path, the method further comprises: obtaining a second image frame, wherein a decoding cost of the second image frame is less than a decoding cost of the target display frame; when the second image frame is not an I frame, obtaining one or more third image frames, wherein the third image frames are before the second image frame, and a number of the third image frames is less than or equal to a total number of all image frames before the second image frame in a group of pictures to which the second image frame belongs; and generating the second decoding path by using frame identifiers of the one or more third image frames and a frame identifier of the second image frame, wherein the reconstructed image of the target display frame is a reconstructed image of the second image frame. The method according to claim 6, characterized in that The method further comprises: when the second image frame is an I frame, generating the second decoding path by using the frame identifier of the second image frame. The method according to claim 3, characterized in that The first type of path comprises a third decoding path, and when one of the plurality of candidate paths is the third decoding path, the method further comprises: determine one or more predicted display frames from the plurality of image frames based on the user operation and the target display frame; obtain a set of image frames, the set of image frames comprising a fourth image frame and one or more fifth image frames, the number of the fifth image frames being the same as the number of the predicted display frames, the decoding cost of the set of image frames being less than the sum of the decoding cost of the predicted display frames and the target display frame; obtain one or more sixth image frames, the sixth image frames being before the fourth image frame, the number of the sixth image frames being less than or equal to the total number of all image frames before the fourth image frame in a group of pictures to which the fourth image frame belongs; generate the third decoding path using the frame identifier of the one or more sixth image frames and the frame identifier of the fourth image frame, the reconstructed image of the target display frame being the reconstructed image of the fourth image frame. The method according to claim 3, characterized in that The first type of path comprises a fourth decoding path, when a candidate path in the plurality of candidate paths is the fourth decoding path, the method further comprises: determine a decoded image frame closest to the target display frame in the memory in the decoder; generate the fourth decoding path using the frame identifier of one or more seventh image frames and the frame identifier of the target display frame, the seventh image frames being image frames after the decoded image frame closest to the target display frame and before the target display frame. The method according to any one of claims 3 to 9, characterized in that The obtaining of the reconstructed image of the target display frame based on the target path and the code stream comprises: when the target path is a first type of path, determining a to-be-decoded image frame based on the target path; decoding a part corresponding to the to-be-decoded image frame in the code stream to obtain the reconstructed image of the target display frame. The method according to claim 4, characterized in that The video file further comprises decoding dependency information of each image frame in the plurality of image frames, and the obtaining of the decoding dependency information of each image frame in the group of pictures to which the target display frame belongs based on the video file comprises: decapsulating the decoding dependency information of each image frame in the group of pictures to which the target display frame belongs from the video file. The method according to claim 4, characterized in that The code stream further comprises decoding dependency information of each image frame in the plurality of image frames, and the obtaining of the decoding dependency information of each image frame in the group of pictures to which the target display frame belongs based on the video file comprises: decapsulating the video file to obtain the code stream; parsing the decoding dependency information of each image frame in the group of pictures to which the target display frame belongs from the code stream. The method according to claim 4, characterized in that The obtaining of the decoding dependency information of each image frame in the group of pictures to which the target display frame belongs based on the video file comprises: decapsulating the video file to obtain the code stream; parsing the code stream to determine the reference relationship of each image frame in the group of pictures to which the target display frame belongs; determining the decoding dependency information of each image frame in the group of pictures to which the target display frame belongs according to the reference relationship of all image frames in the group of pictures to which the target display frame belongs. The method according to any one of claims 1 to 13, characterized in that The evaluation information comprises first evaluation information, and the first evaluation information is used to describe video watching experience; the method further comprises: acquire a video playback frame rate and / or a stall information within a preset time length; determine first evaluation information corresponding to each candidate path in the plurality of candidate paths based on the video playback frame rate and / or the stall information within the preset time length. The method of claim 14, wherein The stall information within the preset time length includes at least one of a stall frequency, a stall duration, or a variable speed amplitude within the preset time length. The method according to claim 14 or 15, characterized in that The evaluation information further includes second evaluation information for describing an interactive experience; the method further includes: acquiring an interactive start delay and an interactive end delay; determining second evaluation information corresponding to each candidate path in the plurality of candidate paths based on the interactive start delay and the interactive end delay. The method of claim 16, wherein The method of selecting a candidate path from the plurality of candidate paths as a target path based on the evaluation information corresponding to each candidate path in the plurality of candidate paths includes: performing weighted calculation on the first evaluation information and the second evaluation information corresponding to each candidate path according to a first weight and a second weight to obtain a weighted result corresponding to each candidate path; selecting a candidate path with the largest weighted result from the plurality of candidate paths as the target path. The method according to any one of claims 1 to 17, characterized in that The method of determining a target display frame from the plurality of image frames includes: determining a target display frame from the plurality of image frames based on a user operation, the user operation including one of a fixed speed playback operation or a variable speed playback operation. A video recording method, characterized by, The method includes: acquiring video data, the video data including original images of a plurality of image frames; encoding the original images of the plurality of image frames to obtain a bitstream; acquiring decoding dependency information corresponding to each image frame in the plurality of image frames, the decoding dependency information of a first image frame being used to indicate a number of second image frames that depend on decoding of the first image frame, the first image being any one of the plurality of image frames, and the second image frame being an image frame after the first image frame in a group of pictures to which the first image frame belongs; generating a video file based on the decoding dependency information of each image frame in the plurality of image frames and the bitstream. The method of claim 19, wherein The decoding dependency information includes any one of a first dependency level, a second dependency level, or a third dependency level; The first dependency level of the first image frame indicates that all the second image frames after the first image frame depend on decoding of the first image frame; The second dependency level of the first image frame indicates that part of the second image frames after the first image frame depend on decoding of the first image frame; The third dependency level of the first image frame indicates that there is no second image frame after the first image frame that depends on decoding of the first image frame. The method of claim 20, wherein The first dependency level includes any one of a first sub-level or a second sub-level; The first sub-level of the first image frame indicates that all the second image frames after the first image frame depend on decoding of the first image frame, and the first image frame is an I frame; The second sub-level of the first image frame indicates that all the second image frames after the first image frame depend on decoding of the first image frame, and the first image frame is a non-I frame. The method according to claim 20 or 21, characterized in that The second dependency level includes any one of a third sub-level or a fourth sub-level; A third sub-level of the first image frame indicates that all the second image frames after the first image frame depend on the first image frame for decoding, and the number of the second image frames depending on the first image frame for decoding is greater than a number threshold; A fourth sub-level of the first image frame indicates that all the second image frames after the first image frame depend on the first image frame for decoding, and the number of the second image frames depending on the first image frame for decoding is less than or equal to the number threshold. The method according to any one of claims 19 to 22, characterized in that The obtaining of the decoding dependency information of each image frame in the plurality of image frames comprises: parsing the code stream to obtain the reference relationship of each image frame in the plurality of image frames; determining the decoding dependency information of each image frame in the plurality of image frames according to the reference relationship of all the image frames in the plurality of image frames. The method according to any one of claims 19 to 23, characterized in that The generating of the video file based on the decoding dependency information of each image frame in the plurality of image frames and the code stream comprises: writing the decoding dependency information of each image frame in the plurality of image frames into the code stream; packaging the code stream to obtain the video file. The method according to any one of claims 19 to 23, characterized in that The generating of the video file based on the decoding dependency information of each image frame in the plurality of image frames and the code stream comprises: packaging the decoding dependency information of each image frame in the plurality of image frames and the code stream to obtain the video file. A video conversion method, characterized in that, The method comprises: obtaining a first video file, the first video file comprising a code stream obtained from original images of a plurality of image frames in encoded video data; unpackaging the first video file to obtain the code stream; parsing the code stream to obtain the reference relationship of each image frame in the plurality of image frames; determining the decoding dependency information of each image frame in the plurality of image frames according to the reference relationship of all the image frames in the plurality of image frames; generating a second video file based on the decoding dependency information of each image frame in the plurality of image frames and the code stream. The method of claim 26, wherein The generating of the second video file based on the decoding dependency information of each image frame in the plurality of image frames and the code stream comprises: writing the decoding dependency information of each image frame in the plurality of image frames into the code stream; packaging the code stream to obtain the second video file. The method of claim 26, wherein The generating of the second video file based on the decoding dependency information of each image frame in the plurality of image frames and the code stream comprises: packaging the decoding dependency information of each image frame in the plurality of image frames and the code stream to obtain the second video file. A video file characterized in that The video file comprises a code stream, the code stream comprising encoded data obtained from original images of a plurality of image frames in encoded video data and decoding dependency information of each image frame in the plurality of image frames, the decoding dependency information of a first image frame being used to indicate the number of second image frames depending on the first image frame for decoding, the first image being any one of the plurality of image frames, and the second image frames being image frames after the first image frame in a picture group to which the first image frame belongs. A video file characterized in that The video file comprises a code stream and decoding dependency information of each image frame in the plurality of image frames; The code stream is obtained by encoding original images of the plurality of image frames, and the decoding dependency information of a first image frame is used to indicate a quantity of second image frames that depend on decoding of the first image frame, the first image is any one of the plurality of image frames, and the second image frame is an image frame after the first image frame in a group of pictures to which the first image frame belongs. An electronic device, characterized by comprising: Comprise: a memory and a processor, the memory being coupled with the processor; the memory stores program instructions, when the program instructions are executed by the processor, the electronic device executes the method as claimed in any one of claims 1 to 28. A computer-readable storage medium, characterized by, The computer readable storage medium stores a computer program, when the computer program runs on a computer or a processor, the computer or the processor executes the method as claimed in any one of claims 1 to 28. A computer program product, characterized in that The computer program product comprises computer instructions, when the computer instructions are executed by a computer or a processor, the steps of the method as claimed in any one of claims 1 to 28 are executed. A computer-readable storage medium, characterized by The computer readable storage medium stores a video file, the video file comprises a code stream, the code stream comprises encoded data obtained by encoding original images of a plurality of image frames in video data and decoding dependency information of each image frame in the plurality of image frames, the decoding dependency information of a first image frame is used to indicate a quantity of second image frames that depend on decoding of the first image frame, the first image is any one of the plurality of image frames, and the second image frame is an image frame after the first image frame in a group of pictures to which the first image frame belongs. A computer-readable storage medium, characterized by The computer readable storage medium stores a video file, the video file comprises a code stream and decoding dependency information of each image frame in a plurality of image frames; The code stream is obtained by encoding original images of the plurality of image frames, and the decoding dependency information of a first image frame is used to indicate a quantity of second image frames that depend on decoding of the first image frame, the first image is any one of the plurality of image frames, and the second image frame is an image frame after the first image frame in a group of pictures to which the first image frame belongs. A video player characterized by comprising: The video player is used to: obtain a video file, the video file comprises a code stream, the code stream comprises encoded data obtained by encoding original images of a plurality of image frames in video data; determine a target display frame from the plurality of image frames; determine a plurality of candidate paths for obtaining a reconstructed image of the target display frame; determine evaluation information corresponding to each candidate path in the plurality of candidate paths, the evaluation information is used to describe a video playback experience; select a candidate path from the plurality of candidate paths as a target path based on the evaluation information corresponding to each candidate path; obtain a reconstructed image of the target display frame based on the target path and the code stream; play the reconstructed image of the target display frame. A video recorder characterized by comprising: The video recorder is used to: obtain video data, the video data comprises original images of a plurality of image frames; encode the original images of the plurality of image frames to obtain a code stream; obtain decoding dependency information corresponding to each of the plurality of image frames, the decoding dependency information of a first image frame being used to indicate a number of second image frames that depend on decoding of the first image frame, the first image being any one of the plurality of image frames, and the second image frames being image frames after the first image frame in a group of pictures to which the first image frame belongs; generate a video file based on the decoding dependency information of each of the plurality of image frames and the code stream.

Citation Information

Patent Citations

  • B frame position decision method and device

    CN105898307A

  • Video frame extraction processing method and device, equipment and medium

    CN112866799A

  • Video coding method and device, storage medium and electronic equipment

    CN113038124A

  • Sek processing method of streaming media video, electronic equipment and storage medium

    CN116684515A

  • Video coding method and device, electronic equipment, storage medium and program product

    CN118354069A