Video encoding method and device, encoding information scheduling method and device

By determining the discarding priority coefficient of video frames in SVC encoding and adding them to the encoding information, the video quality problem caused by direct discarding of the enhancement layer is solved, and video quality is guaranteed and user experience is improved.

CN116074528BActive Publication Date: 2025-08-26BEIJING YUANLI WEILAI SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111277292.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-29
Publication Date
2025-08-26
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

In the existing SVC encoding method, if frame data in the enhancement layer needs to be discarded, the entire enhancement layer is usually discarded directly, resulting in poor video quality and affecting the user experience.

Method used

By acquiring the video to be encoded and encoding it into the basic layer and at least one enhancement layer, the discard priority coefficient of each video frame is determined and added to the initial encoding information, the video frames in the enhancement layer are selectively discarded according to the video data.

Benefits of technology

When discarding frame data of the enhancement layer, frame data with little impact on video quality is preferred to ensure video quality and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116074528B_ABST
    Figure CN116074528B_ABST
Patent Text Reader

Abstract

This specification provides a video encoding method and apparatus, and an encoding information scheduling method and apparatus. The video encoding method includes: obtaining a video to be encoded, encoding the video to be encoded into a base layer and at least one enhancement layer, and obtaining initial encoding information for the video to be encoded; determining, based on the video data of the video to be encoded, a priority coefficient for discarding each video frame of the video to be encoded; and adding the priority coefficient of each video frame to the initial encoding information of the video to be encoded to obtain target encoding information for the video to be encoded. In this way, the priority coefficient for discarding each video frame can also be added to the encoding information. When frame data in the enhancement layer needs to be discarded during the subsequent encoding information scheduling process, the video frames in the enhancement layer can be selectively discarded based on the priority coefficient for discarding each video frame determined during the encoding process, thereby ensuring video quality and improving user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of video processing technology, and more particularly to a video encoding method, a video encoding device, a coding information scheduling method, a coding information scheduling device, a computing device, and a computer-readable storage medium. Background Art

[0002] With the rapid development of computer and network technologies, a wide variety of videos have emerged, and video coding technology has also developed rapidly. As video coding technology has evolved, user needs have become increasingly diverse, demanding not only efficient video encoding but also encoding results that can meet a variety of application scenarios. This has led to the development of the Scalable Video Coding (SVC) encoding method. This encoding method can encode video from three perspectives: temporal, spatial, and quality. It provides scalable encoding information in terms of time, space, and quality, meeting network transmission rates and end-user requirements for video in terms of time, space, and signal-to-noise ratio.

[0003] In the SVC encoding method, the lowest-quality layer is called the base layer, and layers that enhance spatial resolution, temporal resolution, or signal-to-noise ratio are called enhancement layers. Currently, during transmission or decoding, if frame data in an enhancement layer needs to be discarded, the entire enhancement layer is often discarded. This can result in poor video quality and affect the user experience. Summary of the Invention

[0004] In view of this, embodiments of this specification provide a video encoding method, a video encoding device, a coding information scheduling method, a coding information scheduling device, a computing device, and a computer-readable storage medium to address technical deficiencies in the prior art.

[0005] According to a first aspect of the embodiments of this specification, a video encoding method is provided, including:

[0006] Obtaining a video to be encoded, and encoding the video to be encoded into a base layer and at least one enhancement layer to obtain initial encoding information of the video to be encoded;

[0007] Determining, according to video data of the video to be encoded, a priority coefficient for discarding each video frame of the video to be encoded, the priority coefficient being used to represent a probability of discarding the corresponding video frame;

[0008] The priority coefficient of each video frame is added to the initial coding information of the video to be coded to obtain the target coding information of the video to be coded.

[0009] According to a second aspect of the embodiments of this specification, a coding information scheduling method is provided, including:

[0010] Obtaining and parsing target coding information of the video to be decoded, obtaining a base layer, at least one enhancement layer, and a discard priority coefficient of each video frame of the video to be decoded, the priority coefficient being used to indicate a probability of discarding the corresponding video frame;

[0011] Determining, based on a current decoding condition and a priority coefficient fed back by a decoding end, a video frame to be discarded in at least one enhancement layer;

[0012] The remaining video frames in the base layer and at least one enhancement layer except the video frames to be discarded are used as decoding information of the video to be decoded, and the decoding information is scheduled to the decoding end.

[0013] According to a third aspect of the embodiments of this specification, a video encoding apparatus is provided, including:

[0014] The encoding module is configured to obtain a video to be encoded, and encode the video to be encoded into a base layer and at least one enhancement layer to obtain initial encoding information of the video to be encoded;

[0015] A first determining module is configured to determine, based on video data of the video to be encoded, a priority coefficient for discarding each video frame of the video to be encoded, where the priority coefficient is used to represent a probability of discarding the corresponding video frame;

[0016] The adding module is configured to add the priority coefficient of each video frame to the initial coding information of the video to be coded to obtain the target coding information of the video to be coded.

[0017] According to a fourth aspect of the embodiments of this specification, a video encoding apparatus is provided, including:

[0018] an acquisition module configured to acquire and parse target coding information of a video to be decoded, obtain a base layer, at least one enhancement layer, and a discard priority coefficient of each video frame of the video to be decoded, the priority coefficient being used to indicate a probability of discarding the corresponding video frame;

[0019] A second determining module is configured to determine a video frame to be discarded in at least one enhancement layer according to a current decoding condition and a priority coefficient fed back by a decoding end;

[0020] The scheduling module is configured to use the remaining video frames in the base layer and at least one enhancement layer except the video frames to be discarded as decoding information of the video to be decoded, and schedule the decoding information to the decoding end.

[0021] According to a fifth aspect of the embodiments of this specification, there is provided a computing device, including:

[0022] memory and processor;

[0023] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of any video encoding method or encoding information scheduling method.

[0024] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, which, when executed by a processor, implement the steps of any video encoding method or encoding information scheduling method.

[0025] The video encoding method provided in this specification can first obtain a video to be encoded, and encode the video to be encoded into a base layer and at least one enhancement layer to obtain initial encoding information of the video to be encoded; then, based on the video data of the video to be encoded, determine a priority coefficient for each video frame of the video to be encoded to be discarded, and the priority coefficient is used to represent the probability of the corresponding video frame being discarded; thereafter, the priority coefficient of each video frame can be added to the initial encoding information of the video to be encoded to obtain target encoding information of the video to be encoded.

[0026] In this case, after encoding the video to be encoded into a basic layer and at least one enhancement layer, a priority coefficient for discarding each video frame can be determined based on the specific video data. The priority coefficient can be determined based on the impact of the video frame on the video quality, and the priority coefficient for discarding each video frame can also be added to the encoding information. In the subsequent scheduling process of the encoding information, if it is necessary to discard the frame data in the enhancement layer, the priority coefficient for discarding each video frame determined in the encoding process can be used to selectively discard the video frames in the enhancement layer according to the priority coefficient. In this way, frame data with little impact on the video quality can be discarded first, thereby ensuring the video quality and improving the user experience.

[0027] The coding information scheduling method provided in this specification can first obtain and parse the target coding information of the video to be decoded, and obtain a priority coefficient for discarding a base layer, at least one enhancement layer, and each video frame of the video to be decoded, and the priority coefficient is used to represent the probability of the corresponding video frame being discarded; then, based on the current decoding conditions and priority coefficient fed back by the decoding end, the video frames to be discarded in at least one enhancement layer can be determined; thereafter, the remaining video frames in the base layer and at least one enhancement layer except the video frames to be discarded are used as decoding information of the video to be decoded, and the decoding information is scheduled to the decoding end.

[0028] In this case, the target coding information of the video to be decoded may include, in addition to a basic layer and at least one enhancement layer obtained by encoding the video, a priority coefficient for discarding each video frame determined according to specific video data during the encoding process. The priority coefficient can represent the impact of the video frame on the video quality. Therefore, in the scheduling process of the coding information, if it is necessary to discard the frame data in the enhancement layer, the priority coefficient for discarding each video frame determined in the encoding process can be used to selectively discard the video frames in the enhancement layer according to the priority coefficient. In this way, frame data with little impact on the video quality can be discarded first, thereby ensuring the video quality and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 This is a flowchart of a video encoding method provided by an embodiment of this specification;

[0030] Figure 2 This is a flow chart of a coding information scheduling method provided in one embodiment of this specification;

[0031] Figure 3 This is a schematic diagram of time domain coding provided by an embodiment of this specification;

[0032] Figure 4 This is a schematic diagram of a spatial domain / quality domain coding provided in an embodiment of this specification;

[0033] Figure 5 This is a schematic structural diagram of a video encoding device provided in one embodiment of this specification;

[0034] Figure 6 This is a schematic diagram of the structure of a coding information scheduling device provided in one embodiment of this specification;

[0035] Figure 7 This is a structural block diagram of a computing device provided in one embodiment of this specification. DETAILED DESCRIPTION

[0036] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0037] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "an," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0038] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0039] First, the terms involved in one or more embodiments of this specification are explained.

[0040] Scalable Video Coding (SVC), or scalable video coding, is a video coding technology and standard that is generally categorized into temporal scalability, spatial scalability, and quality scalability. During encoding, video is output as a base layer and several enhancement layers. During decoding, the base layer can deliver basic video content at a lower bitrate, while the addition of additional enhancement layers allows for higher video quality. This technology enables flexible and configurable video streams to be created with a single encoding process, adapting to varying network bandwidths and playback scenarios.

[0041] Temporal domain: Also called the time domain, the independent variable is time, meaning the horizontal axis is time and the vertical axis is signal change. In video sequences, this refers to the sequential relationship between multiple images. Spatial domain: Also called the spatial domain, also known as the pixel domain, processing in the spatial domain is pixel-level processing. In video sequences, this refers to the information of a single frame.

[0042] In this specification, a video encoding method is provided. This specification also relates to a video encoding device, a coding information scheduling method, a coding information scheduling device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0043] Figure 1 A flowchart of a video encoding method according to an embodiment of the present disclosure is shown, which specifically includes the following steps:

[0044] Step 102: Obtain a video to be encoded, and encode the video to be encoded into a base layer and at least one enhancement layer to obtain initial encoding information of the video to be encoded.

[0045] It should be noted that the video to be encoded may refer to a video waiting for scalable video encoding. The video to be encoded may be a video in any field and any scenario. The video to be encoded may refer to a video stored locally on the encoding device, or it may be a video obtained from other devices. For example, the video to be encoded may be a chat video obtained in real time during a chat, or a live video obtained in real time during a live broadcast, or it may be a recorded video pre-obtained and stored locally by the encoding device, etc.

[0046] In this specification, after the encoding device obtains the video to be encoded, a scalable video encoding method can be used to encode the obtained video to be encoded into a basic layer and at least one enhancement layer. The layered information obtained by the scalable video encoding method is used as the initial encoding information of the video to be encoded, and then the priority information of the discarded video frames is added to serve as the final encoding information of the video to be encoded.

[0047] In an optional implementation of this embodiment, the video to be encoded may be encoded according to the time domain, spatial domain, and / or quality domain to obtain the base layer and at least one enhancement layer. That is, the video to be encoded is encoded into a base layer and at least one enhancement layer to obtain initial encoding information of the video to be encoded. The specific implementation process may be as follows:

[0048] The video to be encoded is encoded into a base layer and at least one enhancement layer according to the time domain information, spatial domain information and / or quality domain information of the video to be encoded, so as to obtain initial encoding information of the video to be encoded.

[0049] It should be noted that scalable video coding can encode the video in three dimensions: time, space, and quality. This provides scalable coding information, which can be output as a base layer and several enhancement layers during encoding. In other words, during encoding, the video can be encoded in one dimension or all three dimensions, selecting at least one of these three dimensions: time, space, and quality. During decoding, the base layer can deliver basic video content at a lower bitrate, while the addition of additional enhancement layers can achieve higher video quality. This technology enables flexible and configurable video streams to be achieved with a single encoding process, suitable for different network bandwidths and playback scenarios.

[0050] In actual implementation, in order to achieve temporal scalability, a hierarchical bidirectional prediction frame coding method can be used; to achieve spatial scalability, the motion, texture and residual information between layers can be utilized and a layered coding method can be used; to achieve signal-to-noise ratio scalability, coarse-grained scalability and medium-grained scalability methods can be used. These two methods use an inter-layer prediction method similar to spatial scalability.

[0051] In this specification, a video to be encoded can be encoded into a base layer and at least one enhancement layer based on the temporal information, spatial information, and / or quality domain information of the video to be encoded, thereby obtaining initial encoding information of the video to be encoded. This allows the video to be encoded from the three aspects of time, space, and quality, respectively, providing temporally, spatially, and quality scalable encoding information to meet the network transmission rate and the end user's requirements for the video in terms of time, space, and signal-to-noise ratio.

[0052] Step 104: Determine a priority coefficient for discarding each video frame of the video to be encoded based on the video data of the video to be encoded, where the priority coefficient is used to represent a probability of discarding the corresponding video frame.

[0053] Specifically, after obtaining a video to be encoded and encoding it into a base layer and at least one enhancement layer to obtain initial encoding information for the video to be encoded, a priority coefficient for discarding each video frame of the video to be encoded is further determined based on the video data of the video to be encoded. The priority coefficient is used to represent the probability of the corresponding video frame being discarded. The priority coefficient and the probability of discarding can be directly proportional or inversely proportional, that is, the higher the priority coefficient, the greater the probability of the corresponding video frame being discarded, or the higher the priority coefficient, the lower the probability of the corresponding video frame being discarded.

[0054] It should be noted that the smaller the impact of a video frame on the video quality of the complete video, that is, the smaller the impact on the user's viewing experience of the video, the greater the probability of the video frame being discarded; the greater the impact of a video frame on the video quality of the complete video, that is, the greater the impact on the user's viewing experience of the video, the smaller the probability of the video frame being discarded. Therefore, based on the video data of the video to be encoded, the specific video content included in the video frame can be analyzed to determine the impact of the video frame on the video quality of the complete video, thereby determining the priority coefficient of the video frame being discarded.

[0055] In practical applications, it is assumed that the priority coefficient is proportional to the probability of being discarded. That is, the higher the priority coefficient, the greater the probability of the corresponding video frame being discarded. In other words, the greater the impact of a video frame on the video quality of the complete video, the higher the priority coefficient for discarding the video frame can be set, and the smaller the impact of a video frame on the video quality of the complete video, the lower the priority coefficient for discarding the video frame can be set. This allows the priority coefficient of a video frame to represent its impact on the video quality of the complete video, that is, its impact on the user's viewing experience, and thus represents the probability of subsequent discarding.

[0056] In this specification, after encoding the video to be encoded into a base layer and at least one enhancement layer, a priority coefficient for discarding each video frame can be additionally determined based on the specific video data. The priority coefficient can represent the impact of the video frame on the video quality, that is, the probability that it may be discarded later. When the frame data in the enhancement layer needs to be discarded later, the video frames in the enhancement layer can be selectively discarded based on the priority coefficient for discarding each video frame. In this way, frame data with little impact on video quality can be discarded first, thereby ensuring video quality and improving user experience.

[0057] In an optional implementation of this embodiment, for a case where the initial coding information is obtained by coding based on the time domain information of the video to be coded, that is, a base layer and at least one enhancement layer are obtained by coding based on the time domain information of the video to be coded, a priority coefficient of a video frame may be determined based on a change between two preceding and following video frames. That is, a priority coefficient of each video frame included in the video to be coded may be determined based on the video data of the video to be coded. The specific implementation process may be as follows:

[0058] Determining an inter-frame change magnitude between a current video frame and a previous video frame of the video to be encoded;

[0059] According to the magnitude of the inter-frame change, the priority coefficient of the current video frame is determined.

[0060] It should be noted that the greater the inter-frame change between the current video frame and the previous video frame, the more drastic the change in the current video frame compared to the previous video frame, that is, the less similar the current video frame is to the previous video frame. If the current video frame is discarded, it may cause the final decoded and played video to be too abrupt, affecting the user's viewing experience. In other words, the more drastic the change in the video frame, the greater the impact on the video quality of the complete video, and the corresponding impact on the user's viewing experience. Therefore, the more drastic the change in the video frame, the smaller the probability of being discarded.

[0061] In addition, the smaller the inter-frame change amplitude between the current video frame and the previous video frame, the less obvious the change of the current video frame compared with the previous video frame, that is, the more similar the current video frame and the previous video frame are. If the current video frame is discarded, only the latter video frame of the two more similar video frames in the video is discarded, which will not cause obvious changes in the final decoded and played video, nor will it affect the user's viewing experience. In other words, the less drastic the change of the video frame, the smaller the impact on the video quality of the complete video, and the corresponding smaller the impact on the user's viewing experience. Therefore, the less drastic the change of the video frame, the greater the probability of being discarded.

[0062] In this specification, for the case where the initial coding information is obtained by encoding based on the time domain information of the video to be encoded, the inter-frame change amplitude between the current video frame and the previous video frame of the video to be encoded can be determined, and then the priority coefficient of the current video frame can be determined based on the inter-frame change amplitude, so that the determined priority coefficient can represent the degree of change between the current video frame and the previous video frame, and thus the impact of the current video frame on the video quality of the complete video after being discarded can be determined based on the priority coefficient, so that frames with small changes in the time domain can be discarded first, and frames with drastic changes can be retained, so that the video appears to be smoother subjectively.

[0063] In an optional implementation of this embodiment, the inter-frame change amplitude between the current video frame and the previous video frame may be determined based on the pixel difference between the current video frame and the previous video frame, that is, the inter-frame change amplitude between the current video frame and the previous video frame of the video to be encoded may be determined. The specific implementation process may be as follows:

[0064] For each pixel in the current video frame, determine the pixel difference between the pixel in the current video frame and the pixel in the previous video frame;

[0065] Determining a first pixel average and a first pixel standard deviation based on pixel differences between each pixel in a current video frame and a previous video frame;

[0066] The first pixel average value and the first pixel standard deviation are used as the inter-frame variation amplitude.

[0067] It should be noted that a single-frame video is actually an image, which is composed of multiple pixels. For each pixel in the current video frame, if the pixel difference between the current video frame and the previous video frame is large, it means that the difference between the pixel in the current video frame and the previous video frame is greater, that is, the greater the change. Therefore, based on the pixel difference between each pixel in the current video frame and the previous video frame, the difference between each pixel in the current video frame and the previous video frame can be determined, thereby determining the degree of change in the current video frame compared to the previous video frame, that is, the inter-frame change amplitude.

[0068] In practical applications, the difference between each pixel point in the current video frame and the previous video frame can be represented by the first pixel average value and the first pixel standard deviation between each pixel point in the current video frame and the previous video frame. That is, the first pixel average value and the first pixel standard deviation can be used as the inter-frame change amplitude to reflect the degree of change of the current video frame compared with the previous video frame.

[0069] In specific implementation, the inter-frame change amplitude between the current video frame and the previous video frame can be determined by the following formulas (1) and (2):

[0070] TI std =std space [M n (i,j)] (1)

[0071] TI avg =avg space [M n (i,j)] (2)

[0072] Among them, M n (i,j)=F n (i,j)-F n-1 (i, j), represents the pixel difference between the current video frame and the previous video frame at the pixel point (i, j), std space Represents M for all pixels (i, j) of the entire frame image n Find the standard deviation, avg space Represents M for all pixels (i, j) of the entire frame image n Find the mean. std is the first pixel standard deviation, TI avg is the first pixel average value, TI std and TI avg That is, the inter-frame change amplitude between the current video frame and the previous video frame.

[0073] This specification can comprehensively analyze the changes in the pixel value of each pixel point in the current video frame compared with the pixel value of the pixel point in the previous video frame, so as to express the degree of change of the current video frame compared with the previous video frame according to the gap between the pixel points between the current video frame and the previous video frame, and the process of determining the amplitude of inter-frame change is simple and accurate; and, the first pixel average value and the first pixel standard deviation between each pixel point are jointly used as the amplitude of inter-frame change, which improves the accuracy of determining the amplitude of inter-frame change of the current video frame compared with the previous video frame, thereby improving the accuracy of the priority coefficient determined subsequently and ensuring the subsequent video quality.

[0074] In an optional implementation of this embodiment, multiple grading threshold ranges may be pre-set for the first pixel average value and the first pixel standard deviation, so that the priority coefficient of the current video frame is determined based on the corresponding grading threshold ranges. That is, the priority coefficient of the current video frame is determined based on the inter-frame variation amplitude. The specific implementation process may be as follows:

[0075] A priority coefficient of the current video frame is determined according to the first pixel average value and the corresponding multiple grading threshold ranges, and the first pixel standard deviation and the corresponding multiple grading threshold ranges.

[0076] It should be noted that multiple corresponding graded threshold ranges can be pre-set for the first pixel average value, for determining the threshold range within which the first pixel average value lies. For example, the multiple graded threshold ranges corresponding to the first pixel average value can be greater than average threshold 0, less than average threshold 0 and greater than average threshold 1, less than average threshold 1 and greater than average threshold 2, etc. Multiple corresponding graded threshold ranges can also be pre-set for the first pixel standard deviation, for determining the threshold range within which the first pixel standard deviation lies. For example, the multiple graded threshold ranges corresponding to the first pixel standard deviation can be greater than standard deviation threshold 0, less than standard deviation threshold 0 and greater than standard deviation threshold 1, less than standard deviation threshold 1 and greater than standard deviation threshold 2, etc. Among them, average threshold 0, average threshold 1, average threshold 2, and standard deviation threshold 0, standard deviation threshold 1, standard deviation threshold 2 merely represent threshold numbers and do not represent specific numerical values.

[0077] In an optional implementation of this embodiment, the priority coefficient of the current video frame is determined based on the first pixel average value and the corresponding multiple classification threshold ranges, and the first pixel standard deviation and the corresponding multiple classification threshold ranges. The specific implementation process can be as follows:

[0078] If the first pixel average value is within a first threshold range, and / or the first pixel standard deviation is within a second threshold range, the priority coefficient corresponding to the first threshold range and / or the second threshold range is determined as the priority coefficient of the current video frame.

[0079] It should be noted that, among the multiple graded threshold ranges corresponding to the first pixel average value, each graded threshold range is set with a corresponding priority coefficient, and among the multiple graded threshold ranges corresponding to the first pixel standard deviation, each graded threshold range is also set with a corresponding priority coefficient. The first threshold range represents the threshold range within which the first pixel average value lies, and the second threshold range represents the threshold range within which the first pixel standard deviation lies. The first threshold range and the second threshold range are corresponding threshold ranges, that is, the first threshold range and the second threshold range correspond to the same priority coefficient. In other words, the multiple graded threshold ranges corresponding to the first pixel average value correspond to the multiple graded threshold ranges corresponding to the first pixel standard deviation, and a group of threshold ranges corresponds to the same priority coefficient.

[0080] In practical applications, the priority coefficient of the current video frame can be determined based on an AND logic, that is, the first pixel average value is within the first threshold range, and at the same time the first pixel standard deviation is within the second threshold range. At this time, the priority coefficient corresponding to the first threshold range and the second threshold range can be determined as the priority coefficient of the current video frame. Alternatively, the priority coefficient of the current video frame can also be determined based on an OR logic, that is, the first pixel average value is within the first threshold range, or the first pixel standard deviation is within the second threshold range. At this time, the priority coefficient corresponding to the first threshold range or the second threshold range can be determined as the priority coefficient of the current video frame.

[0081] For example, if TI_standard deviation > standard deviation threshold 0, or TI_average value > average value threshold 0, the priority coefficient of the current video frame is the priority coefficient corresponding to standard deviation threshold 0 or average value threshold 0. Among them, the priority coefficients corresponding to standard deviation threshold 0 and average value threshold 0 are the same. If both are 0, the priority coefficient of the current video frame is 0; if standard deviation threshold 1 < TI_standard deviation <= standard deviation threshold 0, or average value threshold 1 < TI_average value <= average value threshold 0, the priority coefficient of the current video frame is the priority coefficient corresponding to standard deviation threshold 1 or average value threshold 1. Among them, the priority coefficients corresponding to standard deviation threshold 1 and average value threshold 1 are the same. If both are 1, the priority coefficient of the current video frame is 1.

[0082] In the embodiments of this specification, the first pixel average value and the first pixel standard deviation between each pixel point of the current video frame and the previous video frame can be jointly used as the inter-frame change amplitude. Therefore, multiple corresponding hierarchical threshold ranges can be preset for the first pixel average value and the first pixel standard deviation, and the hierarchical threshold range where the first pixel average value is located and the hierarchical threshold range where the first pixel standard deviation is located can be determined respectively. Then, according to the combination of the hierarchical threshold ranges where the first pixel average value and the first pixel standard deviation are located, the priority coefficient of the current video frame is determined. By combining the average value and the standard deviation, the priority coefficient of the current video frame is jointly determined, which improves the determination accuracy.

[0083] In an optional implementation manner of this embodiment, for the case where the initial coding information is encoded based on the spatial domain information or the quality domain information of the video to be encoded, that is, the case where one base layer and at least one enhancement layer are encoded based on the spatial domain information or the quality domain information of the video to be encoded, the priority coefficient of each video frame can be determined according to the complexity and upsampling quality of each video frame, that is, according to the video data of the video to be encoded, the priority coefficient of each video frame included in the video to be encoded is determined. The specific implementation process can be as follows:

[0084] For each video frame included in the video to be encoded, determine the complexity of the video frame in the spatial domain or the quality domain;

[0085] A priority coefficient of the video frame is determined according to the complexity of the video frame.

[0086] It should be noted that the principle of SVC spatial scalability is that the base layer transmits low-resolution video, and the enhancement layer can restore high-resolution video frames. If the enhancement layer is discarded, high-resolution video frames can only be restored through upsampling during playback. Compared with high-complexity video frames, when only the base layer is used, low-complexity video frames are more effective in restoring high-resolution through upsampling.

[0087] In practical applications, for video frames with lower complexity, the perceived quality of the video restored based on the base layer and upsampling is better, while for video frames with higher complexity, the perceived quality of the video restored based on the base layer and upsampling is lower. In other words, the higher the complexity of the video frame, the lower the upsampling quality. If this video frame is discarded, it may result in poor clarity in the final decoded and played video, affecting the user's viewing experience. In other words, the more complex the video frame, the greater its impact on the video quality of the complete video, and accordingly, the greater the impact on the user's viewing experience. Therefore, the probability of a video frame being discarded should be lower for more complex video frames.

[0088] In addition, the lower the complexity of the video frame, the better the upsampling quality. If the video frame is discarded, it will have less impact on the clarity of the final decoded and played video, and thus will not overly affect the user's viewing experience. That is, the lower the complexity of the video frame, the smaller the impact on the video quality of the complete video, and the corresponding impact on the user's viewing experience. Therefore, the lower the complexity of the video frame, the greater the probability of being discarded.

[0089] In this specification, for the case where the initial coding information is obtained by encoding based on the spatial domain information or quality domain information of the video to be encoded, the complexity of the video frame can be determined in the spatial domain or quality domain for each video frame included in the video to be encoded, and then the priority coefficient of the video frame can be determined based on the complexity of the video frame, so that the determined priority coefficient can represent the impact of the video frame on the video clarity of the complete video after being discarded, so that frames with low complexity, that is, frames with good upsampling quality, can be discarded first in the subsequent process, thereby ensuring better subjective clarity.

[0090] In an optional implementation of this embodiment, the complexity of each video frame may be determined based on the pixel difference between each pixel point in each video frame, that is, the complexity of the video frame may be determined in the spatial domain or the quality domain. The specific implementation process may be as follows:

[0091] Determine a second pixel average value and a second pixel standard deviation of each pixel point included in the video frame;

[0092] The second pixel average value and the second pixel standard deviation are used as the complexity of the video frame.

[0093] It should be noted that a single frame of video is actually an image, which is composed of multiple pixels. For each video frame, the greater the pixel difference between each pixel in the video frame, the higher the complexity of the video frame, and the smaller the pixel difference between each pixel in the video frame, the lower the complexity of the video frame. Therefore, for each video frame, the second pixel average value and second pixel standard deviation of each pixel in the video frame can be determined, and the second pixel average value and second pixel standard deviation can be used as the complexity of the video frame.

[0094] In specific implementation, for each video frame, the complexity of the video frame can be determined by the following formula (3) and formula (4):

[0095] SI std =std space [Sobel(Fn)] (3)

[0096] SI avg =avg space [Sobel(Fn)] (4)

[0097] Among them, Sobel(Fn) means that the video frame Fn is filtered by Sobel to obtain the pixel-by-pixel result, std space Indicates the standard deviation of the filtering results Fn of all pixels in the entire frame image, avg space Indicates the average value of the filtering results Fn of all pixels in the entire frame image. std is the second pixel standard deviation, SI avg is the second pixel average, SI std and SI avg This is the complexity of the video frame.

[0098] In the embodiments of the present specification, for each video frame, the pixel value differences of each pixel point in the video frame can be comprehensively analyzed, so that the complexity of the video frame can be represented according to the pixel value differences of each pixel point, and the process of determining the complexity of the video frame is simple and accurate; and, the second pixel average value and the second pixel standard deviation between each pixel point are jointly used as the complexity of the video frame, which improves the accuracy of determining the complexity of the video frame, thereby improving the accuracy of the priority coefficient determined subsequently, and ensuring the subsequent video quality.

[0099] In an optional implementation of this embodiment, for each video frame, the second pixel average value and the second pixel standard deviation are used as the complexity of the video frame, and when the priority coefficient of the video frame is determined based on the complexity of the video frame, the specific implementation process is similar to the above-mentioned process of determining the priority coefficient of the current video frame based on the amplitude of the inter-frame change, and will not be repeated in this embodiment of this specification.

[0100] Step 106: Add the priority coefficient of each video frame to the initial coding information of the video to be coded to obtain the target coding information of the video to be coded.

[0101] Specifically, based on determining the priority coefficient of discarding each video frame of the video to be encoded according to the video data of the video to be encoded, the priority coefficient of each video frame can be further added to the initial encoding information of the video to be encoded to obtain the target encoding information of the video to be encoded.

[0102] It should be noted that after encoding the video to be encoded into a basic layer and at least one enhancement layer, the layered information obtained by encoding is not directly used as the final encoding information of the video to be encoded. Instead, the priority coefficient of each video frame in the video to be encoded can be further determined. The priority coefficient represents the probability of the video frame being deleted subsequently, and the priority coefficient of each video frame is added to the initial encoding information of the video to be encoded to obtain the final target encoding information of the video to be encoded. When the target encoding information is subsequently decoded, the video frames in the enhancement layer can be selectively discarded according to the priority coefficient of each video frame included in the target encoding information, so that frame data with little impact on video quality can be discarded first, thereby ensuring video quality and improving user experience.

[0103] In an optional implementation of this embodiment, after adding the priority coefficient of each video frame to the initial coding information of the video to be coded to obtain the target coding information of the video to be coded, the following steps may be further included:

[0104] Parsing the target coding information to obtain a priority coefficient for discarding each video frame of the video to be decoded;

[0105] According to the priority coefficient, decoding information of the video to be decoded is determined and scheduled to the decoding end.

[0106] It should be noted that by adding the priority coefficients of each video frame to the initial encoding information of the video to be encoded, the final target encoding information of the video to be encoded can be obtained. This target encoding information can then be parsed, and based on the priority coefficients of each discarded video frame obtained from the parsing, the final scheduled decoding information is determined and dispatched to the decoding end for decoding and playback. The detailed scheduling process is further described in the encoding information scheduling method below.

[0107] The video encoding method provided in this specification can also determine the inter-frame change amplitude between the current video frame and the previous video frame of the video to be encoded after obtaining a basic layer and at least one enhancement layer through time domain encoding, and then determine the priority coefficient of the current video frame based on the inter-frame change amplitude, so that the determined priority coefficient can represent the degree of change between the current video frame and the previous video frame, and thus, based on the priority coefficient, it is possible to determine the impact of the current video frame on the video quality of the complete video after being discarded, so that frames with small changes in the time domain can be discarded first, and frames with drastic changes can be retained, so that the video appears to be smoother subjectively.

[0108] In addition, after encoding a basic layer and at least one enhancement layer in the spatial domain or quality domain, the complexity of the video frame and the upsampling quality of the video frame can be determined in the spatial domain or quality domain for each video frame included in the video to be encoded, and then the priority coefficient of the video frame can be determined based on the complexity and upsampling quality of the video frame, so that the determined priority coefficient can represent the impact of the video frame on the video clarity of the complete video after being discarded, so that frames with low complexity and good upsampling quality can be discarded first in the subsequent process, thereby ensuring better subjective clarity.

[0109] That is to say, this specification can determine the priority coefficient for discarding each video frame based on the specific video data, and add the priority coefficient for discarding each video frame to the coding information, so that in the subsequent scheduling process of the coding information, if it is necessary to discard the frame data in the enhancement layer, the priority coefficient for discarding each video frame determined in the encoding process can be used to selectively discard the video frames in the enhancement layer according to the level of the priority coefficient. In this way, the frame data with little impact on the video quality can be discarded first, thereby ensuring the video quality and improving the user experience.

[0110] Figure 2 A flowchart of a coding information scheduling method provided according to an embodiment of this specification is shown, which specifically includes the following steps:

[0111] Step 202: Obtain and parse target coding information of the video to be decoded, and obtain a base layer, at least one enhancement layer, and a discard priority coefficient of each video frame of the video to be decoded, wherein the priority coefficient is used to indicate the probability of the corresponding video frame being discarded.

[0112] In practical applications, the target coding information may include not only a base layer and at least one enhancement layer obtained by encoding the video to be decoded, but also a priority coefficient for discarding each video frame, determined based on the specific video data during the encoding process. Therefore, in practical applications, when scheduling the target coding information of the video to be decoded, the scheduling system may first parse the target coding information of the video to be decoded to obtain the priority coefficients for discarding the base layer, at least one enhancement layer, and each video frame of the video to be decoded.

[0113] It should be noted that the priority coefficient can represent the impact of the video frame on the video quality, that is, the probability of the video frame being discarded, so that in the scheduling process of the encoding information, the video frame in the enhancement layer can be selectively discarded based on the priority coefficient of each video frame being discarded, and the decoding information that can be scheduled to the decoding end can be obtained. In this way, the frame data with little impact on the video quality can be discarded first, thereby ensuring the video quality and improving the user experience.

[0114] Step 204: Determine the video frames to be discarded in at least one enhancement layer according to the current decoding condition and priority coefficient fed back by the decoding end.

[0115] Specifically, after acquiring and parsing the target coding information of the video to be decoded, obtaining a basic layer, at least one enhancement layer, and a priority coefficient for discarding each video frame of the video to be decoded, further, based on the current decoding conditions and priority coefficients fed back by the decoding end, the video frames to be discarded in at least one enhancement layer are determined.

[0116] It should be noted that the current decoding condition can refer to the decoding capability of the target encoding information of the video to be decoded, that is, the number of enhancement layers that can be decoded in addition to the base layer. For example, the current decoding condition can be the current network bandwidth, playback scenario, etc. For example, the greater the network bandwidth, the more enhancement layers that can be decoded. The current decoding condition is fed back to the scheduling system by the decoding end, and the scheduling system uses it to make scheduling decisions and determine which video frames to discard.

[0117] In practical applications, the current decoding conditions can be used to determine the decoding capability of the decoder, specifically how many video frames can be decoded. This allows the scheduling system to determine how many frames need to be discarded. The scheduling system can then filter out frames to be discarded from at least one enhancement layer based on the priority coefficients of each video frame included in the target coding information. These frames are then discarded and not transmitted to the decoder for decoding.

[0118] In addition, since the priority coefficient of each video frame to be discarded determined during the encoding process can represent the impact of the video frame on the video quality, that is, the probability of being discarded, when it is necessary to discard frame data in the enhancement layer during the scheduling process, the video frames in the enhancement layer can be selectively discarded based on the priority coefficient of each video frame to be discarded, so that frame data with little impact on video quality can be discarded first, thereby ensuring video quality and improving user experience.

[0119] In an optional implementation of this embodiment, when selecting video frames to be discarded from at least one enhancement layer, it may be first determined which enhancement layers to select from. That is, based on the current decoding conditions and priority coefficients fed back by the decoding end, the video frames to be discarded from the at least one enhancement layer may be determined. The specific implementation process may be as follows:

[0120] determining, according to a current decoding condition, a target enhancement layer to be discarded in at least one enhancement layer;

[0121] According to the priority coefficient, the video frames to be discarded are screened out from the video frames included in the target enhancement layer.

[0122] It should be noted that different decoding conditions result in different numbers of layers that can be decoded. For example, when the network bandwidth is good, in addition to decoding the base layer, multiple enhancement layers can be decoded, or even all enhancement layers can be decoded. The better the network bandwidth, the more enhancement layers that can be decoded. When the network bandwidth is poor, in addition to decoding the base layer, fewer enhancement layers can be decoded, or even only the base layer can be decoded. The worse the network bandwidth, the fewer enhancement layers that can be decoded, or even no enhancement layers can be decoded.

[0123] Therefore, in practical applications, the number of layers to be discarded corresponding to different decoding conditions can be pre-set, so that the subsequent scheduling system can determine the target enhancement layer to be discarded from at least one enhancement layer based on the current decoding conditions. This target enhancement layer may be an enhancement layer that cannot be fully decoded under the current decoding conditions. For example, in the case of time-domain coding, assuming that the target coding information of the video to be decoded includes a base layer T0 and three enhancement layers T1, T2, and T3, then when the network bandwidth is good, the target enhancement layer to be discarded may be T3, and when the network bandwidth is poor, the target enhancement layers to be discarded may be T3 and T2.

[0124] It should be noted that when there are multiple enhancement layers, the target enhancement layer to be discarded in at least one enhancement layer must be determined in the order of the enhancement layers obtained by encoding, starting from the highest enhancement layer and proceeding downwards. That is, during the scheduling process, when an enhancement layer needs to be discarded, it must be discarded in descending order from the highest enhancement layer.

[0125] In addition, after determining the target enhancement layer, it is not necessary to directly discard the entire target enhancement layer. Instead, the target enhancement layer can be further filtered out from the various video frames included in the target enhancement layer based on the priority coefficient. In other words, the priority coefficient of the video frame can be further used to filter out the video frames that have little impact on the video quality and the user viewing experience from the various video frames included in the target enhancement layer and discard them, while retaining the video frames in the target enhancement layer that have a greater impact on the video quality and the user viewing experience. In other words, the video frames that have a smaller impact on the video quality and the user viewing experience can be discarded first, thereby ensuring the video quality of the decoded video and the user experience.

[0126] In an optional implementation of this embodiment, the minimum number of video frames that need to be discarded can be determined based on the current decoding conditions. Then, based on the priority coefficients of the respective video frames, a corresponding number of video frames can be screened out from the determined target enhancement layer as the video frames to be discarded. The corresponding content is subsequently discarded without decoding. That is, based on the priority coefficients, the video frames to be discarded are screened out from the video frames included in the target enhancement layer. The specific implementation process can be as follows:

[0127] Determine the current number of frames to be discarded based on the current decoding conditions;

[0128] Based on the priority coefficients of the discarded video frames included in the target enhancement layer, the target video frames of the current number of frames to be discarded are sequentially screened out in a preset order;

[0129] The filtered target video frames are used as video frames to be discarded.

[0130] It should be noted that different decoding conditions may also pre-set different numbers of frames to be discarded, so that the decoding device can obtain the corresponding preset number of current frames to be discarded according to the current decoding conditions.

[0131] In practical applications, each video frame included in the target enhancement layer has a corresponding discard priority coefficient, which can indicate the order in which the corresponding video frames are discarded. In one possible implementation, the priority coefficient and the probability of being discarded can be directly proportional, that is, the higher the priority coefficient, the greater the probability of the corresponding video frame being discarded. In this case, it is necessary to filter out the target video frames with the current number of frames to be discarded in order from the priority coefficients of the video frames included in the target enhancement layer (i.e., the preset order is from high to low priority coefficients). The filtered target video frames are the video frames to be discarded and are not subsequently decoded.

[0132] In another possible implementation, the priority coefficient and the probability of being discarded can be inversely proportional, that is, the higher the priority coefficient, the smaller the probability of the corresponding video frame being discarded. At this time, it is necessary to filter out the target video frames with the current number of frames to be discarded in order from low to high according to the priority coefficients of each video frame included in the target enhancement layer (that is, the preset order at this time is the priority coefficient from low to high). The filtered target video frames are the video frames to be discarded and no subsequent decoding is performed.

[0133] In an optional implementation of this embodiment, when the preset order is from high to low priority coefficients, based on the priority coefficients of the respective video frames included in the target enhancement layer, the target video frames of the current number of frames to be discarded are sequentially screened in the preset order. The specific implementation process may be as follows:

[0134] Determine a first video frame having the largest priority coefficient among reference video frames, where the reference video frames are the video frames included in the target enhancement layer;

[0135] If the number of the first video frames exceeds the current number of frames to be discarded, the first video frame of the current number of frames to be discarded is selected as the video frame to be discarded, and the target video frame of the current number of frames to be discarded is obtained;

[0136] If the number of first video frames does not exceed the number of video frames to be discarded, each first video frame is used as a video frame to be discarded, and the video frames to be discarded are continuously filtered from the remaining video frames included in the target enhancement layer until the target video frames with the current number of frames to be discarded are obtained.

[0137] It should be noted that the preset order is priority coefficient from high to low, which means that the higher the priority coefficient, the greater the probability that the corresponding video frame will be discarded, that is, the higher the priority coefficient, the smaller the impact on video quality and user experience, and video frames with high priority coefficients can be discarded first.

[0138] In practical applications, we can first determine the first video frame with the largest priority coefficient among the video frames included in the target enhancement layer, and then determine whether the number of first video frames exceeds the current number of frames to be discarded. If it exceeds, it means that the first video frame with the largest priority coefficient can meet the number requirement of video frames to be discarded. At this time, we can select the first video frame of the current number of frames to be discarded as the video frame to be discarded, and we can get the required number (i.e., the current number of frames to be discarded) of target video frames.

[0139] In addition, if the number of first video frames does not exceed the number of video frames to be discarded, it means that the first video frame with the largest priority coefficient cannot meet the requirement of the number of video frames to be discarded, and therefore it is necessary to continue to discard the video frames with the next priority coefficient. At this time, each first video frame can be used as a video frame to be discarded, and then the remaining video frames to be discarded can be screened from the remaining video frames included in the target enhancement layer until the required number of target video frames (i.e., the current number of frames to be discarded) is obtained.

[0140] In an optional implementation of this embodiment, the video frames to be discarded are continuously filtered from the remaining video frames included in the target enhancement layer until the target video frames with the current number of frames to be discarded are obtained. The specific implementation process may be as follows:

[0141] The remaining number of the current number of frames to be discarded, excluding the number of the first video frame, is used as the updated current number of frames to be discarded;

[0142] Using video frames other than the first video frame in each video frame included in the target enhancement layer as reference video frames;

[0143] Return to the step of determining the first video frame with the largest priority coefficient in the reference video frames until the number of determined first video frames exceeds the current number of frames to be discarded, thereby obtaining the target video frames with the current number of frames to be discarded.

[0144] It should be noted that, since the first video frame with the largest current priority coefficient has been determined as the video frame to be discarded, the remaining number of frames to be discarded other than the first video frame is the number of video frames to be discarded that needs to be further screened out. Therefore, the remaining number of frames to be discarded other than the first video frame can be used as the updated current number of frames to be discarded for subsequent screening.

[0145] In addition, the first video frame included in the target enhancement layer has been determined as the video frame to be discarded, and it is necessary to continue screening from the video frames other than the first video frame in the various video frames included in the target enhancement layer. Therefore, the video frames other than the first video frame in the various video frames included in the target enhancement layer can be used as reference video frames, and the determination of the first video frame with the largest priority coefficient in the reference video frames and subsequent operation steps are continued to continue screening the video frames to be discarded until the video frames to be discarded that meet the number requirements are screened out.

[0146] It should be noted that, since when SVC encoding is performed on a video to obtain a base layer and at least one enhancement layer, the higher layers are dependent on the lower layers. The lower layers in the time domain are the layers that are referenced more often, and the video frames in this layer cannot be easily discarded. The upper layers are the layers that are referenced less or not referenced at all, and the video frames in this layer can be discarded at will.

[0147] In actual applications, after determining the target enhancement layer, when filtering out the video frames to be discarded from the various video frames included in the target enhancement layer based on the priority coefficient, if there are video frames with the same priority coefficient, the video frames in the higher enhancement layer can be discarded first, that is, the video frames that are rarely referenced or not referenced are discarded first. In this way, when determining the video frames to be discarded, the hierarchical order of the SVC encoding process is followed, and the priority coefficients of each video frame are incorporated, so that frame data with the least impact on video quality is discarded first, thus ensuring video quality and improving user experience.

[0148] For example, Figure 3 This is a schematic diagram of a time domain coding provided by an embodiment of this specification, such as Figure 3 As shown in the figure, the time domain layer obtained by time domain coding is a base layer T0, three enhancement layers T1, T2, and T3. Figure 3 The reference relationship between T0, T1, T2, and T3 can represent the dependency relationship between the video frames of each layer. The reference relationship can refer to the coding reference relationship between video frames in video coding, that is, when decoding a certain video frame, a reference video frame is required. For example, if you need to decode video frame 4, you need to decode video frame 0 and video frame 8 first. If you need to decode video frame 2, you need to decode video frame 0 and video frame 4 first, and so on.

[0149] Assuming that the number of layers to be discarded corresponding to the current decoding condition is 2, then the target enhancement layers to be discarded can be determined from high to low to be T3 and T2, and assuming that the number of frames to be discarded corresponding to the current decoding condition is 5. Figure 3 As shown, the T2 and T3 layers include video frame 1, video frame 2, video frame 3, video frame 5, video frame 6, and video frame 7, and the priority coefficient for discarding video frame 1 is P2, the priority coefficient for discarding video frame 2 is also P2, the priority coefficient for discarding video frame 3 is P1, the priority coefficient for discarding video frame 5 is P0, the priority coefficient for discarding video frame 6 is P0, and the priority coefficient for discarding video frame 7 is P1.

[0150] First, it is determined that video frames 1 and 2 have the largest priority coefficients in the T2 and T3 layers. The number of these frames is 2, which does not exceed the current number of frames to be discarded, 5. Therefore, video frames 1 and 2 are set as the video frames to be discarded, and the current number of frames to be discarded is updated to 3. Then, it is determined that video frames 3 and 7 have the largest priority coefficients in the T2 and T3 layers, except for video frames 1 and 2. The number of these frames is 2, which does not exceed the current number of frames to be discarded, 3. Therefore, video frames 3 and 7 are set as the video frames to be discarded, and the current number of frames to be discarded is updated to 1.

[0151] Then, it is determined that in the T2 and T3 layers, except for video frame 1, video frame 2, video frame 3 and video frame 7, video frame 5 and video frame 6 have the largest priority coefficients, and the number is 2, which exceeds the current number of frames to be discarded by 1 frame. Therefore, it is necessary to select one from video frame 5 and video frame 6 as the video frame to be discarded. Since video frame 5 is located in the T3 layer and video frame 6 is located in the T2 layer, video frame 6 is located in a lower layer and is more dependent on other video frames. Therefore, video frame 5 in the T3 layer is discarded and video frame 6 in the T2 layer is retained (that is, when the priority coefficients are the same, the high layer is discarded first and the low layer is retained). That is, video frame 5 is determined as the video frame to be discarded at this time. At this time, the final screened video frames to be discarded are video frame 1, video frame 2, video frame 3, video frame 5 and video frame 7.

[0152] As shown above, if you want to drop the T3 layer, you can drop video frame 1 (P2) first, followed by video frame 3 and video frame 7 (P1), and finally video frame 5 (P0). If you want to drop both T3 and T2 layers, you should drop video frame 1 and video frame 2 first, followed by video frame 3 and video frame 7, and finally video frame 5 and video frame 6.

[0153] Another example, Figure 4 This is a schematic diagram of a spatial domain / quality domain coding provided by an embodiment of this specification, such as Figure 4 As shown, the spatial domain / quality domain coding is layered into a basic layer D0 and an enhancement layer D1. The D1 layer includes video frame 1, video frame 2, video frame 3, and video frame 5. The priority coefficient for discarding video frame 1 is P0, the priority coefficient for discarding video frame 2 is also P0, the priority coefficient for discarding video frame 3 is P1, the priority coefficient for discarding video frame 4 is P1, and the priority coefficient for discarding video frame 5 is P0.

[0154] According to the layering results, the layers with higher levels can be discarded first. Assuming that the number of layers to be discarded corresponding to the current decoding conditions is 1 layer, then the target enhancement layer to be discarded can be determined to be D1 from high to low, and assuming that the current number of frames to be discarded corresponding to the current decoding conditions is 3 frames.

[0155] First, it is determined that video frames 3 and 4 have the largest priority coefficients in the D1 layer, and the number is 2, which does not exceed the current number of frames to be discarded, 3. Therefore, video frames 3 and 4 are selected as the video frames to be discarded, and the current number of frames to be discarded is updated to 1. Then, it is determined that video frames 1, 2, and 5 have the largest priority coefficients in the D1 layer, except for video frames 3 and 4. The number is 3, which exceeds the current number of frames to be discarded, 1. Therefore, any one of video frames 1, 2, and 5 is determined as the video frame to be discarded, assuming it is video frame 1. At this time, the final screened video frames to be discarded are video frames 1, 3, and 4.

[0156] It should be noted that, in addition to including a basic layer and at least one enhancement layer obtained by encoding the video, the target coding information of the video to be decoded may also include a priority coefficient for discarding each video frame determined according to specific video data during the encoding process. The priority coefficient can represent the impact of the video frame on the video quality. Therefore, in the scheduling process of the coding information, if it is necessary to discard the frame data in the enhancement layer, the priority coefficient for discarding each video frame determined in the encoding process can be used to selectively discard the video frames in the enhancement layer according to the priority coefficient. In this way, frame data with little impact on the video quality can be discarded first, thereby ensuring video quality and improving user experience.

[0157] Step 206: Use the remaining video frames in the base layer and at least one enhancement layer except the video frames to be discarded as decoding information of the video to be decoded, and schedule the decoding information to the decoding end.

[0158] It should be noted that, based on the current decoding conditions and priority coefficients fed back by the decoding end, on the basis of determining the video frames to be discarded in at least one enhancement layer, the remaining video frames in the basic layer and at least one enhancement layer except the video frames to be discarded can be used as decoding information of the video to be decoded, and the decoding information can be scheduled to the decoding end.

[0159] In actual applications, after filtering out the video frames to be discarded from at least one enhancement layer based on the priority coefficient of the discarded video frames, the video frames to be discarded filtered out from the enhancement layer of the video to be decoded can be abandoned and not scheduled for transmission. Only the remaining video frames are scheduled for transmission to the decoding end for decoding by the decoding end to obtain the corresponding decoded video for playback.

[0160] The coding information scheduling method provided in this specification may include, in addition to a basic layer and at least one enhancement layer obtained by encoding the video, the target coding information of the video to be decoded may also include a priority coefficient for each video frame to be discarded determined according to specific video data during the encoding process. The priority coefficient may represent the impact of the video frame on the video quality. Therefore, during the scheduling process, if it is necessary to discard frame data in the enhancement layer, the priority coefficient for each video frame to be discarded determined during the encoding process may be used to selectively discard video frames from the enhancement layer according to the priority coefficient. In this way, frame data with little impact on the video quality may be discarded first, thereby ensuring video quality and improving user experience.

[0161] Corresponding to the above method embodiment, this specification also provides a video encoding device embodiment, Figure 5 FIG. 1 shows a schematic diagram of the structure of a video encoding device provided by an embodiment of this specification. Figure 5 As shown, the device includes:

[0162] The encoding module 502 is configured to obtain a video to be encoded, and encode the video to be encoded into a base layer and at least one enhancement layer to obtain initial encoding information of the video to be encoded;

[0163] A first determining module 504 is configured to determine, based on the video data of the video to be encoded, a priority coefficient for discarding each video frame of the video to be encoded, where the priority coefficient is used to indicate a probability of discarding the corresponding video frame;

[0164] The adding module 506 is configured to add the priority coefficient of each video frame to the initial coding information of the video to be encoded to obtain the target coding information of the video to be encoded.

[0165] Optionally, the initial coding information is obtained by coding based on time domain information of the video to be coded; the first determining module 504 is further configured to:

[0166] Determining an inter-frame change magnitude between a current video frame and a previous video frame of the video to be encoded;

[0167] According to the magnitude of the inter-frame change, the priority coefficient of the current video frame is determined.

[0168] Optionally, the first determining module 504 is further configured to:

[0169] For each pixel in the current video frame, determining a pixel difference between the pixel in the current video frame and the pixel in the previous video frame;

[0170] Determining a first pixel average and a first pixel standard deviation based on pixel differences between each pixel in the current video frame and the previous video frame;

[0171] The first pixel average value and the first pixel standard deviation are used as the inter-frame variation amplitude.

[0172] Optionally, the first determining module 504 is further configured to:

[0173] The priority coefficient of the current video frame is determined according to the first pixel average value and the corresponding multiple classification threshold ranges, and the first pixel standard deviation and the corresponding multiple classification threshold ranges.

[0174] Optionally, the first determining module 504 is further configured to:

[0175] If the first pixel average value is within a first threshold range, and / or the first pixel standard deviation is within a second threshold range, the priority coefficient corresponding to the first threshold range and / or the second threshold range is determined as the priority coefficient of the current video frame.

[0176] Optionally, the initial coding information is obtained by coding based on spatial domain information or quality domain information of the video to be coded; the first determining module 504 is further configured to:

[0177] For each video frame included in the video to be encoded, determining the complexity of the video frame in a spatial domain or a quality domain;

[0178] According to the complexity of the video frame, the priority coefficient of the video frame is determined.

[0179] Optionally, the first determining module 504 is further configured to:

[0180] Determining a second pixel average value and a second pixel standard deviation of each pixel point included in the video frame;

[0181] The second pixel average value and the second pixel standard deviation are used as the complexity of the video frame.

[0182] The video encoding device provided in this specification can, after encoding the video to be encoded into a base layer and at least one enhancement layer, additionally determine a priority coefficient for discarding each video frame based on specific video data. The priority coefficient can be determined based on the impact of the video frame on the video quality, and the priority coefficient for discarding each video frame is also added to the encoding information. In the subsequent scheduling process of the encoding information, if it is necessary to discard frame data in the enhancement layer, the priority coefficient for discarding each video frame determined in the encoding process can be used to selectively discard video frames in the enhancement layer according to the level of the priority coefficient. In this way, frame data with little impact on video quality can be discarded first, thereby ensuring video quality and improving user experience.

[0183] The above is a schematic scheme of a video encoding device of this embodiment. It should be noted that the technical scheme of the video encoding device and the technical scheme of the above-mentioned video encoding method are based on the same concept. For details not described in detail in the technical scheme of the video encoding device, please refer to the description of the technical scheme of the above-mentioned video encoding method.

[0184] Corresponding to the above method embodiment, this specification also provides an embodiment of a coding information scheduling device, Figure 6 FIG. 1 shows a schematic diagram of the structure of a coding information scheduling device provided by an embodiment of this specification. Figure 6 As shown, the device includes:

[0185] an acquisition module 602 configured to acquire and parse target coding information of a video to be decoded, and obtain a base layer, at least one enhancement layer, and a discard priority coefficient of each video frame of the video to be decoded, where the priority coefficient indicates a probability of discarding the corresponding video frame;

[0186] A second determining module 604 is configured to determine a video frame to be discarded in at least one enhancement layer according to a current decoding condition and a priority coefficient fed back by a decoding end;

[0187] The scheduling module 606 is configured to use the remaining video frames in the base layer and at least one enhancement layer except the video frames to be discarded as decoding information of the video to be decoded, and schedule the decoding information to the decoding end.

[0188] Optionally, the second determining module 604 is further configured to:

[0189] determining, according to a current decoding condition, a target enhancement layer to be discarded in at least one enhancement layer;

[0190] According to the priority coefficient, the video frames to be discarded are screened out from the video frames included in the target enhancement layer.

[0191] Optionally, the second determining module 604 is further configured to:

[0192] Determine the current number of frames to be discarded based on the current decoding conditions;

[0193] Based on the priority coefficients of the respective video frames included in the target enhancement layer, the target video frames of the current number of frames to be discarded are sequentially screened out in a preset order;

[0194] The filtered target video frames are used as video frames to be discarded.

[0195] Optionally, the preset order is from high to low priority coefficients; the second determining module 604 is further configured to:

[0196] Determine a first video frame having the largest priority coefficient among reference video frames, where the reference video frames are the video frames included in the target enhancement layer;

[0197] If the number of the first video frames exceeds the current number of frames to be discarded, the first video frame of the current number of frames to be discarded is selected as the video frame to be discarded, and the target video frame of the current number of frames to be discarded is obtained;

[0198] If the number of first video frames does not exceed the number of video frames to be discarded, each first video frame is used as a video frame to be discarded, and the video frames to be discarded are continuously filtered from the remaining video frames included in the target enhancement layer until the target video frames with the current number of frames to be discarded are obtained.

[0199] Optionally, the second determining module 604 is further configured to:

[0200] The remaining number of the current number of frames to be discarded, excluding the number of the first video frame, is used as the updated current number of frames to be discarded;

[0201] Using video frames other than the first video frame in each video frame included in the target enhancement layer as reference video frames;

[0202] Return to the step of determining the first video frame with the largest priority coefficient in the reference video frames until the number of determined first video frames exceeds the current number of frames to be discarded, thereby obtaining the target video frames with the current number of frames to be discarded.

[0203] The coding information scheduling device provided in this specification may include, in the target coding information of the video to be decoded, a priority coefficient for discarding each video frame determined according to specific video data during the encoding process, in addition to a basic layer and at least one enhancement layer obtained by encoding the video. The priority coefficient may represent the impact of the video frame on the video quality. Therefore, during the scheduling process, if it is necessary to discard frame data in the enhancement layer, the priority coefficient for discarding each video frame determined during the encoding process may be used to selectively discard video frames in the enhancement layer according to the priority coefficient. This allows frame data with little impact on video quality to be discarded first, thereby ensuring video quality and improving user experience.

[0204] The above is a schematic scheme of a coding information scheduling device according to this embodiment. It should be noted that the technical scheme of the coding information scheduling device and the technical scheme of the coding information scheduling method described above are based on the same concept. For details not described in detail in the technical scheme of the coding information scheduling device, please refer to the description of the technical scheme of the coding information scheduling method described above.

[0205] Figure 7 7 shows a block diagram of a computing device 700 according to an embodiment of the present disclosure. Components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and a database 750 is used to store data.

[0206] The computing device 700 also includes an access device 740 that enables the computing device 700 to communicate via one or more networks 760. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of network interface (e.g., a network interface card (NIC)), whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0207] In one embodiment of the present specification, the above components of the computing device 700 and Figure 7 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 7 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0208] Computing device 700 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or PC. Computing device 700 can also be a mobile or stationary server.

[0209] The processor 720 is configured to execute the following computer-executable instructions to implement steps of any video encoding method or encoding information scheduling method.

[0210] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the aforementioned video encoding method or encoding information scheduling method are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned video encoding method or encoding information scheduling method.

[0211] An embodiment of the present specification further provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, are used to implement steps of any video encoding method or encoding information scheduling method.

[0212] The above is an illustrative embodiment of a computer-readable storage medium. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solution of the aforementioned video encoding method or encoding information scheduling method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the aforementioned video encoding method or encoding information scheduling method.

[0213] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0214] Computer instructions include computer program code, which may be in source code form, object code form, executable files, or some intermediate form. Computer-readable media may include any entity or device capable of carrying computer program code, recording media, USB flash drives, removable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunications signals, and software distribution media. It should be noted that the content of computer-readable media may be appropriately expanded or reduced based on the requirements of legislation and patent practice within a jurisdiction. For example, in some jurisdictions, based on legislation and patent practice, computer-readable media does not include electric carrier signals or telecommunications signals.

[0215] It should be noted that for the aforementioned method embodiments, for ease of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that this specification is not limited to the order of the actions described, because according to this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this specification.

[0216] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0217] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to specific embodiments. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of this specification, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A video encoding method, characterized in that: The method comprises: Acquire a video to be encoded, and encode the video to be encoded into a base layer and at least one enhancement layer to obtain initial encoding information of the video to be encoded; Determining, based on the video data of the to-be-encoded video, a priority coefficient for discarding each video frame of the to-be-encoded video, the priority coefficient being used to represent a probability of the corresponding video frame being discarded, wherein the priority coefficient is determined based on an inter-frame variation amplitude of the video frames or a complexity of the video frames; the smaller the inter-frame variation amplitude, the greater the probability of the corresponding video frame being discarded, and the higher the complexity, the lower the probability of the corresponding video frame being discarded; The priority coefficient of each video frame is added to the initial coding information of the video to be encoded to obtain the target coding information of the video to be encoded.

2. The video encoding method according to claim 1, wherein: The initial coding information is obtained by coding based on the time domain information of the video to be coded; The step of determining, based on the video data of the video to be encoded, a priority coefficient for each video frame included in the video to be encoded, comprises: Determining an inter-frame change amplitude between a current video frame and a previous video frame of the video to be encoded; A priority coefficient of the current video frame is determined according to the inter-frame change amplitude.

3. The video encoding method according to claim 2, wherein: The determining of the inter-frame change amplitude between the current video frame and the previous video frame of the video to be encoded includes: For each pixel in the current video frame, determining a pixel difference between the pixel in the current video frame and the pixel in the previous video frame; Determining a first pixel average and a first pixel standard deviation based on pixel differences between each pixel in the current video frame and the previous video frame; The first pixel average value and the first pixel standard deviation are used as the inter-frame variation amplitude.

4. The video encoding method according to claim 3, wherein: The determining the priority coefficient of the current video frame according to the inter-frame change amplitude includes: The priority coefficient of the current video frame is determined according to the first pixel average value and the corresponding multiple classification threshold ranges, and the first pixel standard deviation and the corresponding multiple classification threshold ranges.

5. The video encoding method according to claim 4, wherein: The determining the priority coefficient of the current video frame according to the first pixel average value and the corresponding multiple grading threshold ranges, and the first pixel standard deviation and the corresponding multiple grading threshold ranges, includes: If the first pixel average value is within a first threshold range, and / or the first pixel standard deviation is within a second threshold range, the priority coefficient corresponding to the first threshold range and / or the second threshold range is determined as the priority coefficient of the current video frame.

6. The video encoding method according to claim 2, wherein: The initial coding information is obtained by coding based on the spatial domain information or quality domain information of the video to be coded; The step of determining, based on the video data of the video to be encoded, a priority coefficient for each video frame included in the video to be encoded, comprises: For each video frame included in the video to be encoded, determining the complexity of the video frame in a spatial domain or a quality domain; A priority coefficient of the video frame is determined according to the complexity of the video frame.

7. The video encoding method according to claim 6, wherein: The determining the complexity of the video frame in the spatial domain or the quality domain includes: Determining a second pixel average value and a second pixel standard deviation of each pixel point included in the video frame; The second pixel average value and the second pixel standard deviation are used as the complexity of the video frame.

8. A coding information scheduling method, characterized in that: The method comprises: Obtaining and parsing target coding information of a video to be decoded, and obtaining a priority coefficient for discarding a base layer, at least one enhancement layer, and each video frame of the video to be decoded, the priority coefficient being used to indicate a probability of discarding the corresponding video frame; Determining, based on the current decoding condition fed back by the decoding end and the priority coefficient, a video frame to be discarded in the at least one enhancement layer, wherein the priority coefficient is determined based on an inter-frame variation amplitude of the video frames or a complexity of the video frames, where a smaller inter-frame variation amplitude is, a greater probability that the corresponding video frame is discarded, and a higher complexity video frame is, a lower probability that the corresponding video frame is discarded; The remaining video frames in the base layer and the at least one enhancement layer except the to-be-discarded video frames are used as decoding information of the video to be decoded, and the decoding information is scheduled to the decoding end.

9. The coding information scheduling method according to claim 8, characterized in that: The determining, based on the current decoding condition fed back by the decoding end and the priority coefficient, the video frame to be discarded in the at least one enhancement layer includes: determining, according to the current decoding condition, a target enhancement layer to be discarded in the at least one enhancement layer; The to-be-discarded video frames are screened out from the video frames included in the target enhancement layer according to the priority coefficient.

10. The coding information scheduling method according to claim 9, characterized in that: The step of screening out the to-be-discarded video frame from the video frames included in the target enhancement layer according to the priority coefficient includes: Determining the current number of frames to be discarded according to the current decoding condition; Based on the priority coefficients of the respective video frames included in the target enhancement layer, sequentially screening out the target video frames of the current number of frames to be discarded in a preset order; The filtered target video frames are used as the video frames to be discarded.

11. The coding information scheduling method according to claim 10, characterized in that: The preset order is from high to low priority coefficient; The step of sequentially selecting the target video frames for the current number of frames to be discarded based on the priority coefficients of the respective video frames included in the target enhancement layer in a preset order includes: Determine a first video frame having a maximum priority coefficient among reference video frames, where the reference video frames are the video frames included in the target enhancement layer; If the number of the first video frames exceeds the current number of frames to be discarded, selecting the first video frame of the current number of frames to be discarded as the video frame to be discarded, and obtaining the target video frame of the current number of frames to be discarded; If the number of the first video frames does not exceed the number of video frames to be discarded, each of the first video frames will be used as the video frame to be discarded, and the video frames to be discarded will continue to be filtered from the remaining video frames included in the target enhancement layer until the target video frames with the current number of frames to be discarded are obtained.

12. The coding information scheduling method according to claim 11, characterized in that: The filtering the to-be-discarded video frames from the remaining video frames included in the target enhancement layer until the target video frames having the current number of to-be-discarded frames are obtained includes: The remaining number of the current number of frames to be discarded, excluding the number of the first video frames, is used as the updated current number of frames to be discarded; using the video frames other than the first video frame among the video frames included in the target enhancement layer as the reference video frames; Return to the step of determining the first video frame with the largest priority coefficient in the reference video frames until the number of determined first video frames exceeds the current number of frames to be discarded, and obtain the target video frames with the current number of frames to be discarded.

13. A video encoding device, characterized in that: The device comprises: The encoding module is configured to obtain a video to be encoded, and encode the video to be encoded into a base layer and at least one enhancement layer, to obtain initial encoding information of the video to be encoded; a first determining module configured to determine, based on the video data of the video to be encoded, a priority coefficient for discarding each video frame of the video to be encoded, the priority coefficient being used to represent a probability of the corresponding video frame being discarded, wherein the priority coefficient is determined based on an inter-frame variation amplitude of the video frames or a complexity of the video frames, wherein a smaller inter-frame variation amplitude corresponds to a greater probability of the corresponding video frame being discarded, and a higher complexity video frame corresponds to a lower probability of being discarded; The adding module is configured to add the priority coefficient of each video frame to the initial coding information of the video to be encoded to obtain the target coding information of the video to be encoded.

14. A coding information scheduling device, characterized in that: The device comprises: an acquisition module configured to acquire and parse target coding information of a video to be decoded, and obtain a priority coefficient for discarding a base layer, at least one enhancement layer, and each video frame of the video to be decoded, wherein the priority coefficient is used to represent a probability of discarding the corresponding video frame, wherein the priority coefficient is determined based on an inter-frame variation amplitude of the video frame or a complexity of the video frame; the smaller the inter-frame variation amplitude, the greater the probability of discarding the corresponding video frame, and the higher the complexity, the lower the probability of discarding the video frame; A second determining module is configured to determine a video frame to be discarded in the at least one enhancement layer according to a current decoding condition fed back by a decoding end and the priority coefficient; The scheduling module is configured to use the remaining video frames in the base layer and the at least one enhancement layer except the to-be-discarded video frames as decoding information of the video to be decoded, and schedule the decoding information to the decoding end.

15. A computing device comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the video encoding method according to any one of claims 1 to 7 or the encoding information scheduling method according to any one of claims 8 to 12.

16. A computer-readable storage medium storing computer instructions, which, when executed by a processor, implement the steps of the video encoding method according to any one of claims 1 to 7 or the encoding information scheduling method according to any one of claims 8 to 12.

Citation Information

Patent Citations

  • Method and device for determining priority to schedule packets

    CN101895461A

  • Method and device for transmitting video frame

    CN106792264A