Neural video coding and decoding
By determining the quantization and inverse quantization scaling values in the neural video codec, the upper and lower limit values learned during the training process of the video codec model are solved, and the problem of the limited quality range of neural video codecs is achieved is achieved, and the coordination and efficiency of encoding and quantization are improved.
Patent Information
- Application Number
- CN202410070123.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-18
AI Technical Summary
Existing neural video codecs have limited quality ranges and cannot meet the needs of practical products, making it difficult to support multiple quality levels within a wider quality range.
By determining the quantized scaling value and inverse quantized scaling value of the target frame, the upper and lower limit values learned by the video codec model during the training process can achieve flexible selection of quantized scaling value, improve the coordination between encoding and quantization, and support a wider range of video quality.
It realizes more flexible quantized scaling value selection during the video encoding and decoding process, improves the coordination between encoding and quantization, supports a wider video quality range and code rate range, and improves the efficiency and quality of video encoding and decoding.
Smart Images

Figure CN120343247A_ABST
Abstract
Description
Background Art
[0001] Traditional codecs rely on a hybrid residual-based framework, which has been developed for many years and is still being continuously improved. However, the improvement of the compression ratio of traditional codecs has decreased, while the complexity has increased significantly. This makes further improvement within the traditional framework increasingly challenging. Currently, neural video codecs (NVCs) have received significant attention because they have the potential to break through this development bottleneck. Summary of the Invention
[0002] According to an implementation of the present disclosure, a neural video coding and decoding scheme is proposed. In this scheme, a quantization parameter for a target frame in a video is obtained; based on the quantization parameter and at least a first upper limit value and a first lower limit value for quantization scaling, a quantization scaling value for the target frame is determined, where the quantization scaling value is between the first upper limit value and the first lower limit value, and the first upper limit value and the first lower limit value are determined during the training process of the video codec model; and the video codec model is used to perform quantization on the target frame based on the quantization scaling value during at least one of the frame encoding process or the motion vector encoding process of the target frame to obtain a quantized representation of the target frame. According to an embodiment of the present disclosure, a quantization scaling value uniform sampling mechanism is proposed, and the range of quantization scaling values used for sampling is learned during the model training process. In this way, a single model can be trained to support more flexible selection of quantization scaling values during video coding and decoding, improve the coordination between encoding and quantization, and support a wider range of video quality.
[0003] This section is provided to introduce the selection of objects in a simplified form, which will be further described in the specific implementation manners below. This section is not intended to identify the key features or main features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Brief Description of the Drawings
[0004] Figure 1 A block diagram showing an example environment in which multiple implementations of the present disclosure can be implemented;
[0005] Figure 2 A schematic block diagram showing a partial structure of a video codec model according to some implementations of the present disclosure;
[0006] Figure 3 A schematic block diagram showing an example architecture of a frame coding and decoding model according to some implementations of the present disclosure;
[0007] Figure 4A A diagram showing bitrate control in a high bitrate scenario according to some implementations of the present disclosure;
[0008] Figure 4BShows bitrate control in a low bitrate scenario according to some implementations of the present disclosure;
[0009] Figure 5 Shows a schematic block diagram of an example architecture of a context extraction model according to some implementations of the present disclosure;
[0010] Figure 6 Shows a comparison of the quality range with that of other codecs according to some implementations of the present disclosure;
[0011] Figure 7 Shows a flowchart of a process for video processing according to some implementations of the present disclosure; and
[0012] Figure 8 Shows a schematic block diagram of an electronic device capable of implementing multiple implementations of the present disclosure.
[0013] In these drawings, the same or similar reference signs are used to denote the same or similar elements. Detailed Description
[0014] The present disclosure will now be described with reference to several example implementations. It should be understood that these implementations are described only to enable those of ordinary skill in the art to better understand and thus implement the present disclosure, rather than implying any limitation on the scope of the present disclosure.
[0015] As used herein, the term "comprising" and its variants are to be construed as open-ended terms meaning "including but not limited to". The term "based on" is to be construed as "at least partially based on". The terms "one implementation" and "an implementation" are to be construed as "at least one implementation". The term "another implementation" is to be construed as "at least one other implementation". The terms "first", "second", etc. may refer to different or the same objects. There may be other explicit and implicit definitions hereinafter.
[0016] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning (DL) is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network (NN) model is an example of a model based on deep learning. In this document, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably herein.
[0017] Generally, machine learning can roughly include three stages, namely, the training stage, the testing stage, and the usage stage (also known as the inference stage). In the training stage, a given model can be trained using a large amount of training data and iterated continuously until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also known as the input-to-output mapping) from the training data. The parameter values of the trained model are determined. In the testing stage, test inputs are applied to the trained model to test whether the model can provide correct outputs, thereby determining the performance of the model. In the inference stage, the model can be used to process actual inputs based on the parameter values obtained through training and determine the corresponding outputs.
[0018] Example Environment
[0019] Figure 1 FIG. shows a block diagram of an example environment 100 capable of implementing multiple implementations of the present disclosure. In Figure 1 this environment, the electronic device 110 includes a video codec 112 configured to implement video encoding and / or decoding. The electronic device 120 includes a video codec 122 configured to implement video encoding and / or decoding. The video codec 112 or 122 may include an encoder and / or a decoder. During the encoding process, the encoder may encode the video 130 into a bitstream 132. During the decoding process, the decoder may decode the bitstream 132 into the video 130.
[0020] The electronic devices 110 and 120 can communicate with each other through an appropriate communication network. In some codec environments, the electronic device 110 and the electronic device 120 can perform video communication, and both the video codecs 112 and 122 can have video encoding and decoding functions. For example, the electronic device 110 may provide the bitstream obtained after encoding the video to the electronic device 120 for decoding, and the electronic device 120 may decode the received bitstream to obtain the corresponding video. Additionally, the electronic device 120 may also provide the encoding result of the video to the electronic device 110 for decoding. In some codec environments, the video codec 112 in the electronic device 110 may include an encoder for encoding the video into a bitstream. The electronic device 120 may include a video playback tool, and the video codec 122 therein includes a decoder for decoding the bitstream generated by the video codec 112 to obtain the video for playback.
[0021] It should be understood that Figure 1 the devices and elements shown are merely examples. In practical applications, there may be more electronic devices, and each electronic device may have video encoding and / or decoding functions.
[0022] Early neural video codecs (NVCs) still follow traditional codecs and use a residual coding-decoding based framework, where all sub-modules are replaced by neural networks to enable end-to-end learning. Some later works also proposed stronger sub-modules based on this paradigm to improve performance.
[0023] Conversely, conditional coding-decoding was proposed in the application of NVCs, which shows a lower entropy bound than residual coding-decoding and has greater potential. The conditions in conditional coding-decoding can be freely defined and learned, rather than being limited to predicted frames in the pixel domain. At the same time, conditional coding-decoding can be flexibly used to assist in coding, decoding, and entropy modeling. Currently, NVCs based on conditional coding-decoding have achieved better compression ratios than traditional codec models by mining different spatial and temporal contexts as conditions.
[0024] Although some NVCs have made progress, for example, supporting multiple quality levels in a single model, they are still not practical because their quality range is very limited and cannot meet the requirements of practical products. Therefore, it is expected that NVCs support multiple quality levels within a wider quality range.
[0025] According to an example implementation of the present disclosure, an improved video coding-decoding scheme is proposed. This scheme is implemented based on NVC. In this scheme, for a target frame to be coded and decoded, a quantization scaling value of the target frame is determined based on the quantization parameter of the target frame and a first upper limit value and a first lower limit value for quantization scaling, and the quantization scaling value is between the first upper limit value and the first lower limit value. The determined quantization scaling value is used to obtain a quantized representation of the target frame. In neural video coding-decoding, the first upper limit value and the first lower limit value determined during the training process of the video codec model are used to implement the selection of the quantization scaling value for the target frame. In this way, a single model can be trained to support more flexible selection of quantization scaling values during video coding-decoding, improve the coordination between coding and quantization, and support a wider video quality range.
[0026] The architecture of the video codec model
[0027] Some example implementations of the present disclosure will be described in more detail below with reference to the accompanying drawings.
[0028] Figure 2 A schematic block diagram showing a partial structure of a video codec model 200 according to some implementations of the present disclosure is shown. Each component in the video codec model 200 can be implemented by hardware, software, firmware, or any combination thereof. The video codec model 200 can be implemented in Figure 1 video codecs 112 and / or 122. Note that Figure 2Only a part of the entire video codec model is shown, and there may be many other components in the model.
[0029] Video codec model 200 generally includes a motion vector codec (including motion estimation) model 210, a context extraction model 220, and a frame codec model 230. These models can be implemented based on machine learning techniques, such as based on a neural network architecture. To obtain a higher compression ratio, video codec model 200 is implemented based on more flexible conditional coding and decoding, by extracting context information as a condition to guide the coding and decoding of frames.
[0030] Video codec model 200 can be implemented to convert between individual frames in a video and the bitstream of the video. This conversion includes the encoding process of the video, the decoding process of the video, or both. In the encoding process, video codec model 200 receives a series of frames of the video and can perform video encoding on each frame to obtain the bitstream of the video. In the decoding process, video codec model 200 receives the bitstream of the video and decodes a series of frames of the video from it. In the following, except for the operations specifically described, other operations can be considered to be executable on both the encoding and decoding sides of the video. For ease of understanding, the basic working principle of video codec model 200 will be briefly described below.
[0031] In this article, the target frame x t refers to the current frame in the video to be coded and decoded. As Figure 2 shown, to code and decode the target frame x t with index t, the motion vector codec model 210 is configured to determine the motion information v t of the target frame x t based on the quantization parameter q t of the target frame, encode the motion information v t and then decode it into estimated motion information The foregoing are operations implemented on the encoding side of the video. On the decoding side of the video, the encoded motion information v t can be transmitted to the decoding side for decoding the estimated motion information The processing of the motion vector codec model 210 is represented as f motion .
[0032] The motion information indicates the motion offset of the elements in the target frame x t relative to the reference frame, including the offset size and direction. For example, the motion information may include a motion vector (MV). The reference frame can also be used as the input for the motion vector encoding of the target frame x t . In some implementations, the reference frame of the target frame x t can be the target frame x tOne or more previous frames. In the following implementation, only a single reference frame is taken as an example for explanation, although multiple reference frames are also feasible. As Figure 2 shown, the target frame x t The previous frame x t-1 Or its decoded and reconstructed result is used as a reference frame. In some implementations, the motion vector codec model 210 can be configured to determine the estimated motion information based on an optical flow network In addition to the optical flow network, the motion vector codec model 210 can also be implemented based on any other suitable model capable of determining the motion information of a frame.
[0033] The context extraction model 220 is configured to determine the context information C for the target frame x t of t . In the implementation of the present disclosure, the context extraction model 220 is configured to determine the context information C for the target frame x based on the estimated motion information of the target frame x t as well as the relevant information of the reference frame (frame x ) or its decoded and reconstructed t-1 . The relevant information of the reference frame utilized includes the reference feature information F for the reference frame t as well as the reference reconstructed frame t . The reference feature information F t-1 and the reference reconstructed frame The reference frame x t-1 of the reference feature information F t-1 can characterize the characterization of the feature parameters of the reference frame in the feature space of the reference frame. The reference reconstructed frame is the reconstruction result of the reference frame x t-1 . Each frame is reconstructed during both the encoding and decoding processes. The reference feature information F t-1 and the reference reconstructed frame can complement each other, thereby providing more rich and relevant context information for the target frame x t . In addition, based on the estimated motion information of the target frame x t the context extraction model 220 can extract the motion-aligned context information C . The processing of the context extraction model 220 is represented as f t . Tcontext .
[0034] The frame codec model 230 is configured to generate the target reconstructed frame of the target frame x based on at least the quantization parameter q of the target frame t as well as the context information C t . t In addition, the frame codec model 230 is also configured to generate the target feature information F of the target frame based on at least the context information C . t oft 。The target feature information F t and the target reconstruction frame are cached and passed for encoding and decoding of the next frame until the encoding and decoding of the entire video are completed. On the decoding side, the target reconstruction frame can be output as the decoding target.
[0035] On the encoding side, the frame encoding and decoding model 230 can use the context information C t as a condition to encode the target frame x t into a quantized encoding representation On the decoding side, the quantized encoding representation can be determined from the bitstream of the video After the entropy encoding of the quantized encoding representation the target reconstruction frame and the target feature information F t are reconstructed based on the entropy encoding result. The processing of the frame encoding and decoding model 230 is denoted as f frame .
[0036] Figure 2 shows the encoding and decoding pipeline of the video codec model 200 for each frame x t-2 , x t-1 , x t ... etc. For each frame, the motion vector encoding and decoding model 210, the context extraction model 220, and the frame encoding and decoding model 230 in the model 200 all perform similar operations.
[0037] It should be understood that Figure 2 shows that the video codec model 200 is part of the codec and not a complete codec.
[0038] The above generally introduces the workflow of the video codec model 200. The following will discuss in more detail the example implementations of performing quantization, inverse quantization, and determining quantization scale values and inverse quantization scale values.
[0039] Example implementations of performing quantization and inverse quantization
[0040] In the example implementation of the present disclosure, an improved scheme for neural video encoding and decoding is proposed. During the video encoding and decoding process, it is usually necessary to perform quantization and inverse quantization on video frames.
[0041] The basic quantization and inverse quantization processes can be expressed as follows:
[0042]
[0043] where I represents the input value, QS represents the quantization step size, represents the rounding operation. QS is used to control the reconstruction quality of the output .
[0044] In this scheme, for the target frame x to be encoded and decoded t , based on the target frame x t The quantization parameter q t and the first upper limit value and the first lower limit value for quantization scaling to determine the target frame x t The quantization scaling value is determined to be between a first upper limit value and a first lower limit value, and the first upper limit value and the first lower limit value are determined during the training process of the video codec model 200. The determined quantization scaling value is used to obtain the target frame x t Quantitative representation of .
[0045] Similarly, when we get the target frame x t After the quantization representation, in the process of inverse quantization, it can also be based on the quantization parameter q t and the second upper limit value and the second lower limit value for inverse quantization scaling to determine the target frame x t The inverse quantization scaling value is determined to be between a second upper limit value and a second lower limit value, and the second upper limit value and the second lower limit value are determined during the training process of the video codec model 200. The determined inverse quantization scaling value is used to perform inverse quantization on the quantized encoding representation corresponding to the quantized representation.
[0046] In the neural video codec, the selection of the quantization scaling value for the target frame is implemented by using the first upper limit value, the first lower limit value, the second upper limit value, and the second lower limit value determined during the training process of the video codec model 200. In this way, by sampling the quantization scaling value and the inverse quantization scaling value, the video codec model 200 can experience various trade-offs between bit rate and distortion, which can further improve the coordination between encoding and quantization and support a wider quality range and bit rate range.
[0047] The quantization and dequantization mechanism proposed in the example implementation of the present disclosure can be applied in the frame codec model 230 and the motion vector codec model 210 of the video codec model 200. As an example, the example implementation of the quantization and dequantization mechanism proposed in the present disclosure is described in detail below in conjunction with the frame codec model 230. It can be understood that the quantization and dequantization mechanism proposed in the present disclosure will be similarly applied in the motion vector codec model 210.
[0048] Figure 3 2 shows a schematic block diagram of an example architecture of a frame encoding and decoding model 230 according to some implementations of the present disclosure. Figure 3 In the example of high 312, quantizer 313 and E low 314, of which E highEncoder 312 is configured to process high - resolution, E low Encoder 314 is configured to process low - resolution. Quantizer 313 is configured to perform quantization on the output result of E high and provide the quantization result as an input to Encoder 314 low Decoder 320 may include D high 322, de - quantizer 323 and D low 324, where D high 322 is configured to process high - resolution decoder, D low 324 is configured to process low - resolution decoder. De - quantizer 323 is configured to perform de - quantization on the output result of D low and provide the de - quantization result as an input to D high 322. On the decoding side of the video, Encoder 310 may be omitted. In some implementations, Encoder 310 may include any number of encoders, and Decoder 320 may include a corresponding any number of decoders.
[0049] As Figure 3 shown, to perform quantization on the target frame x t it is necessary to obtain the quantization parameter q for the target frame x in the video t Here, the quantization parameter q t may be set by the user, or may be the default setting in the video codec model 200. t For example, it can be set by the user or be the default setting in the video codec model 200.
[0050] The choice of quantization scaling has an impact on video quality and compression efficiency. The first upper limit value and the first lower limit value can ensure that the video quality and compression efficiency for the target frame x t are within an appropriate range. Based on the quantization parameter q t the first upper limit value and the first lower limit value the quantization scaling value for the target frame x t can be determined The quantization scaling value is between the first upper limit value and the first lower limit value and the first upper limit value and the first lower limit value are determined during the training process of the video codec model 200. The quantization scaling value can be used as a scaling implicit feature.
[0051] After determining the quantization scaling value After that, the frame encoding and decoding model 230 in the video codec model 200 is used to perform operations on the target frame x t During the frame encoding process of, the quantization scaling value can be used to quantize the output result of E high 312 to obtain the quantized representation y t of the target frame x t .
[0052] In some implementations, the quantization scaling value can be used to quantize the output result of E low 314 to obtain the quantized representation of the target frame x t . The quantization scaling value can perform quantization on the target frame x t at any position in the encoder 310, and the implementation of the present disclosure is not limited in this regard.
[0053] Similarly, the above operation of performing quantization on the target frame x t during the frame encoding process can be applied to the motion vector encoding and decoding model 210 in the video codec model 200. Using the motion vector encoding and decoding model 210, during the motion estimation and motion vector encoding process of the target frame x t , quantization can be performed on the target frame x based on the quantization scaling value to obtain the quantized representation y t of the target frame x t t .
[0054] Although Figure 3 shows the quantizer 313 as being between E high 312 and E low 314, according to the specific model design, the quantizer 313 can also be located after E low 314. In addition, although Figure 3 only shows a single quantizer, in the actual model design, the encoder 310 can include multiple neural network processing layers, and similar quantizers can be deployed in multiple network processing layers. For the quantizers deployed in different network processing layers, the upper and lower limit values of their quantization scaling can be configured as a pair of the same values, or multiple pairs of different values can be learned during the training process of the video codec model 200.
[0055] In some implementations, as Figure 3 shown, after performing quantization on the target frame x t , corresponding dequantization needs to be performed to restore the original target frame x t The information. The selection of the inverse quantization scaling has an impact on the recovery of the original video information. Regarding the second upper limit value of the inverse quantization scaling and the second lower limit value can ensure that the information recovery for the target frame x t is within an appropriate range. Based on the quantization parameter q t , the second upper limit value and the second lower limit value can determine the inverse quantization scaling value for the target frame x t The inverse quantization scaling value is between the second upper limit value and the second lower limit value and the second lower limit value and the second upper limit value and the second lower limit value are determined during the training process of the video codec model 200.
[0056] After determining the inverse quantization scaling value , during the frame encoding process of the target frame x t , the inverse quantization scaling value can be used to perform inverse quantization on the output result of D low 324, so as to perform inverse quantization on the quantization encoding representation t corresponding to the quantization representation y Execute inverse quantization.
[0057] In some implementations, the inverse quantization scaling value can be used to perform inverse quantization on the input of D low 324, so as to perform inverse quantization on the quantization encoding representation t corresponding to the quantization representation y Execute inverse quantization. The inverse quantization scaling value can perform inverse quantization on the quantization encoding representation at any position of the decoder 320, and the implementation of the present disclosure is not limited in this regard.
[0058] Similarly, the operation of performing inverse quantization on the quantization encoding representation during the frame encoding process as described above can be applied to the motion vector encoding and decoding model 210 in the video codec model 200. During the motion estimation and motion vector encoding process of the target frame x t , the quantization encoding representation can be inverse quantized based on the inverse quantization scaling value Execute inverse quantization.
[0059] Although Figure 3 shows the inverse quantizer 323 as being in D high 322 and D lowbetween 324, but according to the specific model design, the dequantizer 323 may also be located before D low before 324. In addition, although Figure 3 only a single dequantizer is shown in, in the actual model design, the decoder 320 may include multiple neural network processing layers, and similar dequantizers may be deployed in multiple network processing layers. For the dequantizers deployed in different network processing layers, the upper limit value and the lower limit value of the dequantization scaling can be configured as a pair of the same value, or multiple pairs of different values can be learned during the training process of the video codec model 200.
[0060] Example implementation of determining quantization scaling value and dequantization scaling value
[0061] In some implementations, an interpolation function can be utilized, at least based on the quantization parameter q t between the first upper limit value and the first lower limit value to perform interpolation to obtain the quantization scaling value The interpolation function can be expressed as a function related to q t , and as follows:
[0062]
[0063] where f() represents the interpolation function, and f() can be, for example, an exponential function, a logarithmic function, a linear function, a quadratic function, etc. or any combination thereof, which has monotonicity between the first upper limit value and the first lower limit value In this way, the quantization scaling value can be sampled between the first upper limit value and the first lower limit value so that the video codec model 200 can experience different quantization scaling values and various trade - offs between bitrate and distortion, improving the coordination between encoding and quantization and supporting a wider quality range.
[0064] In some implementations, the quantization parameter q t can be selected from the quantization parameter range [0, q_num - 1], where q_num represents the number of values that the quantization parameter q t can select. q_num can be input by the user or pre - set in the video codec model 200. The value of q_num can be set to 64, for example. After selecting the quantization parameter q t from the quantization parameter range, an interpolation function can be utilized, based on the quantization parameter q tand the upper limit value q_num - 1 of the quantization parameter range at the first upper limit value and the first lower limit value perform interpolation to obtain the quantization scaling value An example of the interpolation function is as follows:
[0065]
[0066] In some implementations, the interpolation function can be configured such that: as the quantization parameter q t linearly increases within the quantization parameter range [0, q_num - 1], the quantization scaling value from the first lower limit value to the first upper limit value exponentially increases. In some implementations, the interpolation function is configured in the form of an exponential function or a logarithmic function. The exponent of the exponential function is determined based on the quantization parameter q t and the quantization parameter range [0, q_num - 1], and the base of the exponential function is determined based on the first upper limit value and the first lower limit value to determine.
[0067] Specifically, the interpolation function f() can be an exponential function, and an example is as follows:
[0068]
[0069] In some implementations, the exponential function of formula (4A) can also be represented in the form of a logarithmic function, and examples are as follows:
[0070]
[0071] Using the interpolation function based on logarithmic operations given by formula (4B) can provide better numerical stability and support uniform sampling in the logarithmic domain.
[0072] Some example interpolation functions are given above, but it can be understood that in different applications, other interpolation functions can also be involved as needed. In some implementations, in addition to exponential functions or logarithmic functions, other functions such as linear interpolation functions can also be used. As an example, the linear interpolation function can be:
[0073]
[0074] In the example implementations above, learning the first upper limit value and the first lower limit value For limiting the range of quantization scaling values. In some embodiments, an interpolation function related to more than two parameters can also be designed, and the values of these parameters can be learned during the learning process of the video codec. For example, the interpolation of the quantization scaling value can be a line segment interpolation function, which includes two or more interpolation functions for characterizing two or more interpolation scaling value intervals falling therein. For quantization parameter q t falling into different value ranges, different interpolation functions can be used to determine.
[0075] Assume that the line segment interpolation includes two line segment interpolation functions, and the parameter values to be learned include where is the intermediate value for quantization scaling. For quantization parameter q mid falling into the quantization parameter range [0, q t , the first line segment interpolation function related to and can be used to determine the quantization scaling value This interpolation function can be in a form similar to the above formula (4A), (4B) or (5), except that is replaced by For quantization parameter q mid falling into the quantization parameter range [q t , q_num - 1], the second line segment interpolation function related to and can be used to determine the quantization scaling value This interpolation function can be in a form similar to the above formula (4A), (4B) or (5), except that is replaced by The types of the first interpolation function and the second interpolation function can be the same or different.
[0076] In some implementations, after using the interpolation function to determine the quantization scaling value , the same interpolation function can be used to perform interpolation between the second upper limit value t and the second lower limit value based on the quantization parameter q to obtain the inverse quantization scaling value For example, if the interpolation function used when determining the quantization scaling value t based on the quantization parameter q is the interpolation function shown in formula (2), then formula (2) can be used to perform interpolation between the second upper limit value t and the second lower limit value based on the quantization parameter q to obtain the inverse quantization scaling value
[0077] In some implementations, after determining the quantization scaling value using an interpolation function the same interpolation function can be used to perform interpolation between a second upper limit value and a second lower limit value based on the quantization parameter q t and the upper limit value q_num - 1 of the quantization parameter range to obtain an inverse quantization scaling value and a second lower limit value For example, if the interpolation function used when determining the quantization scaling value based on the quantization parameter q and the upper limit value q_num - 1 of the quantization parameter range is the interpolation function shown in formula (4B), then formula (4B) can be used to perform interpolation between a second upper limit value t and the upper limit value q_num - 1 of the quantization parameter range to obtain an inverse quantization scaling value The interpolation function used can also satisfy that as the quantization parameter q t linearly increases within the quantization parameter range [0, q_num - 1], the inverse quantization scaling value exponentially increases from the second lower limit value to the second upper limit value The quantization and inverse quantization of the implementation of the present disclosure are performed in the implicit feature domain, and do not require the quantization scaling value used during encoding t to be the same as the inverse quantization scaling value used during decoding. The inverse quantization scaling value corresponding to the quantization scaling value from the second lower limit value to the second upper limit value exponentially increases.
[0078] The quantization and inverse quantization of the implementation of the present disclosure are performed in the implicit feature domain, and do not require the quantization scaling value used during encoding to be the same as the inverse quantization scaling value used during decoding. The inverse quantization scaling value corresponding to the quantization scaling value is independently learned during decoding, so the codec model 200 can have greater flexibility to smoothly adjust the quality. In addition, during the encoding process, the quantization scaling value can be multiplied by the implicit feature instead of being used as a divisor to make the training process of the video codec model 200 more stable by avoiding potential division by zero problems. During the training process of the video codec model 200, the loss function can be used as the objective function for optimizing the video codec model 200. By minimizing the loss function, the parameters of the video codec model 200 can be updated to optimize the video codec model 200. In some implementations, the loss function used by the video codec model 200 is determined by weighting between the bitrate loss and the distortion loss using a weighting factor. The weighting factor can be based on the quantization parameter q a weighting upper limit value λ
[0079] During the training process of the video codec model 200, the loss function can be used as the objective function for optimizing the video codec model 200. By minimizing the loss function, the parameters of the video codec model 200 can be updated to optimize the video codec model 200. In some implementations, the loss function used by the video codec model 200 is determined by weighting between the bitrate loss and the distortion loss using a weighting factor. The weighting factor can be based on the quantization parameter q t a weighting upper limit value λ max and a weighting lower limit value λ minIt is determined by. The weighting factor is used to control the trade-off between bitrate loss and distortion loss.
[0080] Specifically, in the loss function, the weighting factor can be used to weight the distortion loss. An example of the loss function is as follows:
[0081] Loss RD = R + λ D (6)
[0082] Where R represents the bitrate loss, D represents the distortion loss, and λ represents the weighting factor.
[0083] In some implementations, in the loss function, the weighting factor can also be used to weight the bitrate loss. An example of the loss function is as follows:
[0084] Loss RD =λ R +D (7)
[0085] The training of the video codec model 200 is to continuously optimize the model parameters to continuously reduce the value of the loss function until the value of the loss function is minimized or reaches a predetermined target.
[0086] In some implementations, the same interpolation function used in the determination of the quantization scaling value can be used to perform interpolation based on the quantization parameter q t between the weighting upper limit value λ max and the weighting lower limit value λ min to obtain the weighting factor λ. For example, if the interpolation function used in the determination of the quantization scaling value is the interpolation function shown in formula (4B), then the interpolation function used to determine the weighting factor λ is as follows:
[0087]
[0088] To enable the video codec model 200 to support variable quality levels, the weighting factor λ is also variable during training. Exemplarily, the weighting upper limit value λ max and the weighting lower limit value λ min are hyperparameters. The weighting upper limit value λ max can be set to 768, and the weighting lower limit value λ min can be set to 1. For each training step, the quantization parameter q is uniformly sampled from the quantization parameter range [0, q_num - 1] t , and then the weighting factor λ can be obtained via the interpolation function (8), and the value of the loss function can be calculated via the loss function (6) or (7). Through the backpropagation of the value of the loss function, the range between the weighting lower limit value λ min and the weighting upper limit value λ max is [λ min, λ max will guide the first upper limit value and the first lower limit value within the range of learning. Therefore, by controlling the range of the weighting factor λ value, the quantization scaling value range can be adjusted, enabling the video codec model 200 to experience various trade-offs between bitrates and distortion, thereby enhancing the coordination between encoding and quantization and supporting a wider quality range.
[0089] For all spatial positions, the quantization scaling value is the same, which may overlook the spatial characteristics of the video content. Therefore, a spatially-channel-oriented quantization scaling value that is adjusted for different positions is also required such as Figure 3 shown, the entropy model 338 determines the spatially-channel-oriented quantization scaling value based on the quantization-coded representation t-1 of the reference frame x hyperprior information and / or the context information C t of the target frame x t The spatially-channel-oriented quantization scaling value can be used to quantize the quantization representation y t The quantization result can be provided to the arithmetic encoder (AE) 332. The AE 332 generates the bitstream 334 of the video. The bitstream 334 is decoded by the arithmetic decoder (AD) 336 to obtain the quantization-coded representation On the decoding device side of the video, the AE 332 may not be required, and the decoding device can start from the received bitstream 334 and decode the quantization-coded representation with the AD 336 and provide it to the decoder 320. In this way, the spatially-channel-oriented quantization scaling value not only helps to achieve precise adjustment at each spatial position but is also adaptive to the video content of each frame, improving the final compression efficiency.
[0090] Based on the above implementation, the video codec model according to some implementations of the present disclosure can support quality adjustment within a wider quality range, which is also a prerequisite for controlling the bitrate, and controlling the bitrate is a core function of the implemented codec. Figure 4A Figure 4B and Figure 4B respectively illustrate bitrate control in high-bitrate scenarios and low-bitrate scenarios according to some implementations of the present disclosure. As Figure 4A shown, the curve 410 shows the variation of the target bitrate with the frame index, and the curve 412 shows the variation of the actual bitrate with the frame index. As Figure 4BAs shown, curve 420 shows the variation of the target bit rate with the frame index, and curve 422 shows the variation of the actual bit rate with the frame index. Since the video codec model according to some implementations of the present disclosure can adjust the quantization parameter qt for each frame, a fluctuating target bit rate can be supported, and the actual bit rate can be close to the target bit rate. It can be seen that the video codec model according to some implementations of the present disclosure supports bit rate control, can control the video quality, balance the video quality and data transmission requirements, and improve the user experience.
[0091] Example implementation of context extraction
[0092] In addition, in NVC, long prediction chains can cause temporal error accumulation, resulting in a decrease in video quality. To alleviate this problem, one solution is to use a small intra-frame period setting to insert high-quality I (intra) frames more frequently, but this will reduce the compression efficiency. Therefore, it is also desirable to reduce error accumulation in NVC.
[0093] Temporal quality degradation is a fundamental problem for all video codecs, but it is more severe for NVC. Therefore, it is desirable to reduce the temporal quality degradation caused by temporal error accumulation. Figure 5 A schematic block diagram showing an example architecture of a context extraction model 220 according to some implementations of the present disclosure is shown. As Figure 5 shown, the context extraction model 220 may include a feature extractor 510 for the reference reconstructed frame , a feature extractor 520 for the reference feature information F t-1 of the reference reconstructed frame, a refresh controller 530, a motion alignment unit 540, and a context generation unit 550.
[0094] Generally, the context extraction model 220 performs feature extraction on the reference feature information F t-1 through the feature extractor 520, and then generates the context information C t of the target frame x t . The accumulated error may affect the propagated features (e.g., the reference feature information F t-1 ), or the propagated features contain some irrelevant information, resulting in a lower quality of the context information C t .
[0095] Since the reference reconstructed frame of the target frame x t has only three-dimensional information and has less information compared to the reference feature information F of the reference reconstructed frame, therefore, from the reference reconstructed frame t-1 compared to Feature information with less temporal cumulative error can be extracted. In some implementations, using the context extraction model 220 in the video codec model 200, during the context extraction process of the target frame x t it is necessary to determine whether to extract the feature information for the target frame x t from the reference reconstruction frame of the target frame x t or from the reference feature information F t-1 extracted from the target frame x t of the feature information.
[0096] Exemplarily, it can be determined whether to extract the feature information from the reference reconstruction frame or from the reference feature information F t-1 by setting an update period. For example, assuming that after 32 cycles, a large amount of temporal cumulative error will be included in the reference feature information F t-1 the update period can be set to 32. When the update period expires, using the refresh controller 530, via the feature extractor 510, extract the feature information for the target frame x from the reference reconstruction frame t ; when the update period has not expired, using the refresh controller 530, via the feature extractor 520, extract the feature information for the target frame x t-1 from the reference feature information F t of the target frame x
[0097] In some implementations, the context information update period can be determined based on the content type of the video. Exemplarily, if the video has less picture changes and relatively stable content (for example, a landscape video containing natural scenery, a still life video with a static object as the main shooting object, a surveillance video for security monitoring, etc.), then the context information of this video changes less, and a longer context information update period can be set for this video. If the video has relatively frequent picture changes and relatively active content (for example, a sports video, a special effects video, a music video, etc.), then the context information of this video changes more, and a shorter context information update period can be set for this video.
[0098] Exemplarily, it can also be determined whether to extract the feature information from the reference reconstruction frame or from the reference feature information F t-1 by constructing a neural network. By training this neural network, the neural network can determine whether to extract from the reference reconstruction frame according to the input video Extract feature information, still from the reference feature information F t-1 Extract feature information.
[0099] After extracting the feature information for the target frame x t After that, the context information C of the target frame x can be determined based on the feature information t t As Figure 5 shown, after using the refresh controller 530 to extract the feature information of the target frame x t After that, the feature information can be fed to the motion alignment unit 540, and the motion alignment unit 540 is based on the feature information and the target frame x t Estimated motion information Processed, and the motion-aligned feature information is provided to the context generation unit 550 to generate the context information C of the target frame x t t The context information C t Is used during the frame encoding process of the target frame x t In this way, by setting the context information update period and periodically refreshing the feature information for the target frame, the problem of long-term temporal error accumulation can be alleviated, and the quality of video coding and decoding can be improved.
[0100] In some implementations, it can be determined whether to extract the feature information from the reference reconstruction frame t By determining whether the difference between multiple reconstruction frames before the target frame x exceeds a difference threshold Extract feature information, or from the reference feature information F t-1 Extract feature information. The frame index of the target frame x t Is t, and multiple reconstruction frames before the target frame x t For example, can be the reconstruction frame x with index t-1 t-1 、The reconstruction frame x with index t-2 t-2 、The reconstruction frame x with index t-3 t-3 Etc. Exemplarily, the difference between multiple reconstruction frames before the target frame x t Can be the difference between the reconstruction frame x t-1 And the reconstruction frame x t-2 The difference between, or can also be the reconstruction frame x t-2 And the reconstruction frame x t-3 The difference between, etc. If it is determined that the difference exceeds the difference threshold, it can be determined that the context information has changed significantly, and then the feature information for the target frame is extracted from the reference reconstruction frame for the target frame If it is determined that the difference does not exceed the difference threshold, it can be determined that the context information has not changed significantly, and then the reference feature information F of the reference reconstruction frame is used t-1 Extract the feature information for the target frame.
[0101] After extracting the feature information for the target frame x t it is possible to determine the context information C t for the target frame x t based on this feature information. t This context information C t is used during the frame encoding process of the target frame x. In this way, by comparing the difference threshold, the timing for refreshing the feature information of the target frame is determined, improving the efficiency of refreshing the feature information for the target frame, alleviating the long-term temporal error accumulation problem, and improving the quality of video encoding and decoding.
[0102] In some implementations, the method of periodically refreshing the feature information for the target frame and the method of refreshing the feature information for the target frame by comparing the difference threshold can be used simultaneously. For example, if either of these two methods meets the condition, then the feature information for the target frame can be refreshed. In some implementations, these two methods can be alternately used to refresh the feature information for the target frame.
[0103] Exemplary implementation of model training
[0104] In a scenario where long prediction chains lead to temporal error accumulation, in addition to refreshing the feature information for the target frame x t it is also possible to use a longer video to train the video codec model 200 to improve the quality of video encoding and decoding. In some implementations, during the training process of the video codec model 200, the training data of the video codec model 200 includes at least one sample video segment with the number of frames exceeding a predetermined number threshold. Exemplarily, the predetermined number threshold can be set to a number such as 1000, 2000, 3000, etc. After a period of training, it is determined whether the performance of the video codec model 200 has improved, and the predetermined number threshold is adaptively updated accordingly. In this way, the video codec model 200 can identify patterns in longer videos, better explore temporal correlations, and thus improve the quality of video encoding and decoding when the video codec model 200 is actually used.
[0105] To improve the generality of the video codec model 200, it is necessary to train the video codec model 200 to support multiple color spaces simultaneously without additional fine-tuning. In some implementations, during the training process of the video codec model 200, the loss function can be determined based at least on the distortion loss, and the distortion loss can include the distortion loss of the first sample video in the first color space and the distortion loss of the second sample video in the second color space. The first sample video and the second sample video can be the same video or different videos. The first color space can be, for example, the YUV color space, and the second color space can be, for example, the RGB color space. An example of the loss function based on the distortion loss is as follows:
[0106] LossRD = R + λ·(k·D YUV +(1 - k)·D RGB )·D YUV (9)
[0107] where D YUV represents the distortion loss of the YUV color space (i.e., the distortion loss of the first sample video in the first color space), D RGB represents the distortion loss of the RGB color space (i.e., the distortion loss of the second sample video in the second color space), k represents a hyperparameter for weighting, and can be set to 0.8, for example. Of course, the hyperparameter k can also be set to any other value within the range of 0 to 1, and the scope of the implementation of the present disclosure does not limit this.
[0108] In this way, by covering the distortion loss of the YUV color space and the distortion loss of the RGB color space in the loss function simultaneously, the two color spaces can be supported in a single model, improving the generality of the video codec model 200.
[0109] Figure 6 shows a comparison of the quality range according to some implementations of the present disclosure with the quality ranges of other codecs. As Figure 6 shown, curve 610 shows the signal-to-noise ratio (PSNR) range of the video codec model according to some implementations of the present disclosure, curves 612 - 616 show the PSNR ranges of traditional video codec models, and curve 618 shows the PSNR range of a previous neural video codec model. PSNR is a metric for measuring image quality. It can be seen that compared with other video codec models, the video codec model according to some implementations of the present disclosure has a wider quality range, and the PSNR increases with the increase of bits per pixel (BPP), maintaining a monotonically increasing trend, which can meet practical requirements.
[0110] Example process
[0111] Figure 7 FIG. 700 shows a flow chart of a process for video processing according to some implementations of the present disclosure. The process 700 can be implemented at Figure 2 the video codec model 200, which can be applied, for example, by Figure 1 the electronic devices 110 or 120.
[0112] At block 710, the video codec model 200 obtains quantization parameters for a target frame in the video.
[0113] At block 720, the video codec model 200 determines a quantization scaling value for the target frame based on the quantization parameters and at least a first upper limit value and a first lower limit value for quantization scaling, the quantization scaling value being between the first upper limit value and the first lower limit value, and the first upper limit value and the first lower limit value being determined during the training process of the video codec model.
[0114] At block 730, the video codec model 200 performs quantization on the target frame based on the quantization scaling value during at least one of the frame encoding process or the motion vector encoding process of the target frame to obtain a quantized representation of the target frame.
[0115] In some implementations, the process 700 further includes: determining an inverse quantization scaling value for the target frame based on the quantization parameters and at least a second upper limit value and a second lower limit value for inverse quantization scaling, the inverse quantization scaling value being between the second upper limit value and the second lower limit value, and the second upper limit value and the second lower limit value being determined during the training process of the video codec model; and performing inverse quantization on the quantized coding representation corresponding to the quantized representation based on the inverse quantization scaling value during at least one of the frame encoding process of the target frame or the motion vector encoding process of the target frame.
[0116] In some implementations, determining the quantization scaling value for the target frame includes: performing interpolation between the first upper limit value and the first lower limit value based at least on the quantization parameters using an interpolation function, where the interpolation function has monotonicity between the first upper limit value and the first lower limit value.
[0117] In some implementations, the quantization parameters are selected from a quantization parameter range. In some implementations, performing interpolation between the first upper limit value and the first lower limit value based at least on the quantization parameters includes: performing interpolation between the first upper limit value and the first lower limit value based on the quantization parameters and the upper limit value of the quantization parameter range using an interpolation function, where the interpolation function is configured such that as the quantization parameters linearly increase within the quantization parameter range, the quantization scaling value exponentially increases from the first lower limit value to the first upper limit value.
[0118] In some implementations, the interpolation function is configured in the form of an exponential function or a logarithmic function, where the exponent of the exponential function is determined based on the quantization parameter and the quantization parameter range, and the base of the exponential function is determined based on a first upper limit value and a first lower limit value. In some implementations, determining the inverse quantization scaling value of a target frame includes: performing interpolation based on the quantization parameter between a second upper limit value and a second lower limit value using the same interpolation function used in the determination of the quantization scaling value to obtain the inverse quantization scaling value.
[0119] In some implementations, the loss function used in the training process of the video codec model is determined by weighting between a bitrate loss and a distortion loss using a weighting factor, and the weighting factor is determined based on the quantization parameter, a weighting upper limit value, and a weighting lower limit value.
[0120] In some implementations, the weighting factor is determined by: performing interpolation based on the quantization parameter between the weighting upper limit value and the weighting lower limit value using the same interpolation function used in the determination of the quantization scaling value to obtain the weighting factor.
[0121] In some implementations, process 700 further includes: using the video codec model, during the context extraction process of the target frame, if it is determined that the context information update period has expired, extracting feature information for the target frame from the reference reconstruction frame for the target frame; if it is determined that the context information update period has not expired, extracting feature information for the target frame from the reference feature information of the reference reconstruction frame; and determining the context information for the target frame based at least on the feature information for the target frame, and wherein the context information for the target frame is used during the frame encoding process of the target frame.
[0122] In some implementations, the context information update period is determined based on the content type of the video.
[0123] In some implementations, process 700 further includes: using the video codec model, during the context extraction process of the target frame, determining the difference between multiple reconstruction frames before the target frame, if it is determined that the difference exceeds a difference threshold, extracting feature information for the target frame from the reference reconstruction frame for the target frame; if it is determined that the difference does not exceed the difference threshold, extracting feature information for the target frame from the reference feature information of the reference reconstruction frame; and determining the context information for the target frame based at least on the feature information for the target frame, and wherein the context information for the target frame is used during the frame encoding process of the target frame.
[0124] In some implementations, during the training process of the video codec model, the training data of the video codec model includes at least one sample video segment with the number of frames exceeding a predetermined number threshold.
[0125] In some implementations, the loss function used during the training process of the video codec model is determined based at least on a distortion loss, where the distortion loss includes the distortion loss of a first sample video in a first color space and the distortion loss of a second sample video in a second color space.
[0126] Example device
[0127] Figure 8 FIG. shows a schematic block diagram of an electronic device capable of implementing multiple implementations of the present disclosure. It should be understood that Figure 8 The illustrated electronic device 800 is merely exemplary and should not impose any limitation on the functions and scope of the implementations described in the present disclosure. One or more electronic devices 800 may be used, for example, to implement Figure 2 the video codec model 200, Figure 1 the electronic devices 110 or 120.
[0128] As Figure 8 shown, the electronic device 800 includes an electronic device 800 in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to, one or more processors or processing devices 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860.
[0129] In some implementations, the electronic device 800 may be implemented as a computing device, a computing system, a server, a mainframe, or other devices with computing capabilities.
[0130] The processing device 810 may be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 800. The processing device 810 may include a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor, a controller, and / or a microcontroller, etc.
[0131] The electronic device 800 generally includes multiple computer storage media. Such media can be any available media accessible to the electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can include volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can include removable or non-removable media and can include computer-readable media such as a memory stick, flash drive, magnetic disk, or any other media that can be used to store information and / or data and can be accessed within the electronic device 800.
[0132] The electronic device 800 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 8 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces.
[0133] The communication unit 840 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the electronic device 800 can be implemented by a single computing cluster or multiple computer machines that can communicate via a communication connection. Thus, the electronic device 800 can operate in a networked environment using a logical connection to one or more other servers, personal computers (PCs), or another general network node.
[0134] The input device 850 can be one or more various input devices such as a mouse, keyboard, data import device, etc. The output device 860 can be one or more output devices such as a display, data export device, etc. The electronic device 800 can also communicate with one or more external devices (not shown) as needed via the communication unit 840, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 800, or communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (such as a network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0135] In some implementations, in addition to being integrated on a single device, some or all of the various components of the electronic device 800 may also be arranged in the form of a cloud computing architecture. In a cloud computing architecture, these components may be remotely arranged and may work together to implement the functions described in this disclosure. In some implementations, cloud computing provides computing, software, data access, and storage services, which do not require an end user to be aware of the physical location or configuration of the system or hardware providing these services. In various implementations, cloud computing uses appropriate protocols to provide services over a wide area network (such as the Internet). For example, a cloud computing provider provides applications over a wide area network, and they can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on a server at a remote location. The computing resources in a cloud computing environment may be consolidated at a remote data center location or they may be distributed. The cloud computing infrastructure may provide services through a shared data center, even though they appear as a single access point for users. Thus, the components and functions described herein may be provided from a service provider at a remote location using a cloud computing architecture. Alternatively, they may be provided from a conventional server, or they may be directly or otherwise installed on a client device.
[0136] The electronic device 800 may be used to implement resource management in multiple implementations of this disclosure. The memory 820 may include one or more modules having one or more program instructions, which may be accessed and run by the processing unit 810 to implement the functions of the various implementations described herein. For example, the memory 820 may include a video codec module 822 for performing video encoding and decoding using a video codec model. As Figure 8 shown, the electronic device 800 may obtain a video to be encoded or a bitstream to be decoded through the input device 850, and may provide the encoded bitstream or the decoded video through the output device 860. In some implementations, the electronic device 800 may also receive input from other devices (not shown) via the communication unit 840.
[0137] Some example implementation manners of this disclosure are listed below.
[0138] In one aspect, this disclosure provides a computer-implemented method. The method includes: obtaining a quantization parameter for a target frame in a video; determining a quantization scaling value for the target frame based on the quantization parameter and at least a first upper limit value and a first lower limit value for quantization scaling, the quantization scaling value being between the first upper limit value and the first lower limit value, and the first upper limit value and the first lower limit value being determined during the training process of the video codec model; and performing quantization on the target frame based on the quantization scaling value in at least one of a frame encoding process or a motion vector encoding process of the target frame using the video codec model to obtain a quantized representation of the target frame.
[0139] In some implementations, the method further includes: determining an inverse quantization scaling value of a target frame based on a quantization parameter and at least a second upper limit value and a second lower limit value for inverse quantization scaling, the inverse quantization scaling value being between the second upper limit value and the second lower limit value, and the second upper limit value and the second lower limit value being determined during the training process of the video codec model; and performing inverse quantization on the quantization coding representation corresponding to the quantization representation based on the inverse quantization scaling value in at least one of the frame coding process of the target frame or the motion vector coding process of the target frame.
[0140] In some implementations, determining a quantization scaling value of a target frame includes: performing interpolation between a first upper limit value and a first lower limit value based at least on the quantization parameter by using an interpolation function, to obtain the quantization scaling value, where the interpolation function has monotonicity between the first upper limit value and the first lower limit value.
[0141] In some implementations, the quantization parameter is selected from a quantization parameter range. In some implementations, performing interpolation between a first upper limit value and a first lower limit value based at least on the quantization parameter includes: performing interpolation between the first upper limit value and the first lower limit value based on the quantization parameter and the upper limit value of the quantization parameter range by using an interpolation function, where the interpolation function is configured such that as the quantization parameter linearly grows within the quantization parameter range, the quantization scaling value exponentially grows from the first lower limit value to the first upper limit value.
[0142] In some implementations, the interpolation function is configured in the form of an exponential function or a logarithmic function, where the exponent of the exponential function is determined based on the quantization parameter and the quantization parameter range, and the base of the exponential function is determined based on the first upper limit value and the first lower limit value.
[0143] In some implementations, determining an inverse quantization scaling value of a target frame includes: performing interpolation between a second upper limit value and a second lower limit value based on the quantization parameter by using the same interpolation function used in the determination of the quantization scaling value, to obtain the inverse quantization scaling value.
[0144] In some implementations, the loss function used during the training process of the video codec model is determined by weighting between a bitrate loss and a distortion loss by using a weighting factor, and the weighting factor is determined based on the quantization parameter, a weighting upper limit value, and a weighting lower limit value.
[0145] In some implementations, the weighting factor is determined by: performing interpolation between the weighting upper limit value and the weighting lower limit value based on the quantization parameter by using the same interpolation function used in the determination of the quantization scaling value, to obtain the weighting factor.
[0146] In some implementations, the method further includes: using a video codec model, during the context extraction process of a target frame, if it is determined that the context information update period has expired, extracting feature information for the target frame from a reference reconstruction frame for the target frame; if it is determined that the context information update period has not expired, extracting feature information for the target frame from the reference feature information of the reference reconstruction frame; and determining context information for the target frame based at least on the feature information for the target frame, and wherein the context information for the target frame is used during the frame encoding process of the target frame.
[0147] In some implementations, the context information update period is determined based on the content type of the video.
[0148] In some implementations, the method further includes: using a video codec model, during the context extraction process of a target frame, determining the difference between a plurality of reconstruction frames before the target frame, if it is determined that the difference exceeds a difference threshold, extracting feature information for the target frame from a reference reconstruction frame for the target frame; if it is determined that the difference does not exceed the difference threshold, extracting feature information for the target frame from the reference feature information of the reference reconstruction frame; and determining context information for the target frame based at least on the feature information for the target frame, and wherein the context information for the target frame is used during the frame encoding process of the target frame.
[0149] In some implementations, during the training process of the video codec model, the training data of the video codec model includes at least one sample video segment with the number of frames exceeding a predetermined number threshold.
[0150] In some implementations, the loss function used during the training process of the video codec model is determined based at least on a distortion loss, and the distortion loss includes the distortion loss of the first sample video in the first color space and the distortion loss of the second sample video in the second color space.
[0151] In another aspect, the present disclosure provides an electronic device. The electronic device includes: a processor; and a memory coupled to the processor and containing instructions stored thereon, the instructions causing the device to perform the following actions when executed by the processor, the actions including: obtaining quantization parameters for a target frame in a video; determining a quantization scaling value for the target frame based on the quantization parameters and at least a first upper limit value and a first lower limit value for quantization scaling, the quantization scaling value being between the first upper limit value and the first lower limit value, and the first upper limit value and the first lower limit value being determined during the training process of the video codec model; and using the video codec model to perform quantization on the target frame based on the quantization scaling value during at least one of the frame encoding process or the motion vector encoding process of the target frame to obtain a quantized representation of the target frame.
[0152] In some implementations, the operation further includes: determining an inverse quantization scaling value of a target frame based on a quantization parameter and at least a second upper limit value and a second lower limit value for inverse quantization scaling, the inverse quantization scaling value being between the second upper limit value and the second lower limit value, and the second upper limit value and the second lower limit value being determined during the training process of the video codec model; and performing inverse quantization on the quantization-coded representation corresponding to the quantization representation based on the inverse quantization scaling value during at least one of the frame coding process of the target frame or the motion vector coding process of the target frame.
[0153] In some implementations, determining a quantization scaling value of a target frame includes: performing interpolation between a first upper limit value and a first lower limit value based at least on the quantization parameter by using an interpolation function, to obtain the quantization scaling value, where the interpolation function has monotonicity between the first upper limit value and the first lower limit value.
[0154] In some implementations, the quantization parameter is selected from a quantization parameter range. In some implementations, performing interpolation between a first upper limit value and a first lower limit value based at least on the quantization parameter includes: performing interpolation between the first upper limit value and the first lower limit value based on the quantization parameter and the upper limit value of the quantization parameter range by using an interpolation function, where the interpolation function is configured such that as the quantization parameter linearly increases within the quantization parameter range, the quantization scaling value exponentially increases from the first lower limit value to the first upper limit value.
[0155] In some implementations, the interpolation function is configured in the form of an exponential function or a logarithmic function, where the exponent of the exponential function is determined based on the quantization parameter and the quantization parameter range, and the base of the exponential function is determined based on the first upper limit value and the first lower limit value.
[0156] In some implementations, determining an inverse quantization scaling value of a target frame includes: performing interpolation between a second upper limit value and a second lower limit value based on the quantization parameter by using the same interpolation function used in the determination of the quantization scaling value, to obtain the inverse quantization scaling value.
[0157] In some implementations, the loss function used during the training process of the video codec model is determined by weighting between a bitrate loss and a distortion loss by using a weighting factor, and the weighting factor is determined based on the quantization parameter, a weighting upper limit value, and a weighting lower limit value.
[0158] In some implementations, the weighting factor is determined by: performing interpolation between the weighting upper limit value and the weighting lower limit value based on the quantization parameter by using the same interpolation function used in the determination of the quantization scaling value, to obtain the weighting factor.
[0159] In some implementations, the operation further includes: using a video codec model, during the context extraction process of a target frame, if it is determined that the context information update period has expired, extracting feature information for the target frame from a reference reconstruction frame for the target frame; if it is determined that the context information update period has not expired, extracting feature information for the target frame from the reference feature information of the reference reconstruction frame; and determining context information for the target frame based at least on the feature information for the target frame, and wherein the context information for the target frame is used during the frame encoding process of the target frame.
[0160] In some implementations, the context information update period is determined based on the content type of the video.
[0161] In some implementations, the operation further includes: using a video codec model, during the context extraction process of a target frame, determining the difference between a plurality of reconstruction frames before the target frame, if it is determined that the difference exceeds a difference threshold, extracting feature information for the target frame from a reference reconstruction frame for the target frame; if it is determined that the difference does not exceed the difference threshold, extracting feature information for the target frame from the reference feature information of the reference reconstruction frame; and determining context information for the target frame based at least on the feature information for the target frame, and wherein the context information for the target frame is used during the frame encoding process of the target frame.
[0162] In some implementations, during the training process of the video codec model, the training data of the video codec model includes at least one sample video segment with the number of frames exceeding a predetermined number threshold.
[0163] In some implementations, the loss function used during the training process of the video codec model is determined based at least on a distortion loss, and the distortion loss includes the distortion loss of a first sample video in a first color space and the distortion loss of a second sample video in a second color space.
[0164] In another aspect, the present disclosure provides a computer program product, the computer program product being tangibly stored in a computer storage medium and including computer-executable instructions, the computer-executable instructions causing a device to perform the following operations when executed by the device, the operations including: obtaining quantization parameters for a target frame in a video; determining a quantization scaling value for the target frame based on the quantization parameters and at least a first upper limit value and a first lower limit value for quantization scaling, the quantization scaling value being between the first upper limit value and the first lower limit value, and the first upper limit value and the first lower limit value being determined during the training process of the video codec model; and using the video codec model to perform quantization on the target frame based on the quantization scaling value during at least one of the frame encoding process or the motion vector encoding process of the target frame to obtain a quantized representation of the target frame.
[0165] In another aspect, the present disclosure provides a computer-readable medium having computer-executable instructions stored thereon that, when executed by a device, cause the device to perform one or more example implementations of the methods of the above aspects.
[0166] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example and not limitation, example types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0167] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0168] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0169] In addition, although the operations are depicted in a particular order, this should be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details were included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that were described in the context of separate implementations may also be implemented combinatorially in a single implementation. Conversely, various features that were described in the context of a single implementation may also be implemented separately or in any suitable subcombination in multiple implementations.
[0170] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A computer-implemented method, comprising: Obtaining quantization parameters for a target frame in a video; Determining a quantization scaling value for the target frame based on the quantization parameters and at least a first upper limit value and a first lower limit value for quantization scaling, the quantization scaling value being between the first upper limit value and the first lower limit value, and the first upper limit value and the first lower limit value being determined during the training process of a video codec model; And Using the video codec model to perform quantization on the target frame based on the quantization scaling value during at least one of the frame encoding process or the motion vector encoding process of the target frame to obtain a quantized representation of the target frame.
2. The method according to claim 1, further comprising: Determining an inverse quantization scaling value for the target frame based on the quantization parameters and at least a second upper limit value and a second lower limit value for inverse quantization scaling, the inverse quantization scaling value being between the second upper limit value and the second lower limit value, and the second upper limit value and the second lower limit value being determined during the training process of the video codec model; And Performing inverse quantization on the quantized coding representation corresponding to the quantized representation based on the inverse quantization scaling value during at least one of the frame encoding process of the target frame or the motion vector encoding process of the target frame.
3. The method according to claim 1, wherein determining the quantization scaling value for the target frame comprises: Using an interpolation function to perform interpolation based at least on the quantization parameters between the first upper limit value and the first lower limit value to obtain the quantization scaling value, wherein the interpolation function has monotonicity between the first upper limit value and the first lower limit value.
4. The method according to claim 3, wherein the quantization parameters are selected from a quantization parameter range, and wherein performing interpolation based at least on the quantization parameters between the first upper limit value and the first lower limit value comprises: Using the interpolation function to perform interpolation between the first upper limit value and the first lower limit value based on the quantization parameters and the upper limit value of the quantization parameter range, wherein the interpolation function is configured such that as the quantization parameters linearly increase within the quantization parameter range, the quantization scaling value exponentially increases from the first lower limit value to the first upper limit value.
5. The method according to claim 2, wherein determining the inverse quantization scaling value for the target frame comprises: Using the same interpolation function used in the determination of the quantization scaling value to perform interpolation between the second upper limit value and the second lower limit value based on the quantization parameters to obtain the inverse quantization scaling value.
6. The method according to claim 1, wherein the loss function used during the training process of the video codec model is determined by weighting between a bitrate loss and a distortion loss using a weighting factor, the weighting factor being determined based on the quantization parameters, a weighting upper limit value, and a weighting lower limit value.
7. The method according to claim 6, wherein the weighting factor is determined by: Interpolate between the weighted upper limit value and the weighted lower limit value based on the quantization parameter by using the same interpolation function used in the determination of the quantization scaling value to obtain the weighting factor.
8. The method according to claim 1, further comprising: Using the video codec model, during the context extraction process of the target frame, if it is determined that the context information update period has expired, extract the feature information for the target frame from the reference reconstruction frame for the target frame; If it is determined that the context information update period has not expired, extract the feature information for the target frame from the reference feature information of the reference reconstruction frame; and determine the context information for the target frame based at least on the feature information for the target frame, and wherein the context information for the target frame is used during the frame encoding process of the target frame.
9. The method according to claim 8, wherein the context information update period is determined based on the content type of the video.
10. The method according to claim 1, further comprising: Using the video codec model, during the context extraction process of the target frame, determine the difference between a plurality of reconstruction frames before the target frame, If it is determined that the difference exceeds a difference threshold, extract the feature information for the target frame from the reference reconstruction frame for the target frame; If it is determined that the difference does not exceed the difference threshold, extract the feature information for the target frame from the reference feature information of the reference reconstruction frame; and determine the context information for the target frame based at least on the feature information for the target frame, and wherein the context information for the target frame is used during the frame encoding process of the target frame.
11. The method according to claim 1, wherein during the training process of the video codec model, the training data of the video codec model includes at least one sample video segment with the number of frames exceeding a predetermined number threshold.
12. The method according to claim 1, wherein the loss function used during the training process of the video codec model is determined based at least on a distortion loss, and the distortion loss includes the distortion loss of the first sample video in the first color space and the distortion loss of the second sample video in the second color space.
13. An electronic device, comprising: A processor; And A memory coupled to the processor and containing instructions stored thereon, the instructions when executed by the processor cause the device to perform the following actions, the actions including: Obtain the quantization parameter for the target frame in the video; Based on the quantization parameter and at least a first upper limit value and a first lower limit value for quantization scaling, determine the quantization scaling value of the target frame, the quantization scaling value being between the first upper limit value and the first lower limit value, and the first upper limit value and the first lower limit value being determined during the training process of the video codec model; and Using the video codec model, during at least one of the frame encoding process or the motion vector encoding process of the target frame, perform quantization on the target frame based on the quantization scaling value to obtain a quantized representation of the target frame.
14. The apparatus according to claim 13, wherein the operation further comprises: Determining an inverse quantization scaling value of the target frame based on the quantization parameter and at least a second upper limit value and a second lower limit value for inverse quantization scaling, the inverse quantization scaling value being between the second upper limit value and the second lower limit value, and the second upper limit value and the second lower limit value being determined during the training process of the video codec model; And During at least one of the frame encoding process of the target frame or the motion vector encoding process of the target frame, perform inverse quantization on the quantized coding representation corresponding to the quantized representation based on the inverse quantization scaling value.
15. The apparatus according to claim 13, wherein determining the quantization scaling value of the target frame comprises: Using an interpolation function, perform interpolation between the first upper limit value and the first lower limit value based at least on the quantization parameter to obtain the quantization scaling value, Wherein the interpolation function has monotonicity between the first upper limit value and the first lower limit value.
16. The apparatus according to claim 15, wherein the quantization parameter is selected from a quantization parameter range, and wherein performing interpolation between the first upper limit value and the first lower limit value based at least on the quantization parameter comprises: Using the interpolation function, perform interpolation between the first upper limit value and the first lower limit value based on the quantization parameter and the upper limit value of the quantization parameter range, Wherein the interpolation function is configured such that as the quantization parameter linearly increases within the quantization parameter range, the quantization scaling value exponentially increases from the first lower limit value to the first upper limit value.
17. The apparatus according to claim 14, wherein determining the inverse quantization scaling value of the target frame comprises: Using the same interpolation function used in the determination of the quantization scaling value, perform interpolation between the second upper limit value and the second lower limit value based on the quantization parameter to obtain the inverse quantization scaling value.
18. The apparatus according to claim 13, wherein the loss function used during the training process of the video codec model is determined by weighting between a bitrate loss and a distortion loss using a weighting factor, the weighting factor being determined based on the quantization parameter, a weighting upper limit value, and a weighting lower limit value.
19. The apparatus according to claim 18, wherein the weighting factor is determined by: Using the same interpolation function used in the determination of the quantization scaling value, perform interpolation between the weighting upper limit value and the weighting lower limit value based on the quantization parameter to obtain the weighting factor.
20. A computer program product, the computer program product being tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the following operations, the operations comprising: Obtain quantization parameters for a target frame in a video; Based on the quantization parameters and at least a first upper limit value and a first lower limit value for quantization scaling, determine a quantization scaling value for the target frame, the quantization scaling value being between the first upper limit value and the first lower limit value, and the first upper limit value and the first lower limit value being determined during the training process of a video codec model; And Using the video codec model, perform quantization on the target frame based on the quantization scaling value during at least one of the frame encoding process or the motion vector encoding process of the target frame to obtain a quantized representation of the target frame.