A method, a computing device, a non-transitory computer-readable storage medium, and a computer program for encoding a video signal.
By employing new low-pass interpolation filters and disabling RPR-based inter prediction, the inefficiencies in VVC's reference picture resampling are addressed, improving coding efficiency and reducing complexity and bandwidth usage.
Patent Information
- Application Number
- JP2024135770
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-24
- Filing Date
- 2024-08-15
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2040-12-24
AI Technical Summary
Existing video encoding standards like VVC face inefficiencies in reference picture resampling, particularly in affine modes, leading to increased computational complexity, memory bandwidth usage, and aliasing artifacts due to inadequate interpolation filters, especially when reference picture resolution exceeds current picture resolution.
Implementing new low-pass interpolation filters for affine modes and disabling RPR-based inter prediction to reduce computational complexity and memory bandwidth, while allowing dynamic bit depth variation for improved coding efficiency.
Enhances coding efficiency and reduces computational complexity and memory bandwidth by using tailored interpolation filters for affine modes, minimizing aliasing artifacts and enabling flexible bit depth adaptation.
Smart Images

Figure 0007752737000025 
Figure 0007752737000026 
Figure 0007752737000027
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application is based on provisional application No. 62 / 953,471 filed December 24, 2019. and claims priority to, and for all purposes, all of, that provisional application. The contents of which are incorporated herein by reference. [Technical Field]
[0002] This disclosure relates to video encoding and compression. More particularly, this disclosure relates to a method for video encoding. The present invention relates to a method and apparatus for reference picture resampling techniques. [Background technology]
[0003] Various video coding techniques can be used to compress the video data. The encoding is performed according to one or more video coding standards. For example, a video coding standard may be a general-purpose Video Coding (VVC), Joint Exploration Model (JEM), High Efficiency Video Coding (H.26 5 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Experts Group (MPEG) Video coding generally involves exploiting redundancy present in a video image or sequence. It uses a prediction method (e.g., inter-prediction, intra-prediction, etc.) that is suitable for video coding. The key goal of the JPEG2000 is to compress video data into a format that uses a relatively low bit rate while The main objective is to avoid or minimize degradation of video quality. Summary of the Invention
[0004] Examples of this disclosure provide methods and apparatus for resampling reference pictures.
[0005] According to a first aspect of the present disclosure, there is provided a method for decoding a video signal, the method comprising: , the decoder obtains a reference picture I associated with a video block in the video signal. The decoder may also extract video blocks from reference blocks in reference picture I. A reference sample I(i,j) may be obtained, where i and j are the reference samples of one video block. If the video block is coded in non-affine inter mode and If the resolution of reference picture I is greater than the resolution of the current picture, the decoder to generate luma inter-prediction samples and chroma inter-prediction samples for the video block. In order to obtain the first downsampling filter and the second downsampling filter, If the video block is coded in affine mode and the resolution of the reference picture is If the resolution of the picture is larger than the decoder also calculates the luminance intensity of each video block. Third downsampling to generate inter prediction samples and chroma inter prediction samples The decoder may obtain a third downsampling filter and a fourth downsampling filter. Based on the application of a fourth downsampling filter to the reference sample I(i,j), The video block may further include a first prediction sample and a second prediction sample.
[0006] According to a second aspect of the present disclosure, there is provided a computing device, the computing device comprising one or more a processor and a non-volatile memory storing instructions executable by said one or more processors. and a temporary computer-readable memory. arranged to obtain a reference picture I associated with a video block in a video signal; The one or more processors may also The reference sample I(i,j) of the video block may be obtained from i and j. may represent the coordinates of one sample within an image block. The processor further determines whether the video block is coded in non-affine inter mode and is based on a reference picture I. If the resolution of each image block is greater than the resolution of the current picture, the luminance interval of each image block is First downsampling to generate the inter-prediction samples and chroma inter-prediction samples The filter may be arranged to obtain a first downsampling filter and a second downsampling filter. The one or more processors also select a video block coded in affine mode and a reference video block coded in reference mode. If the picture resolution is greater than the current picture resolution, the brightness of each video block is A third downsampling is performed to generate the intensity inter-predicted samples and the chroma inter-predicted samples. The filter may be arranged to obtain a fourth downsampling filter and a fourth downsampling filter. The one or more processors may further include a third and fourth downsampling filter. The inter prediction sample of the video block is applied to the reference sample I(i,j). It may be arranged to obtain a pull.
[0007] According to a third aspect of the present disclosure, a non-transitory computer-readable storage medium having instructions stored thereon is provided. These instructions, when executed by one or more processors of the device, These instructions cause the device to obtain a reference picture I associated with a video block in a video signal. These instructions also cause the device to The reference sample I(i,j) of the video block can be obtained from the image data. These instructions may also represent the coordinates of a single sample within an image block. , the video block is coded in non-affine inter mode and the resolution of the reference picture I is If the resolution is larger than the current picture resolution, the luminance inter prediction samples of each video block are The first downsampling filter and the second downsampling filter are used to generate the saturation and chrominance inter-prediction samples. These instructions also allow the user to obtain the first and second downsampling filters. The device is configured to: If the resolution of the image is larger than the image resolution, the luma inter prediction samples and The third downsampling filter and the fourth downsampling filter are used to generate the inter-prediction samples. These instructions also cause the device to obtain a downsampling filter. , the third and fourth downsampling filters are applied to the reference sample I(i,j). Inter-prediction samples of the video block can be obtained based on the above.
[0008] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory. and are not intended to limit the present disclosure. [Brief explanation of the drawings]
[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with this disclosure. and together with the explanation serve to explain the principles of the present disclosure.
[0010] [Figure 1] FIG. 1 is a block diagram of an encoder according to an example of the present disclosure. [Figure 2] FIG. 2 is a block diagram of a decoder according to an example of the present disclosure. [Figure 3A] FIG. 3A is a diagram illustrating block division of a multi-type tree structure according to an example of the present disclosure. [Figure 3B] FIG. 3B is a diagram illustrating block division of a multi-type tree structure according to an example of the present disclosure. [Figure 3C] FIG. 3C is a diagram illustrating block division of a multi-type tree structure according to an example of the present disclosure. [Figure 3D] FIG. 3D is a diagram illustrating block division of a multi-type tree structure according to an example of the present disclosure. [Figure 3E] FIG. 3E is a diagram illustrating block division of a multi-type tree structure according to an example of the present disclosure. [Figure 4A] FIG. 4A is a diagrammatic representation of a four-parameter affine model according to one example of the present disclosure. [Figure 4B] FIG. 4B is a diagrammatic representation of a four-parameter affine model according to an example of the present disclosure. [Figure 5] FIG. 5 is a diagram of a six-parameter affine model according to an example of the present disclosure. [Figure 6] FIG. 6 is a diagram illustrating adaptive bit depth switching according to an example of the present disclosure. [Figure 7] FIG. 7 is a method for decoding a video signal according to an example of the present disclosure. [Figure 8] FIG. 8 is a method for decoding a video signal according to an example of the present disclosure. [Figure 9] FIG. 9 is a diagram illustrating a computing environment coupled to a user interface according to an example of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0011] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. Reference is made to the drawings in which like numbers in different drawings refer to the same figures unless otherwise indicated. or similar elements. The implementations described in the following description of exemplary embodiments are incorporated herein by reference in their entirety. It does not represent all realizations that are consistent with the and an apparatus consistent with the related aspects of the present disclosure as set forth in the appended claims. This is an example of a method.
[0012] The terminology used in this disclosure is for the purpose of describing particular embodiments only and It is not intended to limit the disclosure. In these cases, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are used interchangeably. Likewise, the plural is intended to be included. As used herein, the term "and / or" is used interchangeably with "and / or." A term refers to any or all possible combinations of one or more of the associated listed items. It should also be understood that the term "common" is intended to represent and encompasses the term "common."
[0013] The terms "first," "second," "third," etc. are used herein to describe various pieces of information. Although the information may be used, it should not be limited by these terms. It is further understood that these terms distinguish one category of information from another. For example, without departing from the scope of this disclosure, the first information may be used only for the purpose of Similarly, the second information may be referred to as the first information. When used in the specification, the term "when" should be interpreted as "when," "in the event of," or may be understood to mean "at one's discretion" depending on the context.
[0014] The first version of the HEVC standard was completed in October 2013 and replaces the previous generation video coding standard, H.26 4 / MPEG AVC, offering approximately 50% bitrate savings or equivalent perceived quality. Although the HEVC standard offers significant coding improvements over its predecessors, Evidence that additional coding tools can be utilized to achieve better coding efficiency than HEVC Based on this, both VCEG and MPEG are working to standardize future video coding. We have begun exploring new coding techniques to achieve significant improvements in coding efficiency. The Joint Visual Exploration Team (JVE) has been tasked with launching major research into advanced technologies that can achieve this. T) was established in October 2015 by ITU-T VECG and ISO / IEC MPEG. One reference software, called the Joint Exploration Model (JEM), is available at JVET. Therefore, by integrating some additional coding tools into the HEVC Test Model (HM), It is supported.
[0015] In October 2017, ITU-T and ISO / IEC announced that they have the capability to go beyond HEVC. The 10th J-SIGMA Joint Call for Proposals (CfP) for video compression was announced in April 2018. At the VET meeting, 23 CfP responses were received and evaluated, which included HEVC Based on these evaluation results, JVET has started a new project to develop a new generation of video coding system named Versatile Video Coding (VVC). In the same month, a VVC test model was developed to demonstrate the reference implementation of the VVC standard. A single reference software code base called VTM was established.
[0016] Like HEVC, VVC is a block-based hybrid video coding framework It is established on the basis of
[0017] Figure 1 shows an overview of a block-based video encoder for VVC. 1 shows a typical encoder 100. The encoder 100 receives a video input 110 and a motion compensation 111. 12, motion estimation 114, intra / inter mode decision 116, block predictor 140, adder 128, transform 130, quantization 132, prediction related information 142, intra prediction 118, Picture buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124 , in-loop filter 122, entropy coding unit 138 and bitstream 1 It has 44.
[0018] In encoder 100, a video frame is divided into multiple video blocks for processing. For each given video block, a prediction method is used to predict the video block based on the inter prediction method or the intra prediction method. Predictions are formed based on this.
[0019] A current video block that is part of the video input 110 and its predictor (block predictor 140) The prediction residual, which represents the difference between the input signal and the output signal (part of the input signal), is sent from the adder 128 to a transform 130. The transform coefficients are then sent from transform 130 to quantization 132 for entropy reduction. The quantized coefficients are then entropy-decoded to produce a compressed video bitstream. The video is then fed to the video encoding unit 138. As shown in FIG. information, motion vectors (MV), reference picture indexes and intra prediction modes Prediction related information 142 from the intra / inter mode decision 116 is also included in the entropy coded units. The compressed data is fed through the bit stream 138 and stored in the compressed bit stream 144. The condensed bitstream 144 includes a video bitstream.
[0020] In the encoder 100, decoder-related circuitry for reconstructing pixels for prediction First, the prediction residual is reconstructed by inverse quantization 134 and inverse transform 136. The reconstructed prediction residual is the unfiltered reconstruction for the current video block. It is combined with the block predictor 140 to generate the pixel.
[0021] Spatial prediction (or "intra prediction") is a prediction of a video block in the same video frame as the current video block. Using pixels from samples of neighboring blocks that have already been coded (called reference samples) The current video block is predicted.
[0022] Temporal prediction (also called "inter-prediction") is the reconstruction of previously coded video pictures. Pixels are used to predict the current video block. Temporal prediction is based on the inherent temporal redundancy in video signals. Temporal prediction for a given coding unit (CU) or coding block The signal is typically signaled by one or more MVs, which are currently connected to the CU and its It indicates the amount and direction of motion between temporal references. In addition, multiple reference pictures are supported. When the temporal prediction signal is received, it is determined which reference picture in the reference picture store the temporal prediction signal comes from. In this case, one reference picture index is additionally transmitted, which is used to identify the reference picture.
[0023] The motion estimation 114 takes signals from the video input 110 and the picture buffer 120, The motion estimation signal is output to the motion compensation 112. The motion compensation 112 receives the video input 110, the picture The signal from the buffer 120 and the motion estimation signal from the motion estimation 114 are taken and motion compensated. The signal is output to the intra / inter mode decision 116 .
[0024] After spatial and / or temporal prediction is performed, the intra / inter prediction in the encoder 100 The mode decision 116 selects the optimal prediction mode based on, for example, a rate-distortion optimization method. Next, the block predictor 140 is subtracted from the current image block, and the transform 130 and Quantization 132 is used to decorrelate the resulting prediction residual. The resulting quantized residual coefficients are is inversely quantized by inverse quantization 134 and inversely transformed by inverse transform 136 to form a reconstructed residual. The inverse transformed reconstructed residual is then added to the predicted block to form the reconstructed signal for the CU. Furthermore, the reconstructed CU is placed in the reference picture store of the picture buffer 120 and is used to store future pictures. Before being used to encode a block, a deblocking filter, a sample adaptation In-loop filters such as a self-adjusting offset (SAO) and / or adaptive in-loop filters (ALF) 22 may be applied to the reconstructed CU to form the output video bitstream 144. , coding mode (inter or intra), prediction mode information, motion information and quantized residual coefficients. All numbers are sent to the entropy coding unit 138 for further compression and packing, A bitstream is formed.
[0025] Figure 1 shows the block diagram of a typical block-based hybrid video encoding system. The input video signal is processed in blocks (called CUs). Therefore, the CU can reach 128x128 pixels. However, it is based only on the quadtree. Unlike HEVC, which divides blocks into blocks by one coding tree unit (C A TU is divided into several CUs based on a 4-ary / 2-ary / 3-ary tree to adapt to the changing local characteristics. Also, the concept of multiple fragmentation unit types in HEVC has been removed, i.e. In VVC, the separation of CU, prediction unit (PU) and transform unit (TU) no longer exists. Instead, each CU is always used as the basic unit for both prediction and transformation. In a multi-type tree structure, one CTU is first divided into a quad-tree structure. The leaf nodes of each quadtree may then be further divided into binary and ternary tree structures.
[0026] As shown in Figures 3A, 3B, 3C, 3D and 3E, there are five division types: That is, there are four-way divisions, horizontal two-way divisions, vertical two-way divisions, horizontal three-way divisions and vertical three-way divisions.
[0027] FIG. 3A shows a diagram illustrating a quaternary division of a block in a multi-type tree structure according to the present disclosure. will be done.
[0028] FIG. 3B is a diagram illustrating vertical binary division of blocks in a multi-type tree structure according to the present disclosure. is shown.
[0029] FIG. 3C illustrates a horizontal binary division of blocks in a multi-type tree structure according to the present disclosure. is shown.
[0030] FIG. 3D illustrates a vertical ternary division of a block in a multi-type tree structure according to the present disclosure. is shown.
[0031] FIG. 3E illustrates a horizontal ternary division of blocks in a multi-type tree structure according to the present disclosure. is shown.
[0032] In FIG. 1, spatial prediction and / or temporal prediction can be performed. “Sequential prediction” (Sequential prediction) is a method of sampling coded neighboring blocks in the same video picture / slice. The current image block is predicted using pixels from the spatial sample (called the reference sample). Prediction reduces the spatial redundancy inherent in video signals. Motion-compensated prediction (also called "motion-compensated prediction") uses reconstructed pixels from a previously coded video picture. The current video block is predicted using temporal prediction. Temporal prediction reduces the temporal redundancy inherent in video signals. The temporal prediction signal for a given CU is typically signaled by one or more MVs. The MV indicates the amount and direction of motion between the current CU and its time reference. Also, when multiple reference pictures are supported, the temporal prediction signal is used to store the reference picture. One reference picture used to identify which reference picture the image comes from. After spatial and / or temporal prediction, the encoder The mode decision block selects the optimal prediction mode based on, for example, a rate-distortion optimization method. The predicted block is then subtracted from the current image block and the prediction is obtained using transformation and quantization. The quantized residual coefficients are dequantized and inverse transformed to form a reconstructed residual. The reconstructed residual is then added to the predicted block to form the reconstructed signal for the CU. , the reconstructed CU is placed in the reference picture store and used to encode future video blocks. Before the loop is over, in-loop filtering such as deblocking filters, SAO and ALF may be applied to the reconstructed CU. The prediction mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all encoded. It is sent to the entropy coding unit for further compression and packing to produce a bitstream is formed.
[0033] FIG. 2 shows the overall block diagram of a video decoder for VVC. Specifically, FIG. 2 shows a representative 2 shows a block diagram of a suitable decoder 200. The decoder 200 receives a bitstream 210, an encoder 212, and a decoder 214. tropy decoding 212, inverse quantization 214, inverse transform 216, adder 218, Inter / intra mode selector 220, intra prediction 222, memory 230, and in-loop filter. 228, motion compensation 224, picture buffer 226, prediction related information 234 and video output It has 232.
[0034] Decoder 200 is similar to the reconstruction-related portion residing in encoder 100 in FIG. In the decoder 200, first, the incoming data is decoded by entropy decoding 212. The video bitstream 210 is decoded to obtain quantization coefficient levels and prediction related information. The quantized coefficient levels are then processed by inverse quantization 214 and inverse transform 216 to produce a reconstruction. The block implemented in the intra / inter mode selector 220 obtains the built-in prediction residual. The block predictor mechanism performs intra prediction 222 or motion prediction based on the decoded prediction information. The set of unfiltered reconstructed pixels is arranged to perform the compensation 224. The adder 218 is used to add the reconstructed prediction residual from the inverse transform 216 and the block predictor mechanism. The output is obtained by summing it with the predicted output generated by the algorithm.
[0035] The reconstructed blocks are further filtered by the in-loop filter 22 before being stored in the picture buffer 226. 8, and the picture buffer 226 acts as a reference picture store. The reconstructed image in the picture buffer 226 is used to drive a display device, The resulting image can be transmitted for use in predicting future video blocks. In situations where the intra-loop filter 228 is activated, filtering is performed on these reconstructed pixels. A filtering operation is performed to obtain the final reconstructed video output 232.
[0036] Figure 2 shows the overall block diagram of a block-based video decoder. The decoding unit entropy decodes the video bitstream. The coding mode and prediction information are then sent to a spatial prediction unit (intra prediction unit) to form a prediction block. the temporal prediction unit (if inter-coding is performed) or the temporal prediction unit (if inter-coding is performed). The residual transform coefficients are passed through an inverse quantization unit and an inverse transform unit to reconstruct the residual block. The predicted block and the residual block are then added together. The reconstructed block is The images may then go through further in-loop filtering before being stored in the reference picture store. The reconstructed image in the reference picture store is then used to drive the display device and, in the future, is transmitted for use in predicting a video block of the image.
[0037] The focus of this disclosure is to improve the existing reference picture resampling design supported by VVC. The present disclosure is based on the following closely related technology. This paper briefly reviews the current coding tools in VVC.
[0038] Affine Mode
[0039] In HEVC, only the translational motion model is applied to motion compensated prediction. There are many kinds of motion in the world, such as zoom, rotation, perspective motion, and other irregular motions. In VVC, affine motion compensation prediction is performed using either translational or affine motion models. For each inter-coded block, a suffix is used to indicate whether the rule is applied to inter prediction. This is applied by signaling one flag. In the current VVC design, one For affine coding blocks, 4-parameter affine mode and 6-parameter affine Two affine modes are supported, including the .
[0040] The four-parameter affine model has the following parameters: horizontal and vertical Two parameters used for translational movement in the direction, and scaling movement in these two directions. One parameter used for horizontal scaling and one parameter used for rotational movement. The horizontal rotation parameter is equal to the vertical rotation parameter. To achieve more efficient affine parameter signaling, VVC In this case, two MVs (control point motion vectors) are located at the upper left and upper right corners of the current block. These affine parameters are obtained by the CPMV (also called CPMV).
[0041] As shown in Figures 4A and 4B, the affine motion field of a block is calculated using two controls. It is described by the point MV(V0,V1).
[0042] Figure 4A shows a diagram of a four-parameter affine model. Figure 4B shows a diagram of a four-parameter affine model. The motion field of one affine coding block is calculated based on the control point motion. (v x ,v y ) is explained as follows:
number
[0043] The six-parameter affine mode has the following parameters: horizontal and vertical Two parameters are used for translational movement in the direction, and one for scaling movement in the horizontal direction. one parameter for rotational movement, one parameter for vertical expansion, One parameter used for large contraction movements and one parameter used for rotation movements. The six-parameter affine motion model is coded using three CPMVs.
[0044] Figure 5 shows a diagram of the six-parameter affine model. As shown in Figure 5, one six-parameter The three control points of a meter-affine block are located at the top left, top right, and bottom left corners of the block. Movement at the top left control point is related to translational movement, and movement at the top right control point is related to horizontal movement. The movement at the bottom left control point corresponds to the vertical rotation and scaling movement. It is related to large contraction motion. Compared to the 4-parameter affine motion model, the 6-parameter horizontal The rotation and scaling movements in the direction do not have to be the same as those in the vertical direction. (V0, Assume that V1, V2) are the MVs at the top left, top right, and bottom left corners of the current block in Figure 5. Then, the motion vector (v x ,v y )but It is obtained as follows:
number
[0045] In VVC, the CPMVs of affine coded blocks are stored in a separate buffer. The stored CPMV is in affine merge mode (i.e., affine merged from adjacent affine blocks). affine explicit modes (i.e., based on prediction-based schemes) and affine explicit modes (i.e., based on prediction-based schemes). It is used only to generate the affine CPMV predictor (signaling affine CPMV via The sub-block MV obtained from the CPMV is used for motion compensation, MV prediction of translational MV, and decode. Used for blocking.
[0046] Similar to regular inter-block motion compensation, the MV of each affine sub-block is a fractional It may refer to a reference sample at a sample location. In such cases, the reference to a fractional pixel location Interpolation filtering is required to generate the samples. To control the memory bandwidth requirements and worst-case computational complexity of the interpolation, a 6-tap interpolator is used. A set of inter-filters is used for motion compensation of affine sub-blocks. Tables 1 and 2 show the This is the interpolation filter used for normal inter-block and affine block motion compensation. As can be seen, the 6-tap interpolation filter for the affine mode is usually The two outermost filter coefficients on each side of the 8-tap filter for the inter-block are 6 For a single tap filter, add directly to the 8 tap filter coefficients. The filter coefficients P0 and P5 in Table 2 are obtained directly from the It is equal to the sum of the filter coefficients P0 and P1 and the sum of the filter coefficients P6 and P7 in Table 1. stomach.
[0047] [Table 1]
[0048] [Table 2]
[0049] Also, for motion compensation of chroma samples, the same 4 taps as for normal inter blocks are used. The interpolation filters (as described in Table 3) are used for the affine blocks. [Table 3]
[0050] Reference Picture Resampling
[0051] Unlike HEVC, the emerging VVC standard allows for rapid conversion of the same content within a single bitstream. Supports spatial resolution switching. Such capability is called reference picture resampling (RPR). ) or Adaptive Resolution Switching (ARC). In this case, a picture or intra random access point ( IRAP) picture (e.g., IDR picture or CRA picture, etc.) is inserted. Allowing resolution changes within a single coded video sequence without requiring Not only can compressed video data be adapted to dynamic communication channel conditions, To avoid a sudden increase in bandwidth consumption due to the relatively large size of IDR or CRA pictures. Specifically, the following typical user situations can benefit from the RPR feature: can be done.
[0052] Rate adaptation in video telephony and conferencing. To adapt the encoded video to changing network conditions, Worse still, if the available bandwidth becomes lower, the encoder will use smaller resolution pixels. We can accommodate this by encoding the picture resolution. The change only happens after the IRAP picture, which causes some problems. A reasonable quality IRAP picture is significantly larger than an inter-coded picture. , and decoding becomes correspondingly more complex, thus consuming time and resources. This is because the decoder requires a change in resolution for loading. This is problematic because it breaks the low latency buffer requirement and forces the audio to resynchronize. and the end-to-end delay of the stream may increase at least temporarily. This results in a poor user experience.
[0053] Changes in active speakers in multi-party video conferences In multi-party video conferences, it is common for active speakers to be seen in front of the rest of the conference participants. When the active speaker changes, the pitch used for each participant is The resolution of the chatter must also be adjusted. Such changes in active speakers occur frequently. In this case, demand with ARC characteristics becomes even more important.
[0054] Faster Start to Streaming For streaming applications, the most common scenario is when the application displays This means that up to a certain length of decoded pictures will be buffered before starting Starting a bitstream with a smaller resolution allows for faster start of display. Ensure that the application has enough pictures in the buffer to Accept.
[0055] Adaptive stream switching in streaming Dynamic Adaptive Streaming over HTTP (DASH) standard @mediaStr It contains a feature called eamStructureId, which is a non-decodable leading Switching between different views at open GOP random access points with pictures and the picture is a RASL picture to which it is associated in, for example, HEVC. Two different views of the same video are at different bit rates. but they have the same spatial resolution and they are the same @mediaStream If it has a StructureId value, the CR with the associated RASL picture It is possible to switch between these two views in the A picture and with acceptable quality. capable of decoding RASL pictures associated with transitions at CRA pictures; This allows for seamless switching. The streamStructureId feature also supports DASH displays with different spatial resolutions. It can be used to switch between
[0056] At the 15th JVET conference, the RPR feature will be officially supported by the VVC standard. The main aspects of existing RPR designs in VVC can be summarized as follows:
[0057] RPR Advanced Signaling
[0058] According to the current RPR design, in the sequence parameter set (SPS), Two syntax elements are used to specify the maximum width and height of the coded picture. pic_width_max_in_luma_samples and pic_heigh t_max_in_luma_samples is signaled. Then, When the resolution varies, the PPS is used to specify different picture resolutions for the pictures it references. The relevant syntax elements for this are pic_width_in_luma_samples and When pic_height_in_luma_samples and pic_height_in_luma_samples are signaled, One new Picture Parameter Set (PPS) must be configured. There is frame compatibility, i.e. pic_width_in_luma_samples and p The value of ic_height_in_luma_samples is x_in_luma_samples and pic_height_max_in_lum The value of a_samples should not be exceeded. Table 4 shows the RPR for SPS and PPS. The associated signaling has been described.
[0059] [Table 4]
[0060] Reference Picture Resampling
[0061] If the resolution changes within one bitstream, the current picture is one of the different dimensions. According to the current RPR design, the number of reference pictures is When the resolution changes, all MVs of the current picture are aligned with the sample grid of the reference picture. This is normalized to the sample grid of the current picture rather than to the current The transformation can be made transparent to the MV prediction process.
[0062] When the picture resolution changes, in addition to the MV, additional data is required during the motion compensation of the current block. The samples in one reference block must be unsampled / downsampled first. In VVC, the scaling ratio, i.e., refPicWidthInLuma Sample / picWidthInLuma and refPicHeightInLum aSample / picHeightInLumaSample is in the range [1 / 8,2] Limited.
[0063] In the current RPR design, the current picture and its reference picture are of different resolutions. In this case, a different interpolation filter is applied to interpolate the reference sample. If the resolution of the current picture is less than or equal to the default 8-tap and 4-tap interpolation The inter-prediction samples of luma and chroma are calculated using the inter-filter. However, the default motion interpolation filter does not exhibit a strong low-pass characteristic. If the picture resolution is higher than the current picture resolution, the default motion interpolation filter is used. Using a filter causes significant aliasing, which is Therefore, the inter-prediction effect of RPR is worse when the ring ratio increases. To improve the efficiency, if the reference picture has a higher resolution than the current picture, Applies a set of three different downsampling filters. If the switching ratio is 1.5:1 or greater, the 8-tap and 4-tap Use the Lanczos filter provided.
[0064] [Table 5]
[0065] [Table 6]
[0066] When the downsampling ratio is 2:1 or more, the cosine window function is used for the 12-tap SHM downsampling. The following 8-tap and 4-tap data are obtained by applying the Use the unsampling filters (shown in Tables 7 and 8). [Table 7]
[0067] [Table 8]
[0068] Finally, the above downsampling filters are used for non-affine inter-block luminance prediction. It is only applied to generate samples and chroma prediction samples. Still downsampling the default 8-tap and 4-tap motion interpolation filters Apply this to g.
[0069] Problems with existing RPR designs
[0070] The goal of this disclosure is to improve the coding efficiency of affine modes when applying RPR. Specifically, the following problems with existing RPR designs in VVC are identified:
[0071] First, as discussed previously, the resolution of the reference picture is higher than the resolution of the current picture. If high, an additional downsampling filter is applied only to non-affine mode motion compensation. In affine mode, 6-tap and 4-tap motion interpolation filters are applied. Assuming that these filters are derived from the default motion interpolation filters, does not exhibit a strong low-pass characteristic, and therefore, compared to non-affine modes, the famous Nyqui Due to the St-Shannon sampling theorem, the predicted samples of affine modes are exhibits worse aliasing artifacts. Thus, better coding performance In order to achieve this, an appropriate low-pass filter is also required if downsampling is required. It is desired to apply it to affine mode motion compensation.
[0072] Second, based on the existing RPR design, the current sample position, MV and reference picture Fractional pixels of the reference sample based on the resolution scaling ratio between the current picture and the current image Therefore, when applying downsampling of the reference block, It interpolates the reference samples of the current block at the expense of memory bandwidth consumption and computational complexity. Assume that the dimensions of a block are M (width) x N (height). The dimensions of the reference picture are the same as the current picture. If the dimensions of the reference picture are the same as those of the image, then the dimensions of the reference picture are (M+7) x (N+7). It is necessary to access an integer sample, and the current block is 8x (M×(N+7))+8×M×N multiplications are required. Downsampling scaling ratio If the rate is s, the corresponding memory bandwidth and multiplications are (s × M + 7) × (s × N + 7) and 8 × (M × (s × N + 7)) + 8 × M × N. In Tables 9 and 10, the RPR When the unsampling scaling ratio is 1.5X and 2X, various blocks are The number of integer samples used for block-dimensional motion compensation and the number of multiplications per sample are compared. In Tables 9 and 10, the column named "RPR 1X" indicates the solution of the reference picture and the current picture. The image quality is the same, i.e., it corresponds to the situation when RPR is not applied. "Ratio of" is the memory bandwidth / multiplication ratio when the RPR downsampling ratio is greater than 1 and the worst case scenario for normal inter-mode without RPR (i.e. As can be seen, the correspondence ratio with normal inter-prediction is plotted. The reference picture has a higher resolution than the current picture compared to the worst-case complexity of the measurement. The memory bandwidth and computational complexity increase significantly when 4 Bidirectional prediction is used, and memory bandwidth and number of multiplications are the worst case scenario. The memory bandwidth and number of multiplications are 231% and 127% of the predictions.
[0073] [Table 9]
[0074] [Table 10]
[0075] Third, in existing RPR designs, VVC is a single, unified stream within the same bitstream. It only supports adaptive switching of picture resolution but does not support the use of a The bit depth will remain the same, but the CfP for the VVC standard will be published. The "Requirements for future video coding standards" states that "This standard does not allow multiple representations of the same content (where each representation is different) to be displayed. adaptive streams that provide different attributes (e.g., spatial resolution or sample bit depth) It is clear that "quick display switching in the case of streaming services should be supported." In a real video application, a single instruction manual Due to the multiple data (SIMD) operations, the coded video The only way to allow bit depth to be changed is through the video encoder / decoder, especially the software. It provides a more flexible performance / complexity tradeoff for codec implementation.
[0076] Improvements to RPR coding
[0077] In this disclosure, we aim to improve the efficiency of RPR coding in VVC and its memory bandwidth. A solution to reduce the computational complexity has been proposed. More specifically, the technology proposed in this disclosure can be summarized as follows:
[0078] First, to improve the RPR coding efficiency in affine mode, a new low-pass interpolation filter is proposed. The filter is used when the reference picture has a higher resolution than the current picture, i.e., when the reference picture has a lower resolution than the current picture. If sampling is required, the existing 8-tap intensity interpolation filter and and a 4-tap chroma interpolation filter.
[0079] Second, in order to simplify RPR, a normal inter-mode For a given CU size, this results in a significant increase in memory bandwidth and computational complexity compared to Therefore, it is proposed to disable RPR-based inter prediction.
[0080] Third, dynamic variation of the internal bit depth for encoding a single video sequence. A method is proposed that allows
[0081] Downsampling filter for affine modes
[0082] As mentioned above, it is necessary to check whether the resolution of the current picture and its reference picture is the same. Regardless of whether the default 6-tap motion interpolation filter or the 4-tap motion interpolation filter is used, The filter is always applied in affine mode. Similar to the interpolation filters used in HEVC, The default motion interpolation filter in VVC does not exhibit a strong low-pass characteristic. If the downscaling ratio is close to 1, the default motion interpolation filter will However, the resolution downsampling from the reference picture to the current picture can be For larger sampling ratios, the Nyquist-Shannon sampling theorem Based on this, the same default motion interpolation filter can be used to eliminate aliasing artifacts. In particular, if the applied MV points to a reference sample at an integer sample position, the performance will be worse. In this case, the default motion interpolation does not apply any filtering operation at all. This can result in a significant degradation in the quality of the predicted samples for the block.
[0083] To mitigate aliasing artifacts due to downsampling, the present disclosure According to
[0046] , a different interpolation method with stronger low-pass characteristics for affine mode motion compensation is Use the filter to replace the existing default 6-tap / 4-tap interpolation filter In addition, the memory bandwidth and computational complexity are the same as those of conventional motion compensation processing. To maintain this, the proposed downsampling filter is It has the same length as the existing interpolation filter, i.e., 6 taps are used for the luminance component and 4 taps for the is used for the saturation component.
[0084] FIG. 7 shows a method for decoding a video signal. The method is applied to, for example, a decoder. That's fine.
[0085] In step 710, the decoder generates a reference image associated with a video block in the video signal. Picture I can be obtained.
[0086] In step 712, the decoder derives a video block from the reference block in reference picture I. A reference sample I(i,j) of the block can be obtained. i and j are, for example, one of the video blocks. The coordinates of the samples may be expressed as
[0087] In step 714, the video block is coded in non-affine inter mode and If the resolution of reference picture I is greater than the resolution of the current picture, the decoder to generate luma inter-prediction samples and chroma inter-prediction samples for the video block. To do this, we can obtain the first and second downsampling filters. .
[0088] In step 716, the video block is coded in affine mode and the reference picture is If the resolution of the decoder is greater than the resolution of the current picture, the decoder A third downlink is used to generate luma inter-predicted samples and chroma inter-predicted samples for the block. A fourth downsampling filter and a fourth downsampling filter can be obtained.
[0089] In step 718, the decoder determines whether the third and fourth downsampling filters are referenced. The inter-prediction sample of the video block based on the sample I(i,j) is applied can be obtained.
[0090] Affine luminance downsampling filter
[0091] There are several ways to derive the luminance downsampling filter for the affine mode.
[0092] Method 1 In one or more embodiments of the present disclosure, a normal inter mode (i.e., a non-affine mode) ) from the existing luminance downsampling filter to affine mode luminance downsampling It is proposed to directly obtain the 8-tap filter by this method. The two leftmost and rightmost filter coefficients of the filter are added together to obtain a 6-tap filter. By making it a single filter coefficient, the following can be obtained: 8 taps in Table 7 (used when scaling ratio is 2X) Luminance downsampling filter to new 6-tap luminance downsampling filter Tables 11 and 12 show the results when the spatial scaling ratios are 1.5:1 and 2:1, respectively. We have described a 6-tap luminance downsampling filter proposed for this case.
[0093] [Table 11]
[0094] [Table 12]
[0095] Method 2 In one or more embodiments of the present disclosure, a SHM obtained based on a cosine windowed sinc function is It is proposed to directly derive a 6-tap affine downsampling filter from the filter. Specifically, in this method, an affine downsampling filter is calculated based on the following formula: Get Ruta.
number
number
[0096] f c is the cutoff frequency, s is the scaling ratio, and w(n) is the cosine window function. Its definition is as shown in the following equation (5).
number
[0097] As an example, assume f is 0.9 and L=6. Tables 13 and 14 show the spatial scaling. The results obtained when the minor ratio was 1.5X (i.e., s=1.5) and 2X (i.e., s=2) Six-tap luminance downsampling was explained. [Table 13]
[0098] [Table 14]
[0099] It should be noted that in Tables 13 and 14, the filter coefficients are 7-bit signed variables. The accuracy is the same as the downsampling filter used in the RPR design. It is maintained so that
[0100] FIG. 8 shows a method for decoding a video signal. The method is applied to, for example, a decoder. That's fine.
[0101] In step 810, the decoder performs a calculation based on the cutoff frequency and the scaling ratio. You can obtain the frequency response of an ideal low-pass filter.
[0102] In step 812, the decoder can obtain a cosine window function based on the filter length. .
[0103] In step 814, the decoder generates a third downconverter based on the frequency response and the cosine window function. You can get a sampling filter.
[0104] Affine chroma downsampling filter
[0105] Below is a summary of how the chroma reference block works when the resolution of the reference picture is higher than the resolution of the current picture. Three methods are proposed to downsample the blocks.
[0106] Method 1 In the first method, existing RPRs designed for use with non-affine modes are used. The 1.5X (Table 6) and 2X (Table 8) 4-tap chroma downsampling filters have been revised. It is proposed to use the affine mode reference samples for downsampling.
[0107] Method 2 In the second method, the default 4-tap chroma interpolation filter (Table 3) is re-adjusted to the affine It is proposed to use the reference sample of the eigenmode for downsampling.
[0108] Method 3 In the third method, the saturation is calculated based on the cosine windowed sinc function shown in (3) to (5). It is proposed to obtain a downsampling filter. Tables 15 and 16 show the cosine window sin Assuming that the cutoff frequency of the c function is 0.9, the values are 1.5X and 2X, respectively. Illustrates the resulting 4-tap chroma downsampling filter for different scaling ratios.
[0109] [Table 15]
[0110] [Table 16]
[0111] Constraint block dimensions for RPR mode
[0112] As analyzed in the "Problem Statement" section, when downsampling occurs, the existing RP R is a significant increase in complexity (e.g., the number of integer samples accessed for motion compensation and Specifically, downsampling the reference block results in a large number of multiplications. The memory bandwidth and number of multiplications required for bidirectional prediction are 231% and and 127%.
[0113] In one or more embodiments, the resolution of the reference picture is higher than the resolution of the current picture. In this case, the inter-prediction for certain block shapes (e.g., 4xN, Nx4, and / or 8x8) Disable bidirectional prediction during the measurement period (but still allow unidirectional prediction) Tables 17 and 18 show that the RPR is used to perform inter prediction. The corresponding sample when bidirectional prediction is disabled for N, Nx4 and 8x8 block sizes. The memory bandwidth and number of multiplications per pull are shown. As can be seen, using the proposed constraints, The memory bandwidth and number of multiplications are 1 for bidirectional prediction in the worst case scenario of 1.5x downsampling. 30% and 107% and 116% of the worst case bidirectional prediction with 2X downsampling and reduced by up to 113%.
[0114] [Table 17]
[0115] [Table 18]
[0116] In the above example, the RPR mode dual is used only for 4xN, Nx4 and 8x8 block sizes. Although directional prediction is prohibited, for engineers who understand the most advanced modern imaging technology, The proposed constraints still apply to other block dimensions and inter-coding modes (e.g., unidirectional This applies to all video streams (forward / backward prediction, merged / non-merged modes, etc.).
[0117] Adaptive Bit Depth Switching
[0118] In the existing RPR design, VVC supports multiple picture resolutions within the same bitstream. supports adaptive switching of only the bit depth for encoding video sequences However, as previously analyzed, one and the same bitstream The actual encoder / decoder device allows switching of the coding bit depth within the This provides greater flexibility and allows for different compromises between coding performance and computational complexity. can be provided.
[0119] In this section, one IRAP picture, such as an immediate decode refresh (IDR) picture, is considered. Adapted to allow for varying the intra-coding bit depth without the requirement to introduce additional An adaptive bit depth switching (ABS) method is proposed.
[0120] FIG. 6 illustrates a hypothetical example in which a current picture 620 and its reference pictures are The blocks 610 and 630 are coded with different internal bit depths. Reference picture 610 Ref0 using 10-bit coding, current picture using 10-bit coding 620 and a reference picture 630 Ref1 that uses 12-bit coding. In order to support the proposed ABS capabilities, the VVC framework currently For this purpose, advanced syntax signaling and modifications to the motion compensation process are proposed.
[0121] Advanced ABS signaling
[0122] For the proposed ABS signaling, in the SPS, coding that refers to the SPS The existing bit depth syntax element specifies the maximum intra-encoded bit depth for a picture. One new syntax element to replace bit_depth_minus8 sps_max_bit_depth_minus8 is proposed. Next, the coding bits When changing depth, specify different coding bit depths for pictures that refer to the PPS. One new PPS syntax, pps_bit_depth_minus8, is sent for It is believed.
[0123] Bitstream conformance exists, i.e., pps_bit_depth_minus8 The value should not exceed the value of sps_max_bit_depth_minus8. 19 described the proposed ABS signaling in SPS and PPS.
[0124] [Table 19]
[0125] Predicted Sample Bit Depth Adjustment
[0126] If a single coding bit depth changes within a single coded video sequence, The current picture can be predicted from other reference pictures, and the reconstructed picture can be The construction samples are represented with different bit depth precision. When this situation occurs, The prediction samples generated from the motion compensation of the reference picture are coded at the bit depth of the current picture. It should be adjusted.
[0127] Interaction with other coding tools
[0128] Assume that ABS is applied to the reference picture and the current picture can be represented with different precision. Then, in VVC, some coding parameters are obtained using reference samples. Existing encoding tools may not work properly. For example, in the current VVC , Bidirectional Optical Flow (BDOF) and Decoder-Side Motion Vector Refinement (DMVR) Two decoder-side techniques are used to improve inter-coding efficiency using temporally predicted samples. Specifically, the BDOF tool uses L0 and L1 to improve the predicted sample quality. Calculates the improvement for each sample using one predicted sample, but DMVR uses L0 and L1 prediction The accuracy of motion vectors is improved at the sub-block level depending on the samples. Based on this, any one of the two prediction signals has a bit depth different from the bit depth of the current picture. When depth coding is performed, BDOF and DMVR processing are always performed for one inter block. It is suggested to avoid this.
[0129] 9 illustrates a computing environment 910 coupled to a user interface 960. 910 may be part of a data processing server. The computing environment 910 includes a processor 920, It includes a memory 940 and an I / O interface 950 .
[0130] The processor 920 controls the overall operation of the computing environment 910, e.g., display, data collection, and data communication. The processor 920 typically controls operations associated with image processing and image processing. One or more processors for performing all or some of the steps in the above method. Additionally, the processor 920 may include a processor for communicating between the processor 920 and other components. The processor may include one or more modules that facilitate the interaction of the It may be a central processing unit (CPU), microprocessor, single chip machine, GPU, etc. stomach.
[0131] Memory 940 stores various types of data to support the operation of computing environment 910. The memory 940 may include predetermined software 942. Examples of such data include any application or method executed in the computing environment 910. The memory 940 may contain instructions for the method, video data sets, image data, etc. This is achieved by using volatile or non-volatile memory devices or a combination thereof. The memory device may be, for example, a static random access memory (SRAM). M), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic data It is a disk or optical disk.
[0132] The I / O interface 950 connects the processor 920 to the keyboard and click wheel. It provides an interface between peripheral interface modules such as buttons. The buttons may include, but are not limited to, a home button, a start scan button, and an end scan button. The I / O interface 950 may be coupled to an encoder and a decoder.
[0133] In an embodiment, a non-transitory computer-readable storage medium containing a plurality of programs is also provided. The plurality of programs are stored in, for example, memory 940, and a computing environment for performing the above method is provided. 910 by the processor 920. For example, the non-transitory computer Data-readable storage media include ROM, RAM, CD-ROM, magnetic tape, floppy disk (registered The device may be a digital recording medium, a digital video recorder, a digital still camera ...
[0134] The non-transitory computer-readable storage medium includes a computational device having one or more processors. a plurality of programs for execution by the device are stored, and the plurality of programs are or when executed by a plurality of processors, the computing device Make the method work.
[0135] In an embodiment, the computing environment 910 may include one or more application-specific integrated circuits for performing the methods described above. Circuits (ASIC), Digital Signal Processors (DSP), Digital Signal Processing Devices (D SPD), programmable logic device (PLD), field programmable gate FPGA, Graphics Processing Unit (GPU), Controller, Microphone may be implemented using a controller, microprocessor, or other electronic components. .
[0136] While the description of the present disclosure has been provided for illustrative purposes, it is not intended that the disclosure be exhaustive or that the present disclosure Many modifications, variations and alternative implementations are possible in the above description. These and other aspects of the present invention will be apparent to one skilled in the art having the benefit of the guidance provided in and the associated drawings.
[0137] The examples have been chosen and described to illustrate the principles of the present disclosure and will be readily apparent to those skilled in the art to various different implementations. By understanding this disclosure in a systematic way, different modifications for specific applications that fit the underlying principles and assumptions are possible. Therefore, it is possible to optimally utilize different realization methods with different characteristics. As such, the scope of the disclosure should not be limited to the specific implementations disclosed, and Modifications and other implementations are expected to fall within the scope of this disclosure. [Explanation of symbols]
[0138] 100...Encoder, 110...Video input, 112...Motion compensation, 114...Motion estimation, 11 6... Intra / Inter mode decision, 118... Intra prediction, 120... Picture buffer , 122...in-loop filter, 124...memory, 126...adder, 128...adder, 13 0...Transform, 132...Quantization, 134...Inverse quantization, 136...Inverse transform, 138...Entropy Encoding unit 140...block predictor 142...prediction related information 144...bitstream Stream, 200...decoder, 210...bitstream, 212...entropy decoder 214...inverse quantization, 216...inverse transform, 218...adder, 220...intra / intra 222...intra prediction; 224...motion compensation; 226...picture buffer 228...in-loop filter, 230...memory, 232...video output, 234 prediction-related information Information, 910... computing environment, 920... processor, 940... memory, 950 I / O interface Ace
Claims
1. 1. A method for encoding a video signal, comprising: determining a reference picture associated with a video block in the video signal; obtaining a reference sample for the video block from the reference picture; determining a luma interpolation filter for the video block to be encoded in an affine motion mode based on a scaling ratio obtained from resolutions of the reference picture and the current picture, wherein one of the filter coefficients of the luma interpolation filter is equal to a sum of first two filter coefficients or last two filter coefficients of a first luma interpolation filter, and the first luma interpolation filter is used for video blocks to be encoded in a non-affine motion mode when a resolution scaling ratio is equal to or greater than a first value; applying the luma interpolation filter to the reference samples to obtain luma inter-predicted samples of the video block; and a method for encoding said video signal, said method comprising:
2. determining the luma interpolation filter for the video block encoded in the affine motion mode based on the scaling ratio, determining a second luma interpolation filter as the luma interpolation filter for the video block encoded in the affine motion mode in response to the scaling ratio being greater than or equal to the first value; the second luma interpolation filter is different from a third luma interpolation filter for the video block encoded in the affine motion mode when the scaling ratio is not greater than or equal to the first value. The method of claim 1.
3. determining a chroma interpolation filter for the video block encoded in the affine motion mode based on a comparison between resolutions of the reference picture and the current picture; and applying the chroma interpolation filter to the reference samples to obtain chroma inter-predicted samples of the video block; and The method of claim 1 further comprising:
4. Determining the luminance interpolation filter comprises:
2. The method of claim 1, comprising obtaining the intensity interpolation filter by applying a cosine window function to the frequency response of an ideal low-pass filter.
5. Determining the luminance interpolation filter comprises: obtaining the frequency response of the ideal low pass filter based on a cutoff frequency and the scaling ratio; obtaining the cosine window function based on a filter length; and deriving a downsampling filter based on the frequency response and the cosine window function.
6. 6. The method of claim 5, wherein the cutoff frequency is equal to 0.9, the filter length is equal to 6, and the scaling ratio is equal to 1.
5.
7. 6. The method of claim 5, wherein the cutoff frequency is equal to 0.9, the filter length is equal to 6, and the scaling ratio is equal to 2.
8. A method for encoding a video signal, comprising: determining a reference picture associated with the video block; obtaining a reference sample for the video block from the reference picture; determining a luma interpolation filter for the video block to be encoded in an affine motion mode based on a scaling ratio obtained from resolutions of the reference picture and the current picture, wherein one of the filter coefficients of the luma interpolation filter is equal to a sum of first two filter coefficients or last two filter coefficients of a first luma interpolation filter, and the first luma interpolation filter is used for video blocks to be encoded in a non-affine motion mode when a resolution scaling ratio is equal to or greater than a first value; applying the luma interpolation filter to the reference samples to obtain luma inter-predicted samples of the video block; and a method for encoding said video signal, said method comprising:
9. The method of claim 8, wherein determining the luma interpolation filter for the video block encoded in the affine motion mode based on the scaling ratio comprises: determining a second luma interpolation filter as the luma interpolation filter for the video block encoded in the affine motion mode in response to the scaling ratio being greater than or equal to the first value; the second luma interpolation filter is different from a third luma interpolation filter for the video block encoded in the affine motion mode when the scaling ratio is not greater than or equal to the first value. The method of claim 8. determining a chroma interpolation filter for the video block encoded in the affine motion mode based on a comparison between resolutions of the reference picture and the current picture; applying the chroma interpolation filter to the reference samples to obtain chroma inter-predicted samples of the video block; and The method of claim 8 further comprising:
11. The step of determining the luminance interpolation filter comprises:
9. The method of claim 8, comprising obtaining the intensity interpolation filter by applying a cosine window function to the frequency response of an ideal low-pass filter.
12. The method of claim 11, wherein determining the luminance interpolation filter comprises: obtaining the frequency response of the ideal low pass filter based on a cutoff frequency and the scaling ratio; obtaining the cosine window function based on a filter length; and deriving a downsampling filter based on the frequency response and the cosine window function.
13. The method of claim 12, wherein the cutoff frequency is equal to 0.9, the filter length is equal to 6, and the scaling ratio is equal to 1.
5.
14. The method of claim 12, wherein the cutoff frequency is equal to 0.9, the filter length is equal to 6, and the scaling ratio is equal to 2.
15. A computing device comprising: one or more processors; a non-transitory computer-readable storage medium storing instructions executable by said one or more processors, said one or more processors performing the steps of the method according to any one of claims 1 to 14. The computing device.
16. A non-transitory computer-readable storage medium on which a plurality of programs for execution by a computing device having one or more processors are stored, comprising: The plurality of programs, when executed by the one or more processors, cause the computing device to perform the method according to any one of claims 1 to 14. The non-transitory computer-readable storage medium.
17. A computer program stored on a computer-readable storage medium, comprising: comprising instructions which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 14, The computer program.
18. Executing the video signal encoding method according to any one of claims 8 to 14 to generate a bitstream; storing the bitstream; A method for storing bitstreams including: