Intercoding for Adaptive Resolution Video Coding
Affine motion prediction and DMVR in video coding formats address high bandwidth costs by enabling adaptive resolution changes, reducing transmission costs through efficient frame upsampling and downsampling.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- クラウド インテリジェンス アセッツ ホールディング (シンガポール)プライベート リミテッド
- Filing Date
- 2019-03-11
- Publication Date
- 2026-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Conventional video coding formats like H.264/AVC and H.265/HEVC require generating new video sequences for resolution changes, leading to high bandwidth costs, and new codecs like AV1's switch_frame and VVC/H.266's motion prediction tools do not effectively reduce bandwidth consumption during adaptive resolution changes.
Implementing motion prediction coding formats that include affine motion prediction and decoder-side motion vector refinement (DMVR) to enable adaptive resolution changes by upsampling and downsampling reconstructed frames, using motion information from multiple reference frames.
Enables efficient adaptive resolution changes with reduced bandwidth consumption by leveraging affine motion prediction and DMVR, allowing for seamless transitions between frames with varying resolutions.
Smart Images

Figure 0007893586000017 
Figure 0007893586000018 
Figure 0007893586000019
Abstract
Description
Technical Field
[0001] This specification relates to inter-coding for adaptive resolution video coding.
Background Art
[0002] In conventional video coding formats, such as the H.264 / AVC (Advanced Video Coding) and H.265 / HEVC (High Efficiency Video Coding) standards, video frames within a sequence have their sizes and resolutions recorded at the sequence level in the header. Therefore, in order to change the frame resolution, a new video sequence has to be generated from intra-coded frames, which involves a considerably larger bandwidth cost for transmission than inter-coded frames. Thus, when the network bandwidth is low, decreasing, or being adjusted, it is desirable to adaptively transmit downsampled low-resolution videos over the network. However, since the bandwidth cost for adaptive downsampling offsets the bandwidth benefit, it is difficult to achieve bandwidth savings while using conventional video coding formats.
[0003] Research has been conducted to support resolution changes while transmitting inter-coded frames. In the implementation of the AV1 codec developed by AOM, a new frame type called switch_frame is provided, which can be transmitted at a different resolution from the previous frame. However, the motion vector coding of switch_frame is restricted because it cannot refer to the motion vectors of the previous frame. Such a reference has conventionally provided another way to reduce the bandwidth cost, so the use of switch_frame still maintains a larger bandwidth consumption that offsets the bandwidth benefit.
[0004] Furthermore, existing motion coding tools perform motion compensation prediction (MCP) based only on the translational motion model.
[0005] The development of the next-generation video codec specification VVC / H.266 provides several new motion prediction coding tools, further supporting motion vector coding that references previous frames, as well as MCP based on irregular types of motion other than translational motion. New techniques are required to implement resolution changes in the bitstream with respect to these new coding tools. [Overview of the project] [Means for solving the problem]
[0006] The means for solving the problems of this disclosure include at least the configuration described in the claims.
[0007] Detailed explanations are provided with reference to the attached drawings. In the drawings, the number(s) at the left of the reference number identify the drawing in which the reference number first appears. The use of the same reference number in different drawings indicates similar or identical articles or features. [Brief explanation of the drawing]
[0008] [Figure 1A] This figure shows the configurations of multiple CMPVs for 4-parameter affine motion models and 6-parameter affine motion models, respectively. [Figure 1B] This figure shows the configurations of multiple CMPVs for 4-parameter affine motion models and 6-parameter affine motion models, respectively. [Figure 2] This figure shows the diagram for deriving motion information for the Luma component of a block. [Figure 3] This figure illustrates the selection of motion candidates for a frame relative to the CU using affine motion prediction coding. [Figure 4] This figure shows an example of deriving a candidate for inheritance affine merge. [Figure 5] This figure shows an example of deriving a candidate for a constructed affine merge. [Figure 6]This figure shows a diagram of the DMVR dual prediction process based on template matching. [Figure 7] This diagram shows an illustrative block diagram of the video encoding process. [Figure 8A] This diagram shows an illustrative flowchart of a video encoding method that implements resolution-adaptive video encoding. [Figure 8B] This diagram shows an illustrative flowchart of a video encoding method that implements resolution-adaptive video encoding. [Figure 8C] This diagram shows an illustrative flowchart of a video encoding method that implements resolution-adaptive video encoding. [Figure 9A] This diagram shows a flowchart illustrating a further example of a video encoding method that implements resolution-adaptive video encoding. [Figure 9B] This diagram shows a flowchart illustrating a further example of a video encoding method that implements resolution-adaptive video encoding. [Figure 9C] This diagram shows a flowchart illustrating a further example of a video encoding method that implements resolution-adaptive video encoding. [Figure 10] This figure shows an exemplary system for implementing a process and method for implementing resolution-adaptive video coding in a motion-predictive coding format. [Figure 11] This figure shows an exemplary system for implementing a process and method for implementing resolution-adaptive video coding in a motion-predictive coding format. [Modes for carrying out the invention]
[0009] The systems and methods described herein are intended to enable adaptive resolution in video encoding, and more specifically, to implement upsampling and downsampling of reconstructed frames to enable adaptive resolution changes between frames based on motion prediction coding tools provided by the VVC / H.266 standard.
[0010] According to exemplary embodiments of this disclosure, a motion prediction coding format may refer to a data format that encodes motion information and PUs of a frame by including motion information and one or more references to prediction units (PUs) of one or more other frames. Motion information may refer to data describing the motion of a block structure of a frame or its unit or subunit, such as motion vectors and references to blocks of the current frame or another frame. A PU may refer to a unit or subunit corresponding to a block structure in a block structure of a frame, such as an coding unit (CU), where the block is divided based on frame data and encoded according to an established video codec. Motion information corresponding to a prediction unit may describe motion prediction encoded by any motion vector coding tool, including but not limited to those described herein.
[0011] According to exemplary embodiments of this disclosure, the motion prediction coding format may include affine motion prediction coding and decoder-side motion vector refinement (DMVR). Features of these motion prediction coding formats relating to exemplary embodiments of this disclosure should be described herein.
[0012] A decoder using affine motion predictive coding may obtain the current frame of a bitstream encoded using an encoding format employing an affine motion model and derive a reconstructed frame ("affine motion predictive coding reconstructed frame"). The current frame may be inter-coded.
[0013] The motion information of the CU in the affine motion prediction coding reconstruction frame may be predicted by affine motion compensation prediction. The motion information may include multiple motion vectors, including multiple control point motion vectors (CPMVs) and derived motion vectors. As shown in Figures 1A and 1B, the multiple CPMVs represent the two motions of the CU, which act as two control points.
number
[0014] The motion vector at the sample position (x, y) may be derived from two control points by the following operation: [Number] The motion vector at the sample position (x, y) may be derived from three control points by the following operation: [Number] The motion information may be further predicted by deriving the motion information of the luma component of the block and also deriving the motion information of the chroma component of the block by applying a block-based affine transformation to the motion information of the block.
[0015] As shown in Figure 2, the luma component of a block may be divided into 4x4 pixel luma subblocks, and for each luma subblock, the luma motion vector at the sample position of the center of the luma subblock may be derived from the control points of the entire CU according to the calculation described above. The derived luma motion vectors of the luma subblocks may be rounded to an accuracy of 1 / 16.
[0016] The chroma component of a block may be divided into 4x4 pixel chroma subblocks, each chroma subblock having four adjacent luma subblocks. For example, adjacent luma subblocks may be below, to the left, to the right, or above the chroma subblock. For each chroma subblock, the motion vector may be derived from the average of the luma motion vectors of the adjacent luma subblocks.
[0017] A motion compensation interpolation filter may be applied to the derived motion vectors of each subblock to generate motion predictions for each subblock.
[0018] The motion information of the CU in the affine motion prediction coding reconstruction frame may include a motion candidate list. The motion candidate list may be a data structure containing references to multiple motion candidates. A motion candidate may be a block structure or its subunit, such as a pixel of the block structure of the current frame or any other suitable subdivision, or it may be a reference to a motion candidate in another frame. A motion candidate may be a spatial motion candidate or a temporal motion candidate. By applying motion vector compensation (MVC), the decoder may select a motion candidate from the motion candidate list and derive the motion vector of the motion candidate as the motion vector of the CU in the reconstruction frame.
[0019] Figure 3 shows an exemplary selection of motion candidates for a frame CU by affine motion predictive coding according to an exemplary embodiment of the present disclosure.
[0020] According to an exemplary embodiment of the present disclosure, where the affine motion prediction mode of the affine motion prediction coding reconstruction frame is the affine merge mode, the CU of the frame has both a width and height of 8 pixels or more. The motion candidate list may be an affine merge candidate list and may include up to 5 CPMVP candidates. The coding of the CU may include a merge index, which may point to a CPMVP candidate for affine merge.
[0021] The current control point motion vector (CPMV) of a CU may be generated based on candidate control point motion vector (CPMVP) factors derived from motion information of spatially adjacent blocks or temporally adjacent blocks to the current CU.
[0022] As shown in Figure 3, there are multiple spatially adjacent blocks of the current CU of a frame. The spatially adjacent blocks of the current CU may be blocks adjacent to the left of the current CU, or blocks adjacent to the top of the current CU. The spatially adjacent blocks have left-right and up-down relationships corresponding to the left-right and up-down directions in Figure 3. As in the example in Figure 3, the affine merge candidate list for a frame encoded according to the affine motion prediction mode, which is the affine merge mode, may include at most the following CPMVP candidates: The spatially adjacent block on the left (A0), The spatially adjacent block above (B0), The spatially adjacent block in the upper right (B1), The spatially adjacent block in the lower left (A1), and The spatially adjacent block in the upper left (B2).
[0023] Of the spatially adjacent blocks shown herein, block A0 may be the block to the left of the current CU302, block A1 may be the block to the left of the current CU302, block B0 may be the block above the current CU302, block B1 may be the block above the current CU302, and block B2 may be the block above the current CU302. The relative positioning of each spatially adjacent block with respect to the current CU302 or relative to each other is not further limited. There are no restrictions on the relative size of each spatially adjacent block with respect to the current CU302 or relative to each other.
[0024] The list of affine merge candidates for CUs of frames encoded according to the affine motion prediction mode, which is an affine merge mode, may include the following CPMVP candidates.
[0025] Up to two inheritance affine merge candidates, Candidates for constructing affine merges, and Zero motion vector.
[0026] Candidate inheritance affine merges may be derived from spatially adjacent blocks that have affine motion information, i.e., spatially adjacent blocks belonging to a CU that have CPMV.
[0027] Candidates for constructing an affine merge may be derived from spatially adjacent blocks and temporally adjacent blocks that do not possess affine motion information; that is, CPMV may be derived from spatially adjacent blocks and temporally adjacent blocks belonging to a CU that possess only translational motion information.
[0028] A zero motion vector may have a motion shift of (0,0).
[0029] Up to one inheritance affine merge candidate may be derived from searching the spatially adjacent blocks to the left of the current CU, and up to one inheritance affine merge candidate may be derived from searching the spatially adjacent blocks above the current CU. The left spatially adjacent blocks may be searched in the order A0 and A1, and the upper spatially adjacent blocks may be searched in the order B0, B1, and B2, in each case, for a first spatially adjacent block having affine motion information. If such a first spatially adjacent block is found in the left spatially adjacent blocks, the CPMVP candidate is derived from the CPMV of the first spatially adjacent block and added to the affine merge candidate list. If such a first spatially adjacent block is found in the upper spatially adjacent blocks, the CPMVP candidate is derived from the CPMV of the first spatially adjacent block and added to the affine merge candidate list. When two CPMVP candidates are derived in this way, a pruning check between the derived CPMVP candidates, that is, a check to see if the two derived CPMVP candidates are the same CPMVP candidate, is not performed.
[0030] Figure 4 shows an example of deriving a candidate for inherited affine merge. The current CU402 has a spatially adjacent block A to its left. Block A belongs to CU404. When block A is encoded according to a four-parameter affine model, CU404 may have the following affine motion information:
number
number
number
number
number
[0031] When block A is encoded according to a 6-parameter affine model, CU404 may additionally have the following affine motion information:
number
number
number
number
[0032] Figure 5 shows an example of deriving a constructive affine merge candidate. The constructive affine merge candidate may also be derived from the four CPMVs of the current CU502, each of which is derived by searching for spatially adjacent blocks of the current CU502 or from the temporally adjacent blocks of the current CU502.
[0033] The following blocks may be referenced in the derivation of CPMV: The spatially adjacent block on the left (A1), The spatially adjacent block on the left (A2), Upper spatially adjacent block (B1), The spatially adjacent block in the upper right (B0), The spatially adjacent block in the lower left (A0), The spatially adjacent block in the upper left (B2), The spatially adjacent block above (B3), and Temporal adjacent blocks (T).
[0034] The following CPMV may be derived for the current CU502: CPMV (CPMV1) in the upper left, CPMV (CPMV2) in the upper right corner, The CPMV (CPMV3) in the lower left, and CPMV (CPMV4) in the bottom right.
[0035] CPMV1 may also be derived by searching for spatially adjacent blocks B2, B3, and A2 in that order and selecting the first available spatially adjacent block according to criteria found in the relevant art, the details of which are not described herein.
[0036] CPMV2 may also be derived by searching for spatially adjacent blocks B1 and B0 in that order and similarly selecting the first available spatially adjacent block.
[0037] CPMV3 may also be derived by searching for spatially adjacent blocks A1 and A0 in that order, and similarly selecting the first available spatially adjacent block.
[0038] CPMV4 may be derived from the temporally adjacent block T, if available.
[0039] The constructor affine merge candidates may be constructed in a given order using the first available combination of CPMV for the current CU502 from the following combinations:
[0040] {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, and {CPMV1, CPMV3}.
[0041] When three CPMV combinations are used, a 6-parameter affine merge candidate is generated. When two CPMV combinations are used, a 4-parameter affine merge candidate is generated. The constructed affine merge candidate is then added to the affine merge candidate list.
[0042] For blocks that do not have affine motion information, for example, blocks belonging to a CU encoded according to the Temporal Motion Vector Predictor (TMVP) coding format, the encoding of the CU may include an interprediction indicator. The interprediction indicator may indicate a List 0 prediction that refers to a first reference picture list pointed to as List 0, a List 1 prediction that refers to a second reference picture list pointed to as List 1, or a biprediction that refers to two reference picture lists pointed to as List 0 and List 1, respectively. In the case of an interprediction indicator indicating a List 0 prediction or a List 1 prediction, the encoding of the CU may include a reference index that points to a reference picture in a reference framebuffer referenced by List 0 or List 1, respectively. In the case of an interprediction indicator indicating a biprediction, the encoding of the CU may include a first reference index that points to a first reference picture in a reference framebuffer referenced by List 0, and a second reference index that points to a second reference picture in a reference frame referenced by List 1.
[0043] The interprediction indicator may be encoded as a flag in the slice header of the intercoded frame. The reference index(s) may be encoded in the slice header of the intercoded frame. One or two motion vector differences (MVDs) corresponding to each reference index(s) may be further encoded.
[0044] In any of the above-mentioned combinations of CPMVs, if the reference indices of the CPMVs are different, i.e., if the CPMV can be derived from a CU that references a different reference picture which may have a different resolution, then that particular combination of CPMVs may be discarded and not used.
[0045] After adding any derived inheritance affine merge candidates and any construct affine merge candidates to the CU's affine merge candidate list, a zero motion vector, i.e., a motion shift of (0,0), is added to any remaining empty positions in the affine merge candidate list.
[0046] According to an exemplary embodiment of the present disclosure, the affine motion prediction mode of the affine motion prediction encoded and reconstructed frame is the affine adaptive motion vector prediction (AMVP) mode, the CU of the frame has both a width and height of 16 pixels or more. The applicability of the AMVP mode and whether a four-parameter affine motion model or a six-parameter affine motion model is used may be signaled by a bit-level flag carried in the video bitstream carrying the encoded frame data. The motion candidate list may be an AMVP candidate list and may contain up to two AMVP candidates.
[0047] The current CU's CPMV may be generated based on AMVP candidates derived from the movement information of spatially adjacent blocks to the current CU.
[0048] The AMVP candidate list for the CU of a frame encoded according to the affine motion prediction mode, which is AMVP mode, may include the following CPMVP candidates: Successor AMVP candidate, Candidates for AMVP in construction, Translational motion vector from adjacent CU, and Zero motion vector.
[0049] Inherited AMVP candidates may be derived in the same way as inherited affine merge candidates, except that each spatially adjacent block searched to derive an inherited AMVP candidate belongs to a CU that references the same reference picture as the current CU. No pruning checks are performed between inherited AMVP candidates and the AMVP candidate list when adding inherited AMVP candidates to the AMVP candidate list.
[0050] The constructed AMVP candidate may be derived in the same way as the constructed affine merge candidate, except that the selection of the first available spatially adjacent block is further carried out according to the criterion that the first available spatially adjacent block is intercoded and has a reference index that refers to the same reference picture as the current CU. Furthermore, according to AMVP implementations that do not support temporal control points, temporally adjacent blocks may not be searched.
[0051] If the current CU is encoded with a 4-parameter affine motion model and CPMV1 and CPMV2 of the current CU are available, then CPMV1 and CPMV2 are added to the AMVP candidate list as one candidate. If the current CU is encoded with a 6-parameter affine motion model and CPMV1, CPMV2, and CPMV3 of the current CU are available, then CPMV1, CPMV2, and CPMV3 are added to the AMVP candidate list as one candidate. Otherwise, the constructed AMVP candidate is not available to be added to the AMVP candidate list.
[0052] A translational motion vector can be a motion vector from a spatially adjacent block belonging to a CU that has only translational motion information.
[0053] A zero motion vector may have a motion shift of (0,0).
[0054] After adding any derived inherited affine merge candidates and any constructed affine merge candidates to the affine merge candidate lists for CU, CPMV1, CPMV2, and CPMV3, they are added to the AMVP candidate list in a given order as translational motion vectors to predict all CPMVs of the current CU, according to their respective availability. Then, zero motion vectors, i.e., motion vectors representing the (0,0) motion shift, are added to any remaining empty positions in the AMVP candidate list.
[0055] Motion information predicted according to DMVR may be predicted by biprediction. Biprediction may be performed on the current frame such that the motion information of the blocks in the reconstructed frame includes references to a first motion vector of a first reference block and a second motion vector of a second reference block, the first reference block having a first temporal distance from the current block and the second reference block having a second temporal distance from the current block. The first and second temporal distances may be in different temporal directions than the current block.
[0056] The first motion vector may be the motion vector of the block of the first referenced picture in the first referenced picture list pointed to as list 0, and the second motion vector may be the motion vector of the block of the second referenced picture in the second referenced picture list pointed to as list 1. The encoding of the CU to which the current block belongs may include a first reference index pointing to the first referenced picture of the referenced frame pointed to by list 0, and a second reference index pointing to the second referenced picture of the referenced frame pointed to by list 1.
[0057] Figure 6 shows a diagram of the DMVR dual prediction process based on template matching. In the first step of the DMVR dual prediction process, the initial first block 602 of the first reference picture 604 in List 0, referenced by the initial first motion vector mv0, and the initial second block 606 of the second reference picture 608 in List 1, referenced by the initial second motion vector mv1, are averaged to generate a weighted combination of the initial first block 602 and the initial second block 606. The weighted combination serves as template 610. Motion prediction of the current block 612 may be performed using the initial first motion vector referencing the initial first block 602 and the initial second motion vector referencing the initial second block 606.
[0058] In the second step of the DMVR biprediction process, the template 610 is compared by cost measurement to a first sample region of a first reference picture 604 adjacent to the initial first block 602 and a second sample region of a second reference image 608 adjacent to the initial second block 606. The cost measurement may utilize a preferred measure of image similarity, such as the sum of absolute differences or the sum of mean-removed absolute differences. If, within the first sample region, the subsequent first block 614 has the minimum cost measured against the template, the subsequent first motion vector mv0' referring to the subsequent first block 614 may replace the initial first motion vector mv0. If, within the second sample region, the subsequent second block 616 has the minimum cost measured against the template, the subsequent second motion vector mv1' referring to the subsequent second block 616 may replace the initial second motion vector mv1. Biprediction may then be performed for the current block 612 using mv0' and mv1'.
[0059] Figure 7 shows an exemplary block diagram of a video encoding process 700 according to an exemplary embodiment of the present disclosure.
[0060] The video encoding process 700 may obtain encoded frames from a source such as a bitstream 710. According to an exemplary embodiment of the present disclosure, considering the current frame 712 having position N in the bitstream, a previous frame 714 having position N-1 in the bitstream may have a resolution greater than or less than the resolution of the current frame, and the next frame 716 having position N+1 in the bitstream may have a resolution greater than or less than the resolution of the current frame.
[0061] The video coding process 700 may decode the current frame 712 to generate a reconstructed frame 718 and output the reconstructed frame 718 to a destination such as a reference frame buffer 790 or display buffer 792. The current frame 712 may also be input to a coding loop 720, which may include repeating the step of inputting the current frame 712 to a video decoder 722, generating a reconstructed frame 718 based on a previous reconstructed frame 794 in the reference frame buffer 790, inputting the reconstructed frame 718 to an in-loop upsampler or downsampler 724, generating an upsampled or downsampled reconstructed frame 796, and outputting the upsampled or downsampled reconstructed frame 796 to the reference frame buffer 790. Alternatively, the reconstructed frame 718 may be output from the loop, which may include inputting the reconstructed frame to a post-loop upsampler or downsampler 726, generating an upsampled or downsampled reconstructed frame 798, and outputting the upsampled or downsampled reconstructed frame 798 to a display buffer 792.
[0062] According to exemplary embodiments of this disclosure, the video decoder 722 may be any decoder that implements a motion prediction coding format, including but not limited to those coding formats described herein. Generating a reconstructed frame based on a previous reconstructed frame in the reference frame buffer 790 may include inter-coded motion prediction as described herein, where the previous reconstructed frame may be an upsampled or downsampled reconstructed frame output by an in-loop upsampler or downsampler 722 during a previous coding loop, and the previous reconstructed frame serves as a reference picture in the inter-coded motion prediction as described herein.
[0063] According to exemplary embodiments of the present disclosure, the in-loop upsampler or downsampler 724 and the post-loop upsampler or downsampler 726 may each implement an upsampling or downsampling algorithm, respectively, that is suitable for at least upsampled or downsampled pixel information of a frame encoded in motion predictive coding format. The in-loop upsampler or downsampler 724 and the post-loop upsampler or downsampler 726 may each implement an upsampling or downsampling algorithm, respectively, that is more suitable for upscaling and downscaling motion information, such as motion vectors.
[0064] The in-loop upsampler or downsampler 724 may utilize a relatively simple and computationally faster upsampling or downsampling algorithm compared to the algorithm utilized by the post-loop upsampler or downsampler 426, such that the upsampled or downsampled reconstructed frame 796 output by the in-loop upsampler or downsampler 724 can be input into the reference frame buffer 790 before the upsampled or downsampled reconstructed frame 796 is needed in a future iteration of the coding loop 720 to function as a previous reconstructed frame, while the upsampled or downsampled reconstructed frame 798 output by the post-loop upsampler or downsampler 726 may not be output in time before the upsampled or downsampled reconstructed frame 796 is thus needed. For example, the in-loop upsampler may utilize a training-independent interpolation, averaging, or bilinear upsampling algorithm, while the post-loop upsampler may utilize a trained upsampling algorithm.
[0065] Therefore, frames that serve as reference pictures when generating the reconstructed frame 718 of the current frame 712, such as the previous reconstructed frame 794, may be upsampled or downsampled according to the resolution of the current frame 712 relative to the resolutions of the previous frame 714 and the next frame 716. For example, if the current frame 712 has a higher resolution than either or both of the previous frame 714 and the next frame 716, the frame serving as the reference picture may be upsampled. The frame serving as the reference picture may be downsampled if the current frame 712 has a lower resolution than either or both of the previous frame 714 and the next frame 716.
[0066] Figures 8A, 8B, and 8C show exemplary flowcharts of a video encoding method 800 that implements resolution-adaptive video encoding according to an exemplary embodiment of the present disclosure, in which frames are encoded by affine motion prediction coding.
[0067] In step 802, the video decoder may obtain the current frame of the bitstream encoded by affine motion prediction coding, and affine merge mode or AMVP mode may be further enabled according to the bitstream signal. The current frame may have position N. A previous frame having position N-1 in the bitstream may have a resolution greater or less than that of the current frame, and a next frame having position N+1 in the bitstream may have a resolution greater or less than that of the current frame.
[0068] In step 804, the video decoder may retrieve one or more reference pictures from the reference frame buffer and compare the resolution of one or more reference pictures with the resolution of the current frame.
[0069] In step 806, if the video decoder determines that one or more of the reference pictures have a different resolution than the current frame, it may select a frame from the reference frame buffer that has the same resolution as the current frame, if available.
[0070] According to exemplary embodiments of the present disclosure, a frame having the same resolution as the current frame may be the most recent frame in a reference frame buffer having the same resolution as the current frame, and may not be the most recent frame in the reference frame buffer.
[0071] In step 808, the loop upsampler or downsampler may determine the ratio of the resolution of the current frame to the resolution of one or more reference pictures, and may scale the motion vectors of one or more reference pictures according to that ratio.
[0072] According to exemplary embodiments of the present disclosure, scaling motion vectors may include increasing or decreasing the magnitude of motion vectors.
[0073] In step 810A, the loop-upsampler or downsampler may further resize the interpreters of one or more reference pictures according to the ratio.
[0074] According to exemplary embodiments of the present disclosure, the interpreter may be, for example, motion information for motion prediction that references other reference pictures which may have different resolutions.
[0075] Alternatively, in step 810B, the in-loop upsampler or downsampler may detect the signaled upsample or downsample filter coefficients within the sequence header or picture header of the current frame and transmit the difference between the signaled filter coefficients and the filter coefficients of the current frame to the video decoder. The filter coefficients can be considered as the coefficients of the interpredictor. Thus, the difference between the interpredictor filter coefficients and the filter coefficients of the current frame allows the predicted motion information to be applied to the filter of the current frame.
[0076] In step 812, the video decoder may derive an affine merge candidate list or an AMVP candidate list for the current frame block. The derivation of the affine merge candidate list or AMVP candidate list may be carried out in accordance with the steps described herein. The derivation of CPMVP candidates or AMVP candidates in the derivation of the affine merge candidate list or AMVP candidate list may be further carried out in accordance with the steps described herein, respectively.
[0077] In step 814, the video decoder may select a CPMVP candidate or AMVP candidate from the affine merge candidate list or AMVP candidate list in accordance with the steps described herein, and derive the motion vector of the CPMVP candidate or AMVP candidate as the motion vector of the block of reconstructed frames.
[0078] In step 816, the video decoder may generate a reconstructed frame from the current frame based on one or more reference pictures and selected CPMVP or AMVP candidates.
[0079] The reconstructed frame may be predicted by referencing a selected reference picture having the same resolution as the current frame, by scaling or resizing the motion vectors or interpreters of other frames in the reference frame buffer according to the same resolution as the current frame, or by applying the difference between the signaled filter coefficients transmitted from the in-loop upsampler or downsampler to the filter of the current frame and the filter coefficients of the current frame while encoding the filter.
[0080] In step 818, the reconstructed frame may be input to at least one of the in-loop upsampler or downsampler, and the post-loop upsampler or downsampler.
[0081] In step 820, at least one of the intra-loop upsampler or downsampler, or the post-loop upsampler or downsampler, may generate an upsampled or downsampled reconstructed frame based on the reconstructed frame.
[0082] Multiple upsampled or downsampled reconstruction frames may be generated according to different resolutions of the multiple resolutions supported by the bitstream.
[0083] In step 822, the reconstructed frame and at least one of the upsampled or downsampled reconstructed frames may be input to at least one of the reference frame buffer and the display buffer.
[0084] If a reconstructed frame is input to the reference frame buffer, the reconstructed frame may be acquired as a reference picture and then upsampled or downsampled in a subsequent iteration of the coding loop as described with respect to step 806 above. If one or more upsampled or downsampled reconstructed frames are input to the reference frame buffer, one of the one or more upsampled or downsampled frames may be selected in a subsequent iteration of the coding loop as a frame having the same resolution as the current frame.
[0085] Figures 9A, 9B, and 9C show exemplary flowcharts of a video coding method 900 that implements resolution-adaptive video coding according to an exemplary embodiment of the present disclosure, in which motion information is predicted by a DMVR.
[0086] In step 902, the video decoder may obtain the current frame of the bitstream. The current frame may have position N. A previous frame at position N-1 in the bitstream may have a resolution greater than or less than the resolution of the current frame, and the next frame at position N+1 in the bitstream may have a resolution greater than or less than the resolution of the current frame.
[0087] In step 904, the video decoder may retrieve one or more reference pictures from the reference frame buffer and compare the resolution of one or more reference pictures with the resolution of the current frame.
[0088] In step 906, if the video decoder determines that one or more of the reference pictures have a different resolution than the current frame, the in-loop upsampler or downsampler may, if available, select a frame from the reference frame buffer that has the same resolution as the current frame.
[0089] According to exemplary embodiments of the present disclosure, a video decoder may select a frame from a reference frame buffer having the same resolution as the current frame. The frame having the same resolution as the current frame may be the most recent frame in the reference frame buffer having the same resolution as the current frame, and may not be the most recent frame in the reference frame buffer.
[0090] In step 908, the loop upsampler or downsampler may determine the ratio of the resolution of the current frame to the resolution of one or more reference pictures, and may resize the pixel pattern of one or more reference pictures according to that ratio.
[0091] According to exemplary embodiments of the present disclosure, resizing the pixel patterns of one or more reference pictures can facilitate a vector refinement process at different resolutions by a DMVR, such as the aforementioned step of comparing a template with a first sample region of a first reference picture adjacent to an initial first block and a second sample region of a second reference picture adjacent to an initial second block, by cost measurement.
[0092] In step 910, the video decoder may perform biprediction and vector refinement on the current frame based on the first and second reference frames of the reference frame buffer, in accordance with the steps described herein.
[0093] In step 912, the video decoder may generate a reconstructed frame from the current frame based on the first reference frame and the second reference frame.
[0094] The reconstructed frame may be predicted by a reference to a selected reference picture having the same resolution as the current frame, or by resizing the pixel patterns of other frames in the reference frame buffer according to the same resolution as the current frame.
[0095] In step 914, the reconstructed frame may be input to at least one of the in-loop upsampler or downsampler, and the post-loop upsampler or downsampler.
[0096] In step 916, at least one of the in-loop upsampler or downsampler, or the post-loop upsampler or downsampler, may generate an upsampled or downsampled reconstructed frame based on the reconstructed frame.
[0097] Multiple upsampled or downsampled reconstruction frames may be generated according to different resolutions of the multiple resolutions supported by the bitstream.
[0098] In step 918, the reconstructed frame and at least one of the upsampled or downsampled reconstructed frames may be input to at least one of the reference frame buffer and the display buffer.
[0099] If a reconstructed frame is input to the reference frame buffer, the reconstructed frame may be acquired as a reference picture and then upsampled or downsampled in a subsequent iteration of the coding loop as described with respect to step 906 above. If one or more upsampled or downsampled reconstructed frames are input to the reference frame buffer, one of the one or more upsampled or downsampled frames may be selected in a subsequent iteration of the coding loop as a frame having the same resolution as the current frame.
[0100] Figure 10 shows an exemplary system 1000 for implementing the processes and methods described above for implementing resolution-adaptive video coding in motion prediction coding format.
[0101] The techniques and mechanisms described herein may be implemented by multiple instances of System 1000, as well as by any other computing devices, systems, and / or environments. System 1000 shown in Figure 10 is merely an example of a system and is not intended to imply any limitations on the scope or functionality of any computing device used to perform the processes and / or procedures described above. Other well-known computing devices, systems, environments, and / or configurations that may be suitable for use with the embodiments include, but are not limited to, personal computers, server computers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, game consoles, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and / or implementations using field-programmable gate arrays ("FPGAs") and application-specific integrated circuits ("ASICs").
[0102] The system 1000 may include one or more processors 1002 and a system memory 1004 communicably coupled to the processor(s) 1002. The processor(s) 1002 may execute one or more modules and / or processes to perform various functions on the processor(s) 1002. In some embodiments, the processor(s) 1002 may include a central processing unit (CPU), a graphics processing unit (GPU), both a CPU and a GPU, or other processing units or components known in the art. Additionally, each of the processor(s) 1002 may have its own local memory that can also store program modules, program data, and / or one or more operating systems.
[0103] Depending on the exact configuration and type of system 1000, system memory 1004 may be volatile such as RAM, non-volatile such as ROM, flash memory, miniature hard drives, and memory cards, or a combination of these. System memory 1004 may include one or more computer executable modules 1006 that can be run by processors 1002.
[0104] Module 1006 may include, but is not limited to, a decoder module 1008 and an upsampler or downsampler module 1010. Decoder module 1008 may include a frame acquisition module 1012, a reference picture acquisition module 1014, a frame selection module 1016, a candidate list derivation module 1018, a motion prediction module 1020, a reconstructed frame generation module 1022, and an upsampler or downsampler input module 1024. Upsampler or downsampler module 1010 may include a ratio determination module 1026, a scaling module 1030, an interpreter resizing module 1032, a filter coefficient detection and difference transmission module 1034, an upsampling or downsampling reconstructed frame generation module 1036, and a buffer input module 1038.
[0105] The frame acquisition module 1012 may be configured to acquire the current frame of a bitstream encoded in affine motion predictive coding format, as described above with reference to Figure 8.
[0106] The reference picture acquisition module 1014 may be configured to acquire one or more reference pictures from the reference frame buffer and compare the resolution of the one or more reference pictures with the resolution of the current frame, as described above with reference to Figure 8.
[0107] The frame selection module 1016 may be configured, as described above with reference to Figure 8, to select a frame from the reference frame buffer that has the same resolution as the current frame when the reference picture acquisition module 1014 determines that one or more of the resolutions of one or more reference pictures are different from the resolution of the current frame.
[0108] The candidate list derivation module 1018 may be configured to derive an affine merge candidate list or AMVP candidate list for the blocks of the current frame, as described above with reference to Figure 8.
[0109] The motion prediction module 1020 may be configured to select a CPMVP or AMVP candidate from the derived affine merge candidate list or AMVP candidate list, as described above with reference to Figure 8, and to derive the motion vector of the CPMVP or AMVP candidate as the motion vector of the block in the reconstructed frame.
[0110] The reconstructed frame generation module 1022 may be configured to generate a reconstructed frame from the current frame based on one or more reference pictures and selected motion candidates.
[0111] The upsampler or downsampler input module 1024 may be configured to input the reconstructed frame to the upsampler or downsampler module 1010.
[0112] The ratio determination module 1026 may be configured to determine the ratio between the resolution of the current frame and the resolution of one or more reference pictures.
[0113] The scaling module 1030 may be configured to scale the motion vectors of one or more reference pictures according to a ratio.
[0114] The interpreter resizing module 1032 may be configured to resize the interpreters of one or more reference pictures according to a ratio.
[0115] The filter coefficient detection and difference transmission module 1034 may be configured to detect the upsample or downsample filter coefficients signaled within the sequence header or picture header of the current frame and to transmit the difference between the signaled filter coefficients and the filter coefficients of the current frame to the video decoder.
[0116] The upsampling or downsampling reconstructed frame generation module 1036 may be configured to generate an upsampling or downsampling reconstructed frame based on the reconstructed frame.
[0117] The buffer input module 1038 may be configured to input an upsampled or downsampled reconstructed frame to at least one of the reference frame buffer and the display buffer, as described above with reference to Figure 8.
[0118] System 1000 may also include an input / output (I / O) interface 1040 for receiving bitstream data to be processed and for outputting reconstructed frames to a reference frame buffer and / or display buffer. System 1000 may also include a communication module 1050 that enables System 1000 to communicate with other devices (not shown) over a network (not shown). The network may include wired media such as the Internet, a wired network or direct wired connection, as well as wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0119] Figure 11 shows an exemplary system 1100 for implementing the processes and methods described above for implementing resolution-adaptive video coding in a motion-predictive coding format.
[0120] The techniques and mechanisms described herein may be implemented by multiple instances of System 1100, as well as by any other computing devices, systems, and / or environments. System 1100 shown in Figure 11 is merely an example of a system and is not intended to imply any limitations on the scope or functionality of any computing device used to perform the processes and / or procedures described above. Other well-known computing devices, systems, environments, and / or configurations that may be suitable for use with the embodiments include, but are not limited to, personal computers, server computers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, game consoles, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and / or implementations using field-programmable gate arrays ("FPGAs") and application-specific integrated circuits ("ASICs").
[0121] The system 1100 may include one or more processors 1102 and a system memory 1104 communicably coupled to the processors 1102. The processors 1102 may execute one or more modules and / or processes to perform various functions on the processors 1102. In some embodiments, the processors 1102 may include a central processing unit (CPU), a graphics processing unit (GPU), both a CPU and a GPU, or other processing units or components known in the art. Additionally, each of the processors 1102 may have its own local memory that can also store program modules, program data, and / or one or more operating systems.
[0122] Depending on the exact configuration and type of system 1100, the system memory 1104 may be volatile such as RAM, non-volatile such as ROM, flash memory, miniature hard drives, and memory cards, or a combination of these. The system memory 1104 may include one or more computer executable modules 1106 that can be run by the processor(s) 1102.
[0123] Module 1106 may include, but is not limited to, a decoder module 1108 and an upsampler or downsampler module 1110. Decoder module 1108 may include a frame acquisition module 1112, a reference picture acquisition module 1114, a dual prediction module 1116, a vector refinement module 1118, an upsampling or downsampling reconstruction frame generation module 1120, and an upsampler or downsampler input module 1122. Upsampler or downsampler module 1110 may include a ratio determination module 1124, a pixel pattern resizing module 1128, an upsampling or downsampling reconstruction frame generation module 1130, and a buffer input module 1132.
[0124] The frame acquisition module 1112 may be configured to acquire the current frame of a bitstream encoded in BIO encoding format, as described above with reference to Figure 9.
[0125] The reference picture acquisition module 1114 may be configured to acquire one or more reference pictures from the reference frame buffer and compare the resolution of the one or more reference pictures with the resolution of the current frame, as described above with reference to Figure 9.
[0126] The dual prediction module 1116 may be configured to perform dual prediction for the current frame based on the first and second reference frames of the reference frame buffer, as described above with reference to Figure 9.
[0127] The vector refinement module 1118 may be configured to perform vector refinement during the biprediction process based on the first and second reference frames of the reference frame buffer, as described above with reference to Figure 6.
[0128] The reconstructed frame generation module 1120 may be configured to generate a reconstructed frame from the current frame based on a first reference frame and a second reference frame.
[0129] The upsampler or downsampler input module 1122 may be configured to input the reconstructed frame to the upsampler or downsampler module 1110.
[0130] The ratio determination module 1124 may be configured to determine the ratio between the resolution of the current frame and the resolution of one or more reference pictures.
[0131] The pixel pattern resizing module 1128 may be configured to resize the pixel patterns of one or more reference pictures according to their ratios.
[0132] The upsampling or downsampling reconstructed frame generation module 1130 may be configured to generate an upsampling or downsampling reconstructed frame based on the reconstructed frame.
[0133] The buffer input module 1132 may be configured to input an upsampled or downsampled reconstructed frame to at least one of the reference frame buffer and the display buffer, as described above with reference to Figure 9.
[0134] System 1100 may also include an input / output (I / O) interface 1140 for receiving bitstream data to be processed and for outputting reconstructed frames to a reference frame buffer and / or display buffer. System 1100 may also include a communications module 1150 that enables System 1100 to communicate with other devices (not shown) over a network (not shown). The network may include wired media such as the Internet, a wired network or direct wired connection, as well as wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0135] Some or all of the operations of the methods described above may be performed by executing computer-readable instructions stored on a computer-readable storage medium, as defined below. As used in the specification and claims, the term “computer-readable instructions” includes routines, applications, application modules, program modules, programs, components, data structures, algorithms, and the like. Computer-readable instructions may be implemented in a variety of system configurations, including single-processor or multi-processor systems, minicomputers, mainframe computers, personal computers, handheld computing devices, microprocessor-based programmable consumer electronics, and combinations thereof.
[0136] Computer-readable storage media may include volatile memory (such as random-access memory (RAM)) and / or non-volatile memory (such as read-only memory (ROM) and flash memory). Computer-readable storage media may also include, but are not limited to, additional removable and / or non-removable storage devices that can provide non-volatile storage devices such as computer-readable instructions, data structures, and program modules, including flash memory, magnetic storage devices, optical storage devices, and / or tape storage devices.
[0137] Non-temporary computer-readable storage media are an example of computer-readable media. Computer-readable media include at least two types of computer-readable media: computer-readable storage media and computer-readable communication media. Computer-readable storage media include volatile and non-volatile, removable and non-removable media, implemented by any process or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, phase-change memory (PRAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, read-only memory for compact discs (CD-ROM), digital versatile discs (DVDs) or other optical storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or any other non-transmission media that may be used to store information for access by computing devices. In contrast, communication media can embody computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms. Computer-readable storage media, as defined herein, do not include communication media.
[0138] Computer-readable instructions stored in one or more non-temporary computer-readable storage media, which, when executed by one or more processors, can perform the operations described above with reference to Figures 1 to 11. Generally, computer-readable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular abstract data type. The order in which the operations are described is not intended to be interpreted as a restriction, and any number of the operations described may be combined in any order and / or in parallel to implement a process.
[0139] Through the technical solutions described above, this disclosure provides inter-coded resolution-adaptive video coding supported by motion-predictive coding formats, improving the video coding process in multiple motion-predictive coding formats by enabling the coding of resolution changes between frames while allowing motion vectors to reference previous frames. Thus, bandwidth savings from inter-coding are maintained, bandwidth savings from motion-predictive coding are achieved, allowing reference frames to be used to predict the motion vectors of subsequent frames, and bandwidth savings from adaptive downsampling and upsampling depending on bandwidth availability are also achieved simultaneously, resulting in a significant improvement in network costs during video coding and content delivery, while reducing additional data transmission that would offset or impair these savings. [Examples]
[0140] A. A method comprising: obtaining the current frame of a bitstream; obtaining one or more reference pictures from a reference frame buffer having a resolution different from the resolution of the current frame; resizing the interpreters of one or more reference pictures; and generating a reconstructed frame from the current frame based on the one or more reference pictures and motion information of one or more blocks of the current frame, wherein the motion information includes at least one interpreter.
[0141] B. The method according to paragraph A, further comprising: comparing the resolution of one or more reference pictures with the resolution of the current frame; determining that the resolution of one or more of the one or more reference pictures differs from the resolution of the current frame, selecting a frame from the reference frame buffer having the same resolution as the current frame; determining the ratio between the resolution of the current frame and the resolution of one or more reference pictures; resizing one or more reference pictures according to the ratio to match the resolution of the current frame; upsampling or downsampling the interpreters of one or more reference pictures according to the ratio; and scaling the motion vectors of one or more reference pictures according to the ratio.
[0142] C. The method of paragraph A, further comprising deriving an affine merge candidate list or AMVP candidate list for the current frame, selecting a CPMVP candidate or AMVP candidate from the affine merge candidate list or AMVP candidate list, respectively, and deriving the motion vectors of motion candidates as the block motion vectors of the reconstructed frame.
[0143] D. The method of paragraph C, further comprising deriving at least one of the inheritance affine merge candidates and the construction affine merge candidates, and adding at least one of the inheritance affine merge candidates and the construction affine merge candidates to an affine merge candidate list or an AMVP candidate list.
[0144] E. The method according to paragraph A, further comprising generating a reconstructed frame from the current frame based on one or more reference pictures and at least one interpreter; inputting the reconstructed frame into at least one of an in-loop upsampler or downsampler and a post-loop upsampler or downsampler; generating an upsampled or downsampled reconstructed frame based on the reconstructed frame; and inputting the upsampled or downsampled reconstructed frame into at least one of a reference frame buffer and a display buffer.
[0145] F. A method comprising: obtaining the current frame of a bitstream; obtaining one or more reference pictures from a reference frame buffer; comparing the resolution of one or more reference pictures with the resolution of the current frame; and, if it is determined that the resolution of one or more of the one or more reference pictures is different from the resolution of the current frame, resizing the pixel patterns of the one or more reference pictures according to the resolution of the current frame.
[0146] G. The method of paragraph F, further comprising performing a biprediction for the current frame based on a first and second reference frame of the reference frame buffer.
[0147] H. The method of paragraph G, further comprising performing a biprediction on the current frame, which further includes performing a vector refinement on the current frame based on a first reference frame and a second reference frame of a reference frame buffer.
[0148] I. The method according to paragraph H, further comprising: generating a reconstructed frame from the current frame based on a first reference frame and a second reference frame; inputting the reconstructed frame into at least one of an in-loop upsampler or downsampler and a post-loop upsampler or downsampler; generating an upsampled or downsampled reconstructed frame based on the reconstructed frame; and inputting the upsampled or downsampled reconstructed frame into at least one of a reference frame buffer and a display buffer.
[0149] J. A method comprising: acquiring the current frame of a bitstream, the bitstream comprising frames having multiple resolutions; acquiring one or more reference pictures from a reference frame buffer; generating a reconstructed frame from the current frame based on one or more reference pictures and motion information of one or more blocks of the current frame, wherein the motion information comprises at least one interpreter; and upsampling or downsampling the current reconstructed frame for each of the multiple resolutions to generate an upsampled or downsampled reconstructed frame that matches the respective resolution.
[0150] K. The method according to paragraph J, further comprising detecting signaled up-sampling or down-sampling filter coefficients for at least one of one or more reference pictures.
[0151] The method according to paragraph K, further comprising applying the difference between the filter coefficients of the L. inter predictor and the filter coefficients of the current frame to the encoding of the filter for the current frame.
[0152] The method of paragraph J, further comprising inputting M. reconstructed frames and each upsampling or downsampling reconstructed frame into a reference frame buffer.
[0153] N. A system comprising one or more processors and memory communicably coupled to one or more processors, wherein the memory stores computer executable modules that can be executed by one or more processors, and the computer executable modules perform related operations when executed by one or more processors, and the computer executable modules further comprising a frame acquisition module configured to acquire the current frame of a bitstream and a reference picture acquisition module configured to acquire one or more reference pictures from a reference frame buffer and compare the resolution of one or more reference pictures with the resolution of the current frame.
[0154] The system described in paragraph N, further comprising a frame selection module configured to select a frame from a reference frame buffer having the same resolution as the current frame when the reference picture acquisition module determines that one or more of the resolutions of one or more reference pictures are different from the resolution of the current frame.
[0155] P. The system described in paragraph O further comprises a candidate list derivation module configured to derive an affine merge candidate list or an AMVP candidate list for a block in the current frame.
[0156] Q. The system described in paragraph P, further comprising a motion prediction module configured to select a CPMVP or AMVP candidate, respectively, from a derived list of affine merge candidates or AMVP candidates.
[0157] The system described in paragraph Q, wherein the motion prediction module is further configured to derive the motion vectors of the CPMVP or AMVP candidate as the motion vectors of the blocks in the reconstructed frame.
[0158] S. The system according to paragraph N, further comprising: a reconstructed frame generation module configured to generate a reconstructed frame from the current frame based on one or more reference pictures and selected motion candidates; an upsampler or downsampler input module configured to input the reconstructed frame to an upsampler or downsampler module; a ratio determination module configured to determine the ratio between the resolution of the current frame and the resolution of one or more reference pictures; an interpredictor resizing module configured to resize the interpredictors of one or more reference pictures according to the ratio; a filter coefficient detection and difference transmission module configured to detect signaled upsample or downsample filter coefficients in the sequence header or picture header of the current frame and to send the difference between the signaled filter coefficients and the filter coefficients of the current frame to a video decoder; a scaling module configured to scale the motion vectors of one or more reference pictures according to the ratio; an upsampling or downsampled reconstructed frame generation module configured to generate an upsampled or downsampled reconstructed frame based on the reconstructed frame; and a buffer input module configured to input the upsampled or downsampled reconstructed frame to at least one of a reference frame buffer and a display buffer.
[0159] A system comprising one or more processors and memory communicatively coupled to one or more processors, wherein the memory stores computer executable modules that can be executed by one or more processors, and the computer executable modules perform related operations when executed by one or more processors, and the computer executable modules include a frame acquisition module configured to acquire the current frame of a bitstream, and a reference picture acquisition module configured to acquire one or more reference pictures from a reference frame buffer and compare the resolution of one or more reference pictures with the resolution of the current frame.
[0160] The system described in paragraph T, further comprising a biprediction module configured to perform biprediction on the current frame based on a first reference frame and a second reference frame of a reference frame buffer.
[0161] V. The system according to paragraph U, further comprising a vector refinement module configured to perform vector refinement during a biprediction process based on a first and second reference frame of a reference frame buffer.
[0162] The system described in paragraph V further comprises: a reconstructed frame generation module configured to generate a reconstructed frame from the current frame based on a first reference frame and a second reference frame; an upsampler or downsampler input module configured to input the reconstructed frame to an upsampler or downsampler module; an upsampling or downsampled reconstructed frame generation module configured to generate an upsampled or downsampled reconstructed frame based on the reconstructed frame; and a buffer input module configured to input the upsampled or downsampled reconstructed frame to at least one of a reference frame buffer and a display buffer.
[0163] While the subject matter has been described in language specific to its structural features and / or methodological behavior, it should be understood that the subject matter as defined in the attached claims is not necessarily limited to the specific features or behaviors described. Rather, specific features and behaviors are disclosed as exemplary forms of implementing the claims. [Explanation of symbols]
[0164] 302 Current CU 402 Current CU 404 CU 502 Current CU 602 The first block of the initial period 604 First reference picture 606 The first second block 608 Second reference picture 610 Template 612 Current Block 614 The first subsequent block 616 The second block that follows 700 Video Encoding Processes 710 bitstream 712 Current Frame 714 Previous frames 716 Next frame 718 Reconstruction Frame 720 encoding loop 722 Video Decoder 724 Loop Upsampler or Downsampler 726 Loop upsampler or downsampler 790 Reference frame buffer 792 Display buffer 794 Previous Reconstruction Frames 796 upsampling or downsampling reconstruction frames 798 upsampling or downsampling reconstruction frames 1000 System 1002 processors (multiple processors allowed) 1004 System Memory 1006 Modules 1008 Decoder Module 1010 Upsampler or Downsampler Module 1012 Frame Acquisition Module 1014 Reference Picture Acquisition Module 1016 Frame Selection Module 1018 Candidate List Derivation Module 1020 Motion Prediction Module 1022 Reconstruction Frame Generation Module 1024 Upsampler or Downsampler Input Module 1026 Ratio Determination Module 1030 Scaling Module 1032 Interpreter Resize Module 1034 Filter coefficient detection and differential transmission module 1036 Upsampling or Downsampling Reconstruction Frame Generation Module 1038 Buffer Input Module 1040 Input / Output (I / O) Interface 1050 Communication Module 1100 System 1102 processors (multiple processors possible) 1104 System Memory 1106 Module 1108 Decoder Module 1110 Upsampler or Downsampler Module 1112 Frame Acquisition Module 1114 Reference Picture Acquisition Module 1116 Dual Prediction Module 1118 Vector Refinement Module 1120 Upsampling or Downsampling Reconstruction Frame Generation Module 1122 Upsampler or Downsampler Input Module 1124 Ratio Determination Module 1128 Pixel Pattern Resizing Module 1130 Upsampling or Downsampling Reconstruction Frame Generation Module 1132 Buffer Input Module 1140 Input / Output (I / O) Interfaces 1150 Communication Module
Claims
1. A method that is executed by one or more processors, To get the current frame of the bitstream, Obtaining one or more reference pictures from a reference frame buffer, wherein one or more of the resolutions of the one or more reference pictures are different from the resolution of the current frame. Selecting a picture from the one or more reference pictures that has the same resolution as the current frame, Predicting the reconstructed frame by referencing the selected reference picture having the same resolution as the current frame, A method comprising: at least one of an in-encoding loop upsampler or downsampler, or a post-encoding loop upsampler or downsampler, generating an upsampled or downsampled reconstructed frame based on the reconstructed frame.
2. The method according to claim 1, further comprising deriving an affine merge candidate list or AMVP candidate list for the blocks of the current frame, wherein the affine merge candidate list or the AMVP candidate list each comprises a plurality of CPMVP candidates or AMVP candidates.
3. The method according to claim 2, wherein deriving the affine merge candidate list or the AMVP candidate list includes deriving up to two inheritance affine merge candidates.
4. The method according to claim 2, wherein deriving the affine merge candidate list or the AMVP candidate list includes deriving a construct affine merge candidate.
5. From the derived list of affine merge candidates or AMVP candidates, select either a CPMVP candidate or an AMVP candidate, respectively. The method according to claim 2, further comprising deriving motion information of the CPMVP candidate or the AMVP candidate as motion information of the block in the current frame.
6. The motion information includes a reference to a reference picture, and the motion information of the CPMVP candidate or the AMVP candidate is derived. The method according to claim 5, further comprising generating a plurality of CPMVs based on the reference to motion information of a reference picture.
7. A computer program configured to cause one or more processors to execute the method described in any one of claims 1 to 6.
8. One or more processors, The system comprises a memory communicatively coupled to one or more processors, the memory storing computer executable modules that can be executed by the one or more processors, the computer executable modules performing associated operations when executed by the one or more processors, and the computer executable modules A frame acquisition module configured to acquire the current frame of a bitstream, A reference frame acquisition module configured to acquire one or more reference pictures from a reference frame buffer, wherein one or more of the one or more reference pictures have a resolution different from the resolution of the current frame, A picture having the same resolution as the current frame is selected from the one or more reference pictures. The reconstructed frame is predicted by a reference to the selected reference picture having the same resolution as the current frame. A system in which upsampled or downsampled reconstructed frames are generated based on the reconstructed frames by at least one of the following: an in-encoding loop upsampler or downsampler, or a post-encoding loop upsampler or downsampler.
9. The system according to claim 8, further comprising a candidate list derivation module configured to derive an affine merge candidate list or an AMVP candidate list for the current frame block, wherein the affine merge candidate list or the AMVP candidate list each comprises a plurality of CPMVP candidates or AMVP candidates.
10. The system according to claim 9, wherein deriving the affine merge candidate list or the AMVP candidate list includes deriving up to two inheritance affine merge candidates.
11. The system according to claim 9, wherein deriving the affine merge candidate list or the AMVP candidate list includes deriving a construct affine merge candidate.
12. The system according to claim 9, further comprising a motion prediction module configured to select a CPMVP candidate or an AMVP candidate from the derived affine merge candidate list or AMVP candidate list, respectively, and to derive motion information of the CPMVP candidate or the AMVP candidate as motion information of the block in the current frame.
13. The motion information includes a reference to the motion information of a reference picture, and the motion prediction module, The system according to claim 12, further configured to generate a plurality of CPMVs based on the aforementioned reference to motion information of a reference picture.