Video decoding device and video encoding device
Patent Information
- Application Number
- JP2022108070
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-07-05
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video encoding methods, such as VVC, lack effective processes for handling wraparound and reference picture resampling (RPR) processing simultaneously, leading to reduced encoding efficiency and incorrect OOB determination when applied to 360-degree panoramic images or fluctuating transmission rates.
A video decoding device that includes an OOB determination unit to assess whether a reference block is subject to out-of-boundary processing, applying wraparound and RPR processing as needed, and deriving mask data to handle OOB conditions, thereby reducing calculation overhead.
Enhances encoding efficiency by allowing simultaneous wraparound and RPR processing, reducing calculation requirements and improving image quality in videos with spatial continuity and variable transmission rates.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] An embodiment of the present invention relates to a video decoding device and a video encoding device. [Background technology]
[0002] In order to efficiently transmit or record moving images, a moving image encoding device is used that generates encoded data by encoding moving images, and a moving image decoding device is used that generates a decoded image by decoding the encoded data.
[0003] Specific examples of video coding methods include H.264 / AVC, High-Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC).
[0004] In such a video coding method, images (pictures) constituting a video are divided into slices obtained by dividing the images, coding tree units (CTUs) obtained by dividing the slices, coding tree units (CTUs) obtained by dividing the coding tree units, and so on. The coding unit (sometimes called a coding unit (CU)) that is to be encoded, and The coding unit is divided into transform units (TUs), which are managed in a hierarchical structure, and the coding unit is encoded / decoded for each CU.
[0005] In such a video coding method, a predicted image is usually generated based on a locally decoded image obtained by coding / decoding an input image, and a prediction error (sometimes called a "difference image" or "residual image") obtained by subtracting the predicted image from the input image (original image) is coded. Methods for generating predicted images include inter-frame prediction (inter prediction) and intra-frame prediction (intra prediction). In VVC, as shown in FIG. 8(a), a wrap-around process is performed on the coordinates of a motion vector to make the left and right ends of a picture continuous in the horizontal coordinate system, thereby making it possible to perform motion compensation. Therefore, by applying the wrap-around process to a video in which the left and right ends of a picture are spatially continuous, such as a 360-degree panoramic image or a 360-degree image, it is possible to improve the coding efficiency. In addition, in VVC, as shown in FIG. 8(b), a reference picture resampling (RPR) process is possible, which changes the resolution on a picture-by-picture basis to perform motion compensation. By applying the RPR process to a service with a variable transmission rate, such as video distribution on the Internet, it is possible to improve the image quality.
[0006] In addition, non-patent document 1 discloses an OOB (Out-Of-Boundary) processing technology that replaces part or all of an area of one bi-predictive reference block that extends outside the picture of a reference image with part or all of an area of the other reference block. [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] “AHG12: Enhanced bi-directional motion compensation”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 25th Meeting, by teleconference, JVET-Y0125 Summary of the Invention [Problem to be solved by the invention]
[0008] In Non-Patent Document 1, there is no OOB determination process that corresponds to the wraparound process. There is a problem that pour-around processing and OOB processing cannot be applied at the same time. When the rounding process is applied, if the reference block is outside the left edge of the picture, as shown in Figure 8, The horizontal coordinate system refers to the area to the right of the picture, so that the horizontal coordinate system cannot be outside the picture. In addition, there is a problem that the coding efficiency may decrease by applying OOB processing. In the literature 1, there is no OOB judgment process corresponding to the RPR process, and there is a problem that the RPR process and the OOB process cannot be applied simultaneously. The image is reduced by a predetermined scale and stored in the reference picture buffer. This causes a problem that the coordinate system of the reference picture differs from the coordinate system of the target picture, and therefore, if the RPR process and the OOB process are used together, the OOB determination process cannot be performed correctly. [Means for solving the problem]
[0009] In order to solve the above problem, a video decoding device according to an aspect of the present invention determines whether a reference block is a target of OOB processing by comparing coordinates of a reference block with coordinates of a picture. An OOB determination unit that determines whether the pixel coordinates of the reference block are correct, and a pixel coordinates of the pixel coordinates of the reference block are compared with the pixel coordinates of the picture. A video having an OOB mask derivation unit that derives mask data indicating whether each pixel can be used or not by 1. An image decoding device, comprising: The OOB determination unit performs wraparound processing on the reference picture including the reference block. Detects whether the above reference block is subject to OOB processing by detecting whether it is applied The present invention is characterized by determining The OOB determination unit performs a reference picture resampling on the reference picture including the reference block. Detect whether ring processing is applied and whether the above reference block is subject to OOB processing The present invention is characterized in that it is determined whether or not Effect of the Invention
[0010] According to an aspect of the present invention, the amount of calculation required for OOB processing in video encoding / decoding processing is reduced. It can be reduced. [Brief description of the drawings]
[0011] [Figure 1] 1 is a schematic diagram showing a configuration of an image transmission system according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a diagram showing a hierarchical structure of data in an encoded stream. [Diagram 3] FIG. 1 is a schematic diagram showing a configuration of a video decoding device. [Figure 4] FIG. 13 is a schematic diagram showing a configuration of an inter-prediction image generating unit. [Diagram 5] FIG. 13 is a schematic diagram showing a configuration of an inter-prediction parameter derivation unit. [Figure 6] 11 is a flowchart illustrating a schematic operation of the video decoding device. [Figure 7] FIG. 1 is a block diagram showing a configuration of a video encoding device. [Figure 8] FIG. 1 is a diagram illustrating an example of wraparound processing and reference picture resampling processing of the prior art. [Figure 9] 11 is a flowchart illustrating a schematic operation of OOB processing. [Figure 10] FIG. 11 is a schematic diagram illustrating an example of an OOB determination process. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] (First embodiment) Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0013] FIG. 1 is a schematic diagram showing the configuration of an image transmission system 1 according to this embodiment.
[0014] The image transmission system 1 is a system that transmits an encoded stream obtained by encoding a target image, decodes the transmitted encoded stream, and displays an image. The image transmission system 1 includes a video encoding device (image encoding device) 11, a network 21, a video decoding device (image decoding device) 31, and a video display device (image display device) 41.
[0015] An image T is input to the video encoding device 11 .
[0016] The network 21 transmits the encoded stream Te generated by the video encoding device 11 to the video decoding device 31. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 21 is not necessarily limited to a bidirectional communication network, and may be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. The network 21 may also be replaced by a storage medium on which the encoded stream Te is recorded, such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark).
[0017] The video decoding device 31 decodes each of the coded streams Te transmitted by the network 21, and generates one or more decoded images Td.
[0018] The moving image display device 41 displays all or part of one or more decoded images Td generated by the moving image decoding device 31. The moving image display device 41 includes, for example, a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. Examples of the form of the display include a stationary type, a mobile type, and an HMD. Also, when the moving image decoding device 31 has high processing power, an image with high image quality is displayed, and when it has only low processing power, an image that does not require high processing power and display ability is displayed.
[0019] <Operator> The operators used in this specification are described below.
[0020] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, and | is a bitwise OR , |= is an OR assignment operator, and || indicates a logical OR.
[0021] x?y:z is a ternary operator that takes y when x is true (non-zero) and z when x is false (0).
[0022] Clip3(a,b,c) is a function that clips c to a value between a and b (inclusive), returns a if c < a, returns b if c > b, and returns c otherwise (where a <= b).
[0023] ClipH(o, W, x) is a function that returns x if x < 0, returns x - o if x > W - 1, and returns x otherwise .
[0024] sign(a) is a function that returns 1 if a > 0, 1 if a == 0, and -1 if a < 0.
[0025] abs(a) is a function that returns the absolute value of a.
[0026] Int(a) is a function that returns the integer value of a.
[0027] floor(a) is a function that returns the largest integer less than or equal to a.
[0028] ceil(a) is a function that returns the smallest integer greater than or equal to a.
[0029] a / d represents the division of a by d (truncated to an integer).
[0030] <Structure of the coding stream Te> Before describing in detail the video encoding device 11 and the video decoding device 31 according to this embodiment, the data structure of the encoded stream Te generated by the video encoding device 11 and decoded by the video decoding device 31 will be described.
[0031] FIG. 2 is a diagram showing a hierarchical structure of data in the coded stream Te. The frame Te illustratively includes a sequence and a plurality of pictures constituting the sequence. (a) to (f) of Fig. 2 are diagrams showing a coded video sequence that defines the sequence SEQ, a coded picture that defines the picture PICT, a coded slice that defines the slice S, coded slice data that defines the slice data, a coding tree unit included in the coded slice data, and a coding unit included in the coding tree unit, respectively.
[0032] (Coded Video Sequence) In the case of a coded video sequence, a video decoder is used to decode the sequence SEQ to be processed. The standard specifies a set of data to be referenced by the device 31. As shown in Fig. 2, the sequence SEQ includes a video parameter set, a sequence parameter set SPS (Sequence Parameter Set), a picture parameter set PPS (Picture Parameter Set), an adaptation parameter set (APS), a picture PICT, and supplemental enhancement information SEI (Supplemental Enhancement Information).
[0033] The video parameter set VPS is used to determine the number of layers of a video image. A set of coding parameters common to multiple video images and a set of coding parameters related to multiple layers and individual layers included in the video images are defined.
[0034] The sequence parameter set SPS is used to decode the target sequence. A set of coding parameters to be referenced by the PPS is specified. For example, the width and height of a picture are specified. Note that there may be multiple SPSs. In that case, one of the multiple SPSs can be selected from the PPS. Select .
[0035] The picture parameter set PPS specifies the number of pictures to be decoded for each picture in the target sequence. A set of coding parameters to be referred to by the video decoding device 31 is specified. For example, the set includes a reference value of the quantization width used in decoding a picture (pic_init_qp_minus26) and a flag indicating application of weighted prediction (weighted_pred_flag). Note that there may be multiple PPSs. In that case, one of the multiple PPSs is selected from each picture in the target sequence.
[0036] (Encoded Picture) A coded picture defines a set of data to be referenced by the video decoding device 31 in order to decode a picture PICT to be processed. As shown in FIG. 2, the picture PICT includes slices 0 to NS-1 (NS is the total number of slices included in the picture PICT).
[0037] In the following, when there is no need to distinguish between slices 0 to NS-1, the symbols are The subscripts may be omitted in the description, and the same applies to other data to which subscripts are added that are included in the coded stream Te described below.
[0038] (Coded Slice) In the case of the coded slice, the video decoding device 31 refers to the slice S to be processed in order to decode the slice S. As shown in Figure 2, a slice consists of a slice header, and includes slice data.
[0039] The slice header includes a group of coding parameters to be referred to by the video decoding device 31 in order to determine a decoding method for the current slice. Slice type designation information (slice_type) that designates the slice type is an example of a coding parameter included in the slice header.
[0040] The slice types that can be specified by the slice type specification information include (1) an I slice that uses only intra prediction during encoding, (2) a P slice that uses unidirectional prediction or intra prediction during encoding, and (3) a P slice that uses unidirectional prediction, bidirectional prediction, or intra prediction during encoding. Examples of inter prediction include B slices using intra prediction. Note that inter prediction is not limited to single prediction and bi-prediction, and a predicted image may be generated using more reference pictures. When called a B slice, it is a slice that includes blocks that can use inter prediction. This refers to the s.
[0041] In addition, the slice header may include a reference to a picture parameter set PPS (pic_parameter_set_id).
[0042] (Encoded slice data) The coded slice data specifies a set of data to be referenced by the video decoding device 31 in order to decode the slice data to be processed. As shown in Fig. 2(d), the slice data includes a CTU. A CTU is a block of a fixed size (e.g., 64x64) that constitutes a slice, and is also called a Largest Coding Unit (LCU).
[0043] (coding tree unit) 2 specifies a set of data to be referenced by the video decoding device 31 in order to decode a CTU to be processed. The CTU is encoded by recursive quad tree (QT) division, binary tree (BT) division, or ternary tree (TT) division. The image is divided into coding units (CU), which are the basic units of processing. BT division and TT division are collectively called multi-tree division (MT (Multi Tree) division). A node in the tree structure obtained by recursive quadtree division is called a coding node. The intermediate nodes of the tree are coding nodes, and the CTU itself is also defined as the top coding node. The lowest coding node is defined as a coding unit.
[0044] (Encoding Unit) FIG. 2 shows data to be referenced by the video decoding device 31 in order to decode the coding unit to be processed. Specifically, a CU consists of a CU header CUH, prediction parameters, and conversion parameters. The CU header includes information such as a prediction mode, a quantized transformation coefficient, and the like.
[0045] The prediction process may be performed on a CU basis, or on a sub-CU basis by further dividing a CU. If the size of a CU and a sub-CU are the same, there is one sub-CU in the CU. If the size of a CU is larger than that of a sub-CU, the CU is divided into sub-CUs. For example, if the CU is 8x8 and the sub-CU is 4x4, the CU is divided into 2 parts horizontally and 2 parts vertically, into 4 sub-CUs.
[0046] Prediction types (prediction modes) include intra prediction (MODE_INTRA), inter prediction (MODE_INTER), and intra block copy (MODE_IBC). Intra prediction is a prediction within the same picture, and inter prediction refers to a prediction process performed between different pictures (for example, between display times or between layer images).
[0047] Transformation and quantization are performed in units of CUs, but quantization coefficients are stored in units of subblocks such as 4x4. It may be entropy coded.
[0048] (Prediction parameters) The predicted image is derived from prediction parameters associated with the block, which include intra-prediction and inter-prediction parameters.
[0049] (Inter prediction parameters) The prediction parameters of inter prediction will be described. The inter prediction parameters are composed of prediction list use flags predFlagL0 and predFlagL1, reference picture indexes refIdxL0 and refIdxL1, and motion vectors mvL0 and mvL1. predFlagL0 and predFlagL1 are flags indicating whether or not a reference picture list (L0 list, L1 list) is used, and when the value is 1, the corresponding reference picture list is used. Note that in this specification, when the "flag indicating whether or not XX" is written, a flag other than 0 (for example, 1) is XX, and 0 is not XX, and in logical negation, logical product, etc., 1 is treated as true and 0 is treated as false (similar below). However, in an actual device or method, other values can also be used as true and false values.
[0050] Syntax elements for deriving inter prediction parameters include, for example, a merge flag merge_flag (general_merge_flag), a merge index merge_idx, merge_subblock_flag, regulare_merge_flag, ciip_flag, merge_gpm_partition_idx, merge_gpm_idx0, merge_gpm_idx1, inter_pred_idc, a reference picture index refIdxLX, mvp_LX_idx, a difference vector mvdLX, and a motion vector precision mode amvr_mode. merge_subblock_flag is a flag indicating whether to use inter prediction in subblock units. regulare_merge_flag is a flag indicating whether to use a normal merge mode or MMVD. ciip_flag is a flag indicating whether to use a CIIP (combined inter-picture merge and intra-picture prediction) mode. merge_gpm_partition_idx is an index indicating a partition shape of the GPM mode. merge_gpm_idx0 and merge_gpm_idx1 are indices that indicate merge indices in the GPM mode. inter_pred_idc is an inter prediction identifier for selecting a reference picture to be used in the AMVP mode. mvp_LX_idx is a predicted vector index for deriving a motion vector.
[0051] (Reference picture list) The reference picture list is a list of reference pictures stored in the reference picture memory 306. Each CU can select any picture in the reference picture list RefPicListX (X=0 or 1). Specify whether to actually refer to the L0 list or L1 list with refIdxLX. Note that LX is a description method used when there is no distinction between L0 prediction and L1 prediction. In the following, LX will be replaced with L0 and L1 to distinguish between parameters for the L0 list and parameters for the L1 list.
[0052] (Merge prediction and AMVP prediction) The prediction parameter decoding (encoding) method can be merged prediction mode. Adaptive Motion Vector Prediction (AMVP) mode and general_merge_flag is a flag for identifying these. The merge mode is a prediction mode in which some or all of the motion vector difference is omitted, and the prediction list usage flag predFlagLX, the reference picture index refIdxLX, and the motion vector mvLX are not included in the encoded data, but are derived from the prediction parameters of the neighboring blocks that have already been processed. The AMVP mode is a mode in which inter_pred_idc, refIdxLX, and mvLX are included in the encoded data. Note that mvLX is encoded as mvp_LX_idx, which identifies the prediction vector mvpLX, and the difference vector mvdLX. The general name for the prediction modes in which the motion vector difference is omitted or simplified is called the general merge mode, and the general merge mode and the AMVP prediction may be selected by the general_merge_flag.
[0053] If general_merge_flag is 1, regular_merge_flag may be transmitted separately. If regular_merge_flag is 1, normal merge mode or MMVD may be selected, otherwise CIIP mode or GPM mode may be selected. In CIIP mode, a predicted image is generated by a weighted sum of an inter predicted image and an intra predicted image. In GPM mode, a predicted image is generated by dividing a target CU into two non-rectangular prediction units by a line segment.
[0054] The inter_pred_idc is a value indicating the type and number of reference pictures, and takes one of the values PRED_L0, PRED_L1, and PRED_BI. PRED_L0 and PRED_L1 are managed by the L0 list and the L1 list, respectively. PRED_BI is managed by the L0 list and the L1 list. 1 shows bi-prediction using two reference pictures.
[0055] merge_idx is the prediction parameter candidate (merge candidate) derived from the processed block. (comment) is to be used as a prediction parameter for the current block.
[0056] (Motion Vector) mvLX indicates the amount of shift between blocks on two different pictures. A prediction vector and a difference vector related to mvLX are called mvpLX and mvdLX, respectively.
[0057] (Inter prediction identifier inter_pred_idc and prediction list usage flag predFlagLX) The relationship between inter_pred_idc, predFlagL0, and predFlagL1 is as follows, and they are mutually convertible.
[0058] inter_pred_idc = (predFlagL1<<1)+predFlagL0 predFlagL0 = inter_pred_idc & 1 predFlagL1 = inter_pred_idc >> 1 In addition, the inter prediction parameter may use a prediction list usage flag or an inter prediction identifier. Furthermore, the determination using the prediction list usage flag may be replaced with a determination using the inter prediction identifier. Conversely, the determination using the inter prediction identifier may be replaced with a determination using the prediction list usage flag.
[0059] (Bi-predictive biPred decision) The flag biPred indicating whether or not the prediction is bi-predictive can be derived based on whether or not two prediction list usage flags are both 1. For example, the flag can be derived by the following formula.
[0060] biPred = (predFlagL0==1 && predFlagL1==1) Alternatively, biPred can be derived based on whether the inter prediction identifier is a value indicating the use of two prediction lists (reference pictures). For example, biPred can be derived using the following formula:
[0061] biPred = (inter_pred_idc==PRED_BI) ? 1 : 0 (Configuration of a video decoding device) The configuration of a video decoding device 31 (FIG. 3) according to this embodiment will be described.
[0062] The video decoding device 31 includes an entropy decoding unit 301, a parameter decoding unit (prediction image decoding device ) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a prediction image generation unit (prediction image generation device) 308, an inverse quantization and inverse transformation unit 311, and an addition unit 312, a prediction parameter The meter derivation unit 320 is also included. The image decoding device 31 may also be configured without including the loop filter 305 .
[0063] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit. The CU decoding unit 3022 further includes a TU decoding unit 3024. These may be collectively referred to as a decoding module. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, PPS, and APS, and slice header (slice information) from the encoded data. The CT information decoding unit 3021 decodes the CT from the encoded data. The CU decoding unit 3022 decodes the CU from the encoded data. The TU decoding unit 3024 decodes the CU from the encoded data.
[0064] When a prediction error is included in a TU, the TU decoding unit 3024 decodes the QP update information and the quantized transform coefficient from the encoded data. The QP update information is a difference value from the quantization parameter predicted value qPpred, which is a predicted value of the quantization parameter QP.
[0065] The predicted image generating unit 308 generates an inter-predicted image and an intra-predicted image. The device 300 includes a unit 310.
[0066] The prediction parameter derivation unit 320 is a unit including the inter prediction parameter derivation unit 303 (FIG. 5) and the intra prediction parameter derivation unit 304 (FIG. 6). It includes a prediction parameter derivation unit.
[0067] In the following, we will describe an example in which CTU and CU are used as processing units, but this is not limiting. Alternatively, the processing may be performed in units of sub-CUs. and processing may be performed in units of blocks or sub-blocks.
[0068] The entropy decoding unit 301 performs entropy decoding on the encoded stream Te input from the outside. Entropy coding is performed to decode individual codes (syntax elements). There are two types of entropy coding: one is to use a context (probability model) that is adaptively selected according to the type of syntax element and the surrounding circumstances to code the syntax elements into variable-length codes, and the other is to use a predetermined table or formula to code the syntax elements into variable-length codes.
[0069] The entropy decoding unit 301 outputs the decoded code to the parameter decoding unit 302. The decoded code is, for example, a prediction mode predMode, general_merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX, amvr_mode, etc. Control of which code to decode is performed based on an instruction from the parameter decoding unit 302.
[0070] (Basic flow) FIG. 6 is a flowchart illustrating a schematic operation of the video decoding device 31.
[0071] (S1100: Decode Parameter Set Information) The header decoding unit 3020 decodes parameter set information such as the VPS, SPS, and PPS from the encoded data.
[0072] (S1200: Decode slice information) The header decoding unit 3020 decodes slice header information from the encoded data. Decode (slice information).
[0073] Hereinafter, the video decoding device 31 performs steps S1300 to S5000 for each CTU included in the target picture. By repeating the above process, a decoded image of each CTU is derived.
[0074] (S1300: Decode CTU Information) The CT information decoding unit 3021 decodes the CTU from the encoded data.
[0075] (S1400: Decode CT Information) The CT information decoding unit 3021 decodes the CT from the encoded data.
[0076] (S1500: CU Decoding) The CU decoding unit 3022 performs S1510 and S1520 to decode the CU from the encoded data. Issued.
[0077] (S1510: Decode CU information) The CU decoding unit 3022 decodes CU information, prediction information, and TU division from the encoded data. flag, CU residual flag, etc.
[0078] (S1520: TU information decoding) When a TU includes a prediction error, the TU decoding unit 3024 decodes the encoded TU information. The quantized prediction error, etc. is decoded from the data.
[0079] (S2000: Generation of predicted image) The predicted image generation unit 308 generates a predicted image for each block included in the current CU based on the prediction information.
[0080] (S3000: Inverse quantization and inverse transformation) The inverse quantization and inverse transformation unit 311 performs inverse quantization and inverse transformation for each TU included in the target CU. Then, inverse quantization and inverse transform processing is performed.
[0081] (S4000: Decoded image generation) The adder 312 receives a predicted image from the predicted image generator 308 and The prediction error supplied from the inverse quantization and inverse transform unit 311 is added to the A decoded image is generated.
[0082] (S5000: Loop Filter) The loop filter 305 applies a loop filter such as a deblocking filter, SAO, or ALF to the decoded image to generate a decoded image.
[0083] The loop filter 305 is a filter provided in the coding loop, and is used to remove block distortion and ringing. The loop filter 305 is a filter that removes distortion and improves image quality. The loop filter 305 applies a deblocking filter, a sample adaptive offset (SAO), and an adaptive filter to the decoded image of the CU generated by the adder 312. Apply a filter such as an automatic loop filter (ALF).
[0084] The reference picture memory 306 stores the decoded image of the CU generated by the adder 312 in a location that is determined in advance for each current picture and current CU.
[0085] The prediction parameter memory 307 stores prediction parameters at a predetermined position for each CTU or CU to be decoded. Specifically, the prediction parameter memory 307 stores the parameters decoded by the parameter decoding unit 302 and the prediction mode predMode separated by the entropy decoding unit 301. do.
[0086] The predicted image generating unit 308 receives predMode, prediction parameters, etc. The generation unit 308 reads the reference picture from the reference picture memory 306. generates a predicted image of a block or subblock using prediction parameters and the read reference picture (reference block) in the prediction mode indicated by predMode. Here, the reference block is a set of pixels on the reference picture (usually rectangular, so called a block), and is an area to be referenced to generate a predicted image.
[0087] (Configuration of inter-prediction parameter derivation unit) As shown in FIG. 5, the inter-prediction parameter derivation unit 303 receives the input from the parameter decoding unit 302. Based on the input syntax element, the prediction parameters stored in the prediction parameter memory 307 are The inter prediction parameter derivation unit 303 derives inter prediction parameters by referring to the motion vector derivation unit 303. The inter prediction parameter derivation unit 303 also outputs the inter prediction parameters to the inter prediction image generation unit 309 and the prediction parameter memory 307. The inter prediction parameter derivation unit 303 and its internal elements, the AMVP prediction parameter derivation unit 3032, the merge prediction parameter derivation unit 3036, the MMVD prediction unit 30376, and the MV addition unit 3038, are means common to the video encoding device and the video decoding device, and therefore may be collectively referred to as a motion vector derivation unit (motion vector derivation device).
[0088] If general_merge_flag is 1, i.e., indicates merge prediction mode, derive merge_idx. and outputs it to the merge prediction parameter derivation unit 3036.
[0089] When general_merge_flag is 0, that is, when it indicates the AMVP prediction mode, the AMVP prediction parameter derivation unit 3032 derives mvpLX from inter_pred_idc, refIdxLX, or mvp_LX_idx.
[0090] (MV addition section) The MV adder 3038 adds the derived mvpLX and mvdLX to derive mvLX.
[0091] (Merge prediction) The merge prediction parameter derivation unit 3036 calculates the prediction parameters (predFlagLX, mvLX, refIdxLX) The merge prediction parameter derivation unit 3036 derives merge candidates including the merge_idx and the merge_idx of the merge candidates included in the merge candidate list. The motion information (mvLXN[0], mvLXN[1]), predFlagLXN, and refIdxLXN of the merge candidate N indicated by The merge prediction parameter derivation unit 3036 stores the inter prediction parameters of the selected merge candidates in the prediction parameter memory 307. Both are output to the inter-prediction image generation unit 309.
[0092] (Inter-prediction image generation unit 309) When the predMode indicates inter prediction, the inter prediction image generation unit 309 The inter prediction parameter derivation unit 303 calculates the inter prediction parameter and the reference picture. A predicted image of a block or sub-block is generated by super-prediction.
[0093] FIG. 4 is a block diagram of the inter-prediction image generating unit 309 included in the prediction image generating unit 308 according to this embodiment. The inter-prediction image generating unit 309 is a motion compensation unit (prediction image generating device). The synthesizer 3095 includes a weighted prediction unit 3094.
[0094] (Motion Compensation) The motion compensation unit 3091 (interpolated image generation unit 3091) receives the input from the inter prediction parameter derivation unit 303. Based on the input inter prediction parameters (predFlagLX, refIdxLX, mvLX), an interpolated image (motion compensated image) is generated by reading a reference block from the reference picture memory 306. The reference block is a block shifted by mvLX from the position of the target block on the reference picture RefPicLX specified by refIdxLX. If mvLX does not have integer precision, an interpolated image is generated by applying a filter called a motion compensation filter for generating pixels at decimal positions.
[0095] The motion compensation unit 3091 derives the integer position (xInt, yInt) and phase (xFrac, yFrac) corresponding to the top left coordinates (xPb, yPb) of a block of size bW*bH, the coordinates within the prediction block (xL, yL), and the motion vector (mvLX[0], mvLX[1]) using the following formula (MC-P1).
[0096] xInt = xPb+(mvLX[0]>>(log2(MVPREC)))+xL xFrac = mvLX[0]&(MVPREC-1) yInt = yPb+(mvLX[1]>>(log2(MVPREC)))+yL yFrac = mvLX[1]&(MVPREC-1) Here, MVPREC indicates the accuracy of mvLX (1 / MVPREC pixel accuracy), log2MVPREC=(log2(MVPREC)), x=0...bW-1, y=0...bH-1. For example, MVPREC=16. In order to perform RPR, the values may be derived as in (MC-P2) described later. Furthermore, for (xInt, yInt) derived in (MC-P1) and (MC-P2), the motion compensation unit 3091 may perform position adjustment for wraparound. The position may be modified. If the flag to treat subpicture boundaries as picture boundaries is enabled (sps_subpic_treated_as_pic_flag == 1) and the number of subpictures for the reference picture is greater than 1 (sps_num_subpics_minus1 > 0 for reference picture refPicLX) xInt = Clip3( SubpicLeftBoundaryPos, SubpicRightBoundaryPos, refWraparoundEnabledFlag ? ClipH( ( PpsRefWraparoundOffset ) * MinCbSizeY, picW, xInt ) : xInt ) yInt = Clip3( SubpicTopBoundaryPos, SubpicBotBoundaryPos, yInt ) Here, SubpicLeftBoundaryPos, SubpicRightBoundaryPos, SubpicTopBoundaryPos, and SubpicBottomBoundaryPos are the left, right, top, and bottom boundary positions of the subpicture, respectively. In all other cases (when the flag to treat subpicture boundaries as picture boundaries is disabled (sps_subpic_treated_as_pic_flag == 0) or when the number of subpictures in the reference picture is 1 (sps_num_subpics_minus1 for reference picture refPicLX == 0) xInt = Clip3( 0, picW ? 1, refWraparoundEnabledFlag ? ClipH( ( PpsRefWraparoundOffset ) * MinCbSizeY, picW, xInt ) : xInt ) yInt = Clip3( 0, picH ? 1, yInt ) Here, refWraparoundEnabledFlag = pps_ref_wraparound_enabled_flag && !refPicIsScaled, PpsRefWraparoundOffset = pps_pic_width_in_luma_samples / MinCbSizeY ? pps_pic_width_minus_wraparound_offset, where MinCbSizeY is a predetermined constant or variable (e.g., 4), and pps_pic_width_minus_wraparound_offset is an offset decoded from the encoded data indicating the position of the wraparound.
[0097] The motion compensation unit 3091 derives a temporary image temp[][] by performing horizontal interpolation processing on the reference picture refImg using an interpolation filter. In the following, Σ is the sum for k=0..NTAP-1, mcFilter[Frac][k] is the k-th interpolation filter coefficient in phase Frac, shift1 is a normalization parameter that adjusts the value range, and offset1=1<<(shift1-1).
[0098] temp[x][y] = (ΣmcFilter[xFrac][k]*refImg[xInt+k-NTAP / 2+1][yInt]+offset1)>>shift1 Next, the motion compensation unit 3091 derives the interpolated image Pred[][] by vertically interpolating the temporary image temp[][]. In the following, Σ is the sum for k=0..NTAP-1, shift2 is a normalization parameter that adjusts the value range, and offset2=1<<(shift2-1).
[0099] Pred[x][y] = (ΣmcFilter[yFrac][k]*temp[x][y+k-NTAP / 2+1]+offset2)>>shift2 (formula MC-1) (Out of Picture Range (OOB) Processing) The OOB processing unit 3092 includes an OOB determination unit 30921 and an OOB mask derivation unit 30922. Based on the inter prediction parameters (predFlagLX, refIdxLX, mvLX) input from the inter prediction parameter derivation unit 303, determines whether each reference block in the bi-prediction mode includes an area outside the range of the reference picture (OOB: Out-Of-Boundary).
[0100] The following description will be given with reference to FIG. When the OOB determination unit 30921 determines that some pixels of the reference block are outside the picture, that is, the reference block is OOB (isOOB==true, S1002), the OOB mask derivation unit 30922 Derive mask data (mask value, OOB mask) that indicates which areas of the lock are out of range The inter predicted image generating unit 309 applies an OOB mask to the reference block that is OOB to generate a predicted image (S1006). A normal predicted image is generated for the reference block (S1008). This will be described in detail below.
[0101] (OOB detection process) The OOB determination unit 30921 determines whether or not each of the two (first and second) reference blocks is out of picture (OOB) in bi-prediction. The area of the LX reference block (X=0 or 1) corresponding to the reference image refLX is represented by the block top left coordinates (xRefLX, yRefLX) and the block size (xRefLX, yRefLX). The size of the picture is expressed as the width and height (bW, bH). Let the width and height of the picture be picW and picH, respectively. In the following, the coordinate system uses an integer value of mvLX according to the precision MVPREC of mvLX. Using these values, the OOB determination unit 30921 derives a truth value isOOB[X] indicating whether the LX reference block is outside the picture by the following formula.
[0102] An example is shown in FIG. 10. FIG. 10 is a diagram in integer units, and the motion vector mvLX' and the reference pixel position (xRefLX', yRefLX') are parameters obtained by expressing the motion vector mvLX and the reference pixel position (xRefLX, yRefLX) in integer form. xRefLX = (xPb << log2MVPREC) + mvLX[0] yRefLX = (yPb << log2MVPREC) + mvLX[1] picMinX = 0 - MVPREC / 2 picMinY = 0 - MVPREC / 2 picMaxX = ((picW - 1) << log2MVPREC) + MVPREC / 2 picMaxY = ((picH - 1) << log2MVPREC) + MVPREC / 2 isOOB[X] = (xRefLX <= picMinX) || (yRefLX <= picMinY) || (xRefLX + (bW << log2MVPREC) - 1) >= picMaxX) || (yRefLX + (bH << log2MVPREC) - 1) >= picMaxY) The log2MVPREC-bit left shift ("<< log2MVPREC") can be processed by multiplication with MVPREC ("*MVPREC"), or the subsequent "*MVPREC" can be processed by "<< log2MVPREC". <Configuration Example 1> The OOB determination unit 30921 in this configuration example determines whether the target block is a target for OOB processing, including determining whether wraparound processing is applied to the target picture. Similarly, it may be determined whether the reference block is a target for OOB processing by using the presence or absence of wraparound processing applied to the reference picture containing the reference block. The syntax element pps_ref_wraparound_enabled_flag notified in the PPS is for horizontal wraparound ... A flag indicating whether wraparound motion compensation is available. If pps_ref_wraparound_enabled_flag is true (1), it indicates that reference picture wraparound processing is applied (enabled) to the current picture. When wraparound processing is enabled, the OOB determination unit 30921 sets isOOB[X]=false. if (pps_ref_wraparound_enabled_flag == true) { isOOB[X] = false } else { xRefLX = (xPb< <log2MVPREC) + mvLX[0] yRefLX = (yPb< <log2MVPREC) + mvLX[1] picMinX = 0 - MVPREC / 2 picMinY = 0 - MVPREC / 2 picMaxX = ((picW-1)< <log2MVPREC) + MVPREC / 2 picMaxY = ((picH-1)< <log2MVPREC) + MVPREC / 2 isOOB[X] = (xRefLX <= picMinX) || (yRefLX <= picMinY) || (xRefLX + (bW<<log2MVPREC) -1) > = picMaxX) || (yRefLX + (bH<<log2MVPREC) -1) > = picMaxY) } Alternatively, the determination may be made by a logical expression of the OOB determination process, rather than by conditional branching.
[0103] isOOB[X] = (pps_ref_wraparound_enabled_flag == false) && ((xRefLX <= picMinX) || (yRefLX <= picMinY) || (xRefLX + (bW<<log2MVPREC) -1) > = picMaxX) || (yRefLX + (bH<<log2MVPREC) -1) > = picMaxY) In the above example, pps_ref_wraparound_enabled_flag is used to check whether wraparound processing is enabled or disabled. However, the present invention is not limited to this. The value of another variable or function derived based on the presence or absence of wraparound processing in the target picture, for example, refWraparoundEnabledFlag, may be used (the same applies to the following configuration).
[0104] According to the above configuration, the OOB process prevents the coding efficiency from being reduced due to an inappropriate wraparound process. This has the effect of suppressing an increase in the amount of calculation required for generating a predicted image.
[0105] <Configuration example 2> The OOB determination unit 30921 may change the OOB determination method depending on whether the wraparound process is applied to the target picture. The presence or absence of wraparound processing applied to the reference block is used to determine whether the reference block is an OOB processing target. For example, the wraparound process in VVC is Therefore, when wraparound processing is enabled, isOOB[X] is derived by comparing only the vertical coordinate, which is not subject to wraparound processing, as shown below. On the other hand, when wraparound processing is disabled, isOOB[X] is derived by comparing both the vertical and horizontal coordinates (X=0,1). xRefLX = (xPb< <log2MVPREC) + mvLX[0] yRefLX = (yPb< <log2MVPREC) + mvLX[1] picMinX = 0 - MVPREC / 2 picMinY = 0 - MVPREC / 2 picMaxX = ((picW-1)< <log2MVPREC) + MVPREC / 2 picMaxY = ((picH-1)< <log2MVPREC) + MVPREC / 2 if (pps_ref_wraparound_enabled_flag == true) { isOOB[X] = (yRefLX <= picMinY) || (yRefLX + (bH<<log2MVPREC) -1) > = picMaxY) } else { isOOB[X] = (xRefLX <= picMinX) || (yRefLX <= picMinY) || (xRefLX + (bW<<log2MVPREC) -1) > = picMaxX) || (yRefLX + (bH<<log2MVPREC) -1) > = picMaxY) } Alternatively, the determination may be made by a logical expression of the OOB determination process, rather than by conditional branching. isOOB[X] = ((pps_ref_wraparound_enabled_flag == false) && ((xRefLX <= picMinX) || (xRefLX + (bW<<log2MVPREC) -1) > = picMaxX))) || (yRefLX <= picMinY) || (yRefLX + (bH<<log2MVPREC) -1) > = picMaxY) According to the above configuration, even when the wraparound process is enabled, the OOB process is also used, improving the coding efficiency. This has the effect of making the above possible.
[0106] <Configuration example 3> In this configuration example, the OOB determination unit 30921 applies reference picture resampling (RPR) processing. For reference pictures that have been converted, the coordinates are corrected with a specified scale before the OOB determination process is performed. Specifically, the coordinates are corrected as follows:
[0107] Instead of the above-mentioned (MC-P1), the motion compensation unit 3091 derives the luminance coordinates (refxSbL, refySbL) and (refxL, refyL) indicated by the motion vector refMvLX as follows (MC-P2): refxSbL = ( ( ( xPb ? ( SubWidthC * pps_scaling_win_left_offset ) ) << log2MVPREC ) + mvLX[0] ) * scalingRatio[0] refxL = ( ( sign( refxSbL ) * ( ( abs( refxSbL ) + offset1 ) >> shiftP1 ) + xL * ( ( scalingRatio[0] + offsetP2 ) >> shiftP2 ) ) + fRefLeftOffset + offsetP3 ) >> shift3 refySbL = ( ( ( yPb ? ( SubHeightC * pps_scaling_win_top_offset ) ) << log2MVPREC ) + mvLX[1] ) * scalingRatio[1] refyL = ( ( sign( refySbL ) * ( ( abs( refySbL ) + offset1 ) >> shift1 ) + yL * ( ( scalingRatio[1] + offsetP ) >> shiftP ) ) + fRefTopOffset + offsetP3 ) >> shiftP3 Here, offsetP1=(1<<(shiftP1-1), offsetP2=(1<<(shift2-1), offset3=(1<<(shift3-1). For example, shiftP1=8, shiftP3=4, shiftP3=6, offsetP1, offsetP2, offsetP3=128, 8 , 6 xInt = refxL>>log2MVPREC xFrac = refxL &(MVPREC-1) yInt = refyL >>log2MVPREC yFrac = refyL &(MVPREC-1) Here, SubWidthC and SubHeightC are determined by the sampling method of the color difference format. It may be derived as follows by decoding the syntax element sps_chroma_format_idc in the encoded data: sps_chroma_format_idc is a parameter that indicates the chrominance format.
[0108] When sps_chroma_format_idc=0(Monochrome), SubWidthC=1, SubHeightC=1 When sps_chroma_format_idc=1(4:2:0), SubWidthC=2, SubHeightC=2 When sps_chroma_format_idc=2(4:2:2), SubWidthC=2, SubHeightC=1 When sps_chroma_format_idc=3(4:4:4), SubWidthC=1, SubHeightC=1 pps_scaling_win_left_offset, pps_scaling_win_top_offset, pps_scaling_win_right_offset, and pps_scaling_win_bottom_offset are offset values applied during scaling. refMvLX (X=0,1) is a motion vector on the reference picture, with an accuracy of MVPREC (1 / MVPREC). scalingRatio[0] and scalingRatio[1] are the horizontal and vertical scaling factors of the reference picture with respect to the target picture, respectively. scalingRatio[0] = ((fRefWidth<<14) + (CurrPicScalWinWidthL>>1)) / CurrPicScalWinWidthL scalingRatio[1] = ((fRefHeight<<14) + (CurrPicScalWinHeightL>>1)) / CurrPicScalWinHeightL CurrPicScalWinWidthL = pps_pic_width_in_luma_samples ? SubWidthC * (pps_scaling_win_right_offset + pps_scaling_win_left_offset) CurrPicScalWinHeightL = pps_pic_height_in_luma_samples ?SubHeightC * (pps_scaling_win_bottom_offset + pps_scaling_win_top_offset) fRefWidth is the CurrPicScalWinWidthL, fRefHeight of the jth reference picture in RefPicListX. is the CurrPicScalWinHeightL of the j-th reference picture in RefPicListX.
[0109] fRefLeftOffset and fRefTopOffset are values that are set as follows: fRefLeftOffset = (SubWidthC * pps_scaling_win_left_offset) << 10) fRefTopOffset = (SubHeightC * pps_scaling_win_top_offset) << 10) xL, yL are relative coordinates with the top left coordinate of the reference block being (0,0).
[0110] The OOB determination unit 30921 is a unit for determining whether a reference picture is to be scaled by the RPR process. (xRefLX, yRefLX) is derived as (refx, refy) using (xL, yL)=(0, 0) as the upper left coordinate of the reference block. xRefLX = refxL = ((sign(refxSbL) * ((abs(refxSbL) + 128) >> 8)) + fRefLeftOffset + 32) >> 6 yRefLX = refyL = ((sign(refySbL) * ((abs(refySbL) + 128) >> 8)) + fRefTopOffset + 32) >> 6 The width bW and height bH of the reference block after correcting the scaling due to the RPR process are It is derived by the formula: bW = sbWidth * ((scalingRatio[0]+8)>>4) bH = sbHeight * ((scalingRatio[1]+8)>>4) The correction method is not limited to the above formula, and any other method may be used as long as it scales the coordinate system using at least the scaling magnification (scalingRatio) of the RPR process.
[0111] The OOB determination unit 30921 uses the coordinates (xRefLX, xRefLY) corrected by scaling in this way. This is used to determine whether the reference block is OOB. picMinX = 0 - MVPREC / 2 picMinY = 0 - MVPREC / 2 picMaxX = ((picW-1)< <log2MVPREC) + MVPREC / 2 picMaxY = ((picH-1)< <log2MVPREC) + MVPREC / 2 isOOB[X] = (xRefLX <= picMinX) || (yRefLX <= picMinY) || (xRefLX + (bW<<log2MVPREC) - 1) > = picMaxX) || (yRefLY + (bH<<log2MVPREC) - 1) > = picMaxY) According to the above configuration, it is possible to simultaneously apply the RPR process and the OOB process. That is, for a reference picture to which the RPR process has been applied, the OOB determination process is performed based on the scaled coordinates, so that it is possible to improve the coding efficiency by the OOB process. .
[0112] <Configuration Example 4> In the fourth configuration example, an example will be described in which the OOB process is not performed on a reference picture to which the RPR process is applied. The OOB determination unit 30921 in this configuration example performs reference picture resampling (RPR) processing. Determines whether the target block is subject to OOB processing, including whether OOB processing is applied. Determine whether or not. if (refPicIsScaled == true) { isOOB[X] = false } else { isOOB[X] = (xRefLX <= picMinX) || (yRefLX <= picMinY) || (xRefLX + (bW<<log2MVPREC) -1) > = picMaxX) || (yRefLX + (bH<<log2MVPREC) -1) > = picMaxY) } Alternatively, the determination may be made by a logical expression of the OOB determination process, rather than by conditional branching.
[0113] isOOB[X] = !refPicIsScaled && ((xRefLX <= picMinX) || (yRefLX <= picMinY) || (xRefLX + (bW<<log2MVPREC) -1) > = picMaxX) || (yRefLX + (bH<<log2MVPREC) -1) > = picMaxY) Here, refPicIsScaled is the value when RPR processing is enabled for the reference picture and the target picture This variable is true if the size of the reference picture is different from that of the i-th picture. It can be derived as follows: refPicIsScaled[i][j] = ( pps_pic_width_in_luma_samples != refPicWidth || pps_pic_height_in_luma_samples != refPicHeight || pps_scaling_win_left_offset != refScalingWinLeftOffset || pps_scaling_win_right_offset != refScalingWinRightOffset || pps_scaling_win_top_offset != refScalingWinTopOffset || pps_scaling_win_bottom_offset != refScalingWinBottomOffset || sps_num_subpics_minus1 != fRefNumSubpics) sps_num_subpics_minus1+1 and fRefNumSubpics indicate the number of subpictures of the target picture and the number of subpictures of the reference picture, refPicWidth and refPicHeight indicate the width and height of the reference picture, and refScalingWinLeftOffset, refScalingWinLeftOffset, refScalingWinLeftOffset, and refScalingWinLeftOffset indicate the offset of the reference picture.
[0114] According to the above configuration, the OOB process is not applied to a reference picture to which the RPR process is applied, and an effect is achieved in that it is possible to reduce the amount of calculation in generating a predicted image.
[0115] <Configuration Example 5> In this configuration example, a case will be described in which both the wraparound process and the reference picture resampling process are available. The OOB determination unit 30921 in this configuration example determines whether the wraparound process is applied to the target picture, and performs reference picture resampling (RPR ) processing has been applied to the target block. if (pps_ref_wraparound_enabled_flag || refPicIsScaled) { isOOB[X] = false } else { xRefLX = (xPb< <log2MVPREC) + mvLX[0] yRefLX = (yPb< <log2MVPREC) + mvLX[1] picMinX = 0 - MVPREC / 2 picMinY = 0 - MVPREC / 2 picMaxX = ((picW-1)< <log2MVPREC) + MVPREC / 2 picMaxY = ((picH-1)< <log2MVPREC) + MVPREC / 2 isOOB[X] = (xRefLX <= picMinX) || (yRefLX <= picMinY) || (xRefLX + (bW<<log2MVPREC) -1) > = picMaxX) || (yRefLX + (bH<<log2MVPREC) -1) > = picMaxY) } Alternatively, the determination may be made by a logical expression of the OOB determination process, rather than by conditional branching.
[0116] isOOB[X] = (!refWraparoundEnabledFlag && !refPicIsScaled) && ((xRefLX <= picMinX) || (yRefLX <= picMinY) || (xRefLX + (bW<<log2MVPREC) -1) > = picMaxX) || (yRefLX + (bH<<log2MVPREC) -1) > = picMaxY) According to the above configuration, when both the wraparound process and the reference picture resampling process are available, if either one of the processes is applied, the OOB process is not performed. This has the effect of suppressing an increase in the amount of calculation required for generating a predicted image. Note that the OOB determination unit 30921 may be configured using another combination of configuration example 1 or 2 and configuration example 3 or 4, without being limited to this configuration example.
[0117] (OOB mask derivation process) The OOB mask derivation unit 30922 derives an OOB mask (OOBMask[X]) for the reference block. OOBMask[X] is a binary image (mask) of bWxbH pixels, and the value of OOBMask[X][px][py] (px=0..bW-1, py=0..bH-1) indicates that the pixel (px, py) is inside the picture if it is 0 (false), and outside the picture if it is 1 (true). Note that false and true are not limited to 0 and 1, and may be other values, such as bit mask values "0000000000b" and "1111111111b" consisting of a binary string of 0 or 1 with a length equal to or greater than BitDepth. Hexadecimal values "000", "3FF", "000", "FFF", "0000", and "FFFF" may also be used. The OOB mask derivation unit 30922 sets all values of OOBMask[X] to false for blocks that the OOB determination unit 30921 determines are not targets for OOB processing (isOOB[X] == false). In addition, for blocks that the OOB determination unit 30921 determines are targets for OOB processing (isOOB[X] == true), the OOB mask derivation unit 30922 uses the following formula to derive a mask. Note that the inequality sign (">=" or "<=") does not have to include the equality sign (">" or "<"). OOBMask[X][px][py] = ((xRefLX + px <= picMinX) || (yRefLX + py <= picMinY) || (xRefLX + px >= picMaxX) || (yRefLX + py >= picMaxY)) ? 1 : 0 where px = 0..bW-1 and py = 0..bH-1.
[0118] (OOB predicted image generation) In the case of bi-prediction, the motion compensation unit 3091 (interpolated image generation unit 3091) sets Pred[][] derived by the above (Equation MC-1) as the interpolated images PredL0[][] and PredL1[][] for each of the L0 list and the L1 list. Then, an interpolated image Pred[][] is generated from PredL0[][] and PredL1[][]. When a bi-prediction reference block is a target of OOB processing, an OOBMask[X] corresponding to the reference block is used to generate a predicted image. In this case, Pred[][] is derived as follows. if (OOBMask[0][px][py] && !OOBMask[1][px][py]) { (formula WP-1) Pred[px][py] = (PredL1[px][py] + offset1)>> shift1 } else if (!OOBMask[0][px][py] && OOBMask[1][px][py]) { Pred[px][py] = (PredL0[px][py] + offset1)>> shift1 } else { Pred[px][py] = (PredL0[px][py] + PredL1[px][py] + offset2)>> shift2 } Here, px=0..bW-1, py=0..bH-1. Also, shift1=Max(2,14-bitDepth), shift2=Max(3,15-bitDepth), offset1=1<<(shift1-1), offset2=1<<(shift2-1). is the shift (e.g. 14-bitDepth) to return the interpolated image bit depth (e.g. 14-bit) to the original bit depth. shift2 is the right shift amount that returns the image to the original bit depth and also serves as the average. offset1 and offset2 are the rounding values when shifting right by shift1 and shift2.
[0119] In addition, CU-based explicit weighted bi-prediction (BCW) or intra-inter joint prediction (CIIP) Depending on whether or not you use (combined inter-picture merge and intra-picture prediction), The above processing may be switched depending on the situation. bcwIndex is an index indicating the weight of the BCW. ciip_flag is a flag indicating whether CIIP is on or off. When bcwIndex==0 or the CIIP flag (ciip_flag) is 1, the above processing (formula WP-1) is performed. In all other cases, the following processing is performed. w1 = bcwWLut[bcwIdx] w0 = (8 - w1) where bcwWLut[k] = {4, 5, 3, 10, -2} if (OOBMask[0][px][py] && !OOBMask[1][px][py]) { Pred[px][py] = (PredL1[px][py] + offset1)>> shift1 } else if (!OOBMask[0][px][py] && OOBMask[1][px][py]) { Pred[px][py] = (PredL0[px][py] + offset1)>> shift1 } else { Pred[px][py] = ( w0 * PredL0[px][py] + w1 * PredL1[px][py] + offset3 )>>(shift1+3)) } Here, offset3=1<<(shift1+2).
[0120] The flag predFlagLX indicates whether to use the LX reference picture on a block-by-block basis, and the OOBMask By combining the above processes into one branch and processing as follows, the amount of processing can be reduced. if ((predFlagL0 == 1 && predFlagL1 ==0) || (!OOBMask[0][px][py] && OOBMask[1][px][py])) { (formula WP-2) Pred[px][py] = (PredL0[px][py]*we0 + offset1)>> shift1 } else if ((predFlagL0 == 0 && predFlagL1 ==1) || OOBMask[0][px][py] && !OOBMask[1][px][py]) { Pred[px][py] = (PredL1[px][py]*we1 + offset1)>> shift1 } else if ((bcwIdx ==0 || ciip_flag == 1) { Pred[px][py] = (w0 * PredL0[px][py] + w1 * PredL1[px][py] + offset2)>>shift2 } else { / / bcwIdx !=0 and ciip_flag == 0) Pred[px][py] = (we0 * PredL0[px][py] + we1 * PredL1[px][py] + offset3)>>(shift1+3)) } <Exclusive configuration> When using the above exclusive configuration, both OOBMask[0] and OOBMask[1] will never be 1. Therefore, the motion compensation unit 3091 (interpolated image generation unit 3091) can derive an interpolated image by the following branching process. if (OOBMask[0][px][py]) { Pred[px][py] = (PredL1[px][py] + offset1)>> shift1 } else if (OOBMask[1][px][py]) { Pred[px][py] = (PredL0[px][py] + offset1)>> shift1 } else { Pred[px][py] = (PredL0[px][py] + PredL1[px][py] + offset2)>> shift2 } It is also appropriate to calculate PredL0, PredL1, and PredBI in advance and then derive the final Pred by the following mask processing. First, PredL0, PredL1, and PredBI are calculated. PredL0[px][py] = (PredL0[px][py] + offset1) >> shift1 PredL1[px][py] = (PredL1[px][py] + offset1) >> shift1 PredBI[px][py] = (PredL0[px][py] + PredL1[px][py] + offset2) >> shift2 Next, if bit masks of binary "1111111111b" and "0000000000b" with a BitDepth length are used as the values of OOBMask[0][px][py] and OOBMask[1][px][py], the final Pred is derived using the following process. Pred[px][py]= (PredBI[px][py] & (!OOBMask[0][px][py] & !OOBMask[1][px][py])) + (PredL0[px][py] & OOBMask[0][px][py]) + (PredL1[px][py] & OOBMask[1][px][py]) In addition, if the values are true=1 and false=0, the bit mask values "0000000000b" and "1111111111b" can be generated from 0 and 1 and the following processing can be performed. Here, "+" can be replaced with "|" (the following masks The same applies to calculation processing.) Pred[px][py]= (formula excl-1) (PredBI[px][py] & (0-((!OOBMask[0][px][py] & !OOBMask[1][px][py]) ? 1 : 0)))+ (PredL0[px][py] & (0-(OOBMask[0][px][py] ? 1 : 0))) + (PredL1[px][py] & (0-(OOBMask[1][px][py] ? 1 : 0))) The above formula includes a calculation to convert the binary (0 or 1) mask to 0 (all bits are 0, pixel value is invalid) or -1 (all bits are 1, pixel value is valid) in order to create a mask with a bit width equal to or greater than the bit depth of the pixel value. The derivation of the predicted image is not limited to this, and any calculation can be used in which the terms that are invalid due to the OOB mask among the three terms corresponding to PredBI, PredL0, and PredL1 become 0. For example, if the pixel value bit depth is set to a maximum of 16 bits, the following formula may be used: Pred[px][py]= (PredBI[px][py] & ((!OOBMask[0][px][py] & !OOBMask[1][px][py])? 0xFFFF : 0))+ (PredL0[px][py] & (OOBMask[0][px][py] ? 0xFFFF : 0)) + (PredL1[px][py] & (OOBMask[1][px][py] ? 0xFFFF : 0)) Instead of using addition to combine the three terms, we can use bitwise OR as shown below. Pred[px][py]= (PredBI[px][py] & ((!OOBMask[0][px][py] & !OOBMask[1][px][py])? 0xFFFF : 0)) | (PredL0[px][py] & (OOBMask[0][px][py] ? 0xFFFF : 0)) | (PredL1[px][py] & (OOBMask[1][px][py] ? 0xFFFF : 0)) Alternatively, multiplication may be used. Pred[px][py]= (PredBI[px][py] * ((!OOBMask[0][px][py] & !OOBMask[1][px][py])? 1 : 0)) + (PredL0[px][py] * (OOBMask[0][px][py] ? 1 : 0)) + (PredL1[px][py] * (OOBMask[1][px][py] ? 1 : 0)) According to the above, the exclusive process can be simplified. The exclusive process can be simpler than the calculation of (expression excl-1).
[0121] The motion compensation unit 3091 (interpolated image generation unit 3091) may refer to the value of isOOB and perform separate processing in advance. For example, you can do the following. If you do this, when isOOB[X] is false, This has the advantage that there is no need to refer to OOBMask[X][px][py] for each pixel, thereby reducing the amount of calculations. if (isOOB[0] && !isOOB[1]) { if (OOBMask[0][px][py]) { Pred[px][py] = (PredL1[px][py] + offset1)>> shift2 } else { Pred[px][py] = (PredL0[px][py] + PredL1[px][py] + offset2)>> shift } } else if (!isOOB[0] && isOOB[1]) { if (OOBMask[1][px][py]) { Pred[px][py] = (PredL0[px][py] + offset1)>> shift2 } else { Pred[px][py] = (PredL0[px][py] + PredL1[px][py] + offset2)>> shift } } else if (isOOB[0] && isOOB[1]) { if (OOBMask[0][px][py] &&!OOBMask[1][px][py]) { Pred[px][py] = (PredL1[px][py] + offset1)>> shift2 } else if (!OOBMask[0][px][py] && OOBMask[1][px][py]) { Pred[px][py] = (PredL0[px][py] + offset1)>> shift2 } else { Pred[px][py] = (PredL0[px][py] + PredL1[px][py] + offset2)>> shift } } else { Pred[px][py] = (PredL0[px][py] + PredL1[px][py] + offset2)>> shift } (Weight Prediction) The weighted prediction unit 3094 generates a predicted image of the block by multiplying the interpolated image PredLX by a weighting coefficient.
[0122] The inter-prediction image generation unit 309 outputs the generated prediction image of the block to the addition unit 312 .
[0123] (Configuration of a video encoding device) Next, a configuration of the video encoding device 11 according to this embodiment will be described. 1 is a block diagram showing a configuration of a video encoding device 11 according to the present embodiment. The video encoding device 11 includes a prediction image generating unit 101, a subtraction unit 102, a transformation and quantization unit 103, an inverse quantization and inverse transformation unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, an encoding parameter determining unit 110, a parameter encoding unit 111, a prediction parameter derivation unit 120, and an entropy encoding unit 104.
[0124] The predicted image generation unit 101 generates a predicted image for each CU. The predicted image generation unit 101 includes the inter predicted image generation unit 309 and the intra predicted image generation unit already described, and therefore a description thereof will be omitted.
[0125] The subtraction unit 102 generates a prediction error by subtracting the pixel values of the predicted image of the block input from the predicted image generation unit 101 from the pixel values of the image T. The subtraction unit 102 outputs the prediction error to the transformation and quantization unit 103.
[0126] The transform / quantization unit 103 calculates transform coefficients by frequency transforming the prediction error input from the subtraction unit 102, and derives quantized transform coefficients by quantizing the prediction error. The quantized transform coefficients are output to the parameter coding unit 111 and the inverse quantization and inverse transform unit 105 .
[0127] The transform / quantization unit 103 includes a separation transform unit (first transform unit) and a non-separation transform unit (second transform unit). and a scaling unit.
[0128] The separate transform unit applies a separate transform to the prediction error, and the scaling unit scales the transform coefficients with a quantization matrix.
[0129] The inverse quantization and inverse transform unit 105 is the same as the inverse quantization and inverse transform unit 311 in the video decoding device 31, and a description thereof will be omitted.
[0130] The parameter coding unit 111 includes a header coding unit 1110, a CT information coding unit 1111, and a CU coding unit 1112 (prediction mode coding unit). The CU coding unit 1112 further includes a TU coding unit 1114. The following describes an outline of the operation of each module.
[0131] The header encoding unit 1110 performs encoding processing of parameters such as header information, division information, prediction information, and quantized transform coefficients.
[0132] The CT information encoding unit 1111 encodes the QT, MT (BT, TT) division information and the like.
[0133] The CU encoding unit 1112 encodes the CU information, prediction information, division information, and so on.
[0134] When a prediction error is included in a TU, the TU encoding unit 1114 encodes the QP update information and the quantized prediction error.
[0135] The CT information encoding unit 1111 and the CU encoding unit 1112 supply syntax elements such as inter prediction parameters (predMode, general_merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX), intra prediction parameters, and quantized transform coefficients to the parameter encoding unit 111.
[0136] The entropy coding unit 104 receives the quantized transform coefficients and coding parameters (division information, prediction parameters) from the parameter coding unit 111. These are then entropy coded to generate the coded stream Te, which is then output.
[0137] The prediction parameter derivation unit 120 is a means including the inter-prediction parameter coding unit 112 and the intra-prediction parameter coding unit, and uses the parameters input from the coding parameter determination unit 110. The intra prediction parameter and the inter prediction parameter are output to the parameter coding unit 111. do.
[0138] (Configuration of the inter-prediction parameter encoding unit) The inter-prediction parameter encoding unit 112 includes a parameter encoding control unit 1121, an inter-prediction The parameter encoding control unit 1121 includes a parameter derivation unit 303. The inter prediction parameter derivation unit 303 is a common component to that of the video decoding device. The parameter encoding control unit 1121 includes a merge index derivation unit 11211 and a vector candidate index derivation unit 11212.
[0139] The merge index derivation unit 11211 derives merge candidates and the like, and calculates inter prediction parameters The vector candidate index derivation unit 11212 derives predicted vector candidates and the like, and outputs them to the inter prediction parameter derivation unit 303 and the parameter coding unit 111.
[0140] (Configuration of intra-prediction parameter encoding unit) The intra-prediction parameter coding unit includes a parameter coding control unit and an intra-prediction parameter derivation unit. The intra-prediction parameter derivation unit has a common configuration with the video decoding device.
[0141] However, unlike the video decoding device, the inter prediction parameter derivation unit 303 and the intra prediction The inputs to the parameter derivation unit are the encoding parameter determination unit 110 and the prediction parameter memory 108 , and the parameter derivation unit outputs to a parameter encoding unit 111 .
[0142] The adder 106 generates a decoded image by adding, for each pixel, the pixel value of the predicted block input from the predicted image generation unit 101 and the prediction error input from the inverse quantization and inverse transform unit 105. The adder 106 stores the generated decoded image in a reference picture memory 109.
[0143] The loop filter 107 performs deblocking filtering, SAO, and ALF on the decoded image generated by the adder 106. Note that the loop filter 107 does not necessarily include the above three types of filters. For example, the filter may be configured with only a deblocking filter.
[0144] The prediction parameter memory 108 stores the prediction parameters generated by the coding parameter determination unit 110 in a predetermined location for each current picture and CU.
[0145] The reference picture memory 109 stores the decoded image generated by the loop filter 107 at a predetermined position for each current picture and CU.
[0146] The encoding parameter determination unit 110 determines one of the multiple sets of encoding parameters. The coding parameters are the above-mentioned QT, BT or TT division information, prediction parameters, or parameters to be coded that are generated in relation to these. The predicted image generating unit 101 generates a predicted image using these coding parameters.
[0147] The coding parameter determination unit 110 determines the size of the amount of information and the coding parameter for each of the plurality of sets. The RD cost value indicating the error is calculated. The RD cost value is, for example, the sum of the code amount and the value obtained by multiplying the squared error by a coefficient λ. The code amount is the information amount of the coded stream Te obtained by entropy coding the quantization error and the coding parameters. The squared error is calculated in the subtraction unit 102. The cost value is the sum of squares of the prediction errors calculated by the coding unit 110. The coefficient λ is a preset real number greater than zero. The coding parameter determination unit 110 selects a set of coding parameters that minimizes the calculated cost value. The coding parameter determination unit 110 outputs the determined coding parameters to the parameter coding unit 111 and the prediction parameter derivation unit 120.
[0148] In addition, in the above-described embodiment, a part of the video encoding device 11 and the video decoding device 31, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generating unit 306, The control unit 308, the inverse quantization and inverse transform unit 311, the addition unit 312, the prediction parameter derivation unit 320, the prediction image generation unit 101, the subtraction unit 102, the transformation and quantization unit 103, the entropy coding unit 104, the inverse quantization and inverse transform unit 105, the loop filter 107, the coding parameter determination unit 110, the parameter coding unit 111, and the prediction parameter derivation unit 120 may be realized by a computer. In this case, the control functions may be realized by recording a program for realizing the control functions in a computer-readable recording medium, reading the program recorded in the recording medium into a computer system, and executing the program. Note that the "computer system" here refers to a computer system built into either the video coding device 11 or the video decoding device 31, and includes hardware such as an OS and peripheral devices. Also, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, and a storage device such as a hard disk built into a computer system. Furthermore, the term "computer-readable recording medium" may include a medium that dynamically stores a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and a medium that stores a program for a certain period of time, such as a volatile memory inside a computer system that serves as a server or client in such a case. The above program may be one that realizes part of the above-mentioned functions, or may be one that can realize the above-mentioned functions in combination with a program already recorded in the computer system.
[0149] In addition, a part or the whole of the video encoding device 11 and the video decoding device 31 in the above-mentioned embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the video encoding device 11 and the video decoding device 31 may be individually made into a processor, or a part or the whole may be integrated into a processor. The integrated circuit method is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. Furthermore, when an integrated circuit technology that replaces LSI appears due to the progress of semiconductor technology, an integrated circuit based on that technology may be used.
[0150] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design changes, etc. are possible within the scope that does not deviate from the gist of the present invention.
[0151] The present invention is not limited to the above-described embodiment, and various modifications are possible within the scope of the claims. In other words, the technical scope of the present invention also includes embodiments obtained by combining technical means that are appropriately modified within the scope of the claims. [Industrial Applicability]
[0152] The embodiments of the present invention can be suitably applied to a video decoding device that decodes coded data in which image data is coded, and a video coding device that generates coded data in which image data is coded, and can also be suitably applied to the data structure of coded data that is generated by a video coding device and referenced by the video decoding device. [Explanation of symbols]
[0153] 31 Video Decoding Device 301 Entropy Decoding Unit 302 Parameter Decoding Unit 3022 CU Decoding Unit 3024 TU Decoding Unit 303 Inter-prediction parameter derivation unit 305, 107 Loop Filter 306, 109 Reference Picture Memory 307, 108 Prediction parameter memory 308, 101 Prediction image generation unit 309 Inter-prediction image generation unit 3092 OOB Processing Unit 30921 OOB judgment section 30922 OOB mask lead-out part 311, 105 Inverse quantization and inverse transformation unit 312, 106 Addition section 320 Prediction Parameter Derivation Unit 11 Video Encoding Device 102 Subtraction section 103 Transformation and Quantization Section 104 Entropy coding unit 110 Encoding parameter determination unit 111 Parameter Encoding Unit 112 Inter-prediction parameter coding unit 120 Prediction parameter derivation part
Claims
1. An MPEG decoder, comprising: an OOB determination unit that determines whether the reference block is a target for OOB processing by comparing the coordinates of the reference block with the coordinates of the picture; and an OOB mask derivation unit that derives mask data indicating whether each pixel can be used by comparing the coordinates of the pixels belonging to the reference block with the coordinates obtained by adding an offset to the coordinates of the picture, wherein the coordinates of the reference block are the coordinates obtained by adding a motion vector to the coordinates of the target block, and the OOB determination unit determines whether the reference block is a target for OOB processing by using whether or not there is a process applied to the reference picture including the reference block.
2. The MPEG decoder according to claim 1, wherein the OOB determination unit determines whether the reference block is a target for the OOB processing by using whether or not there is a wrap-around process applied to the reference picture including the reference block.
3. The MPEG decoder according to claim 2, wherein when the wrap-around process is applied to the reference picture including the reference block, the OOB determination unit determines that the OOB processing for the reference block is invalid.
4. The MPEG decoder according to claim 2, wherein when the wrap-around process is applied to the reference picture including the reference block, the OOB determination unit changes a method of determining whether the reference block is a target for the OOB processing.
5. The MPEG decoder according to claim 4, wherein when the wrap-around process is applied to the reference picture including the reference block, the OOB determination unit determines whether the reference block is a target for the OOB processing by using coordinate values outside the application target of the wrap-around process.
6. The MPEG decoder according to claim 1, wherein the OOB determination unit determines whether the reference block is a target for the OOB processing by using whether or not there is a reference picture resampling process applied to the reference picture including the reference block.
7. The moving image decoding apparatus according to claim 6, wherein the OOB determination unit determines to invalidate the OOB process on the reference block when the reference picture resampling process is applied to the reference picture including the reference block.
8. The moving image decoding apparatus according to claim 6, wherein the OOB determination unit changes a determination method as to whether the reference block is an OOB process target when the reference picture resampling process is applied to the reference picture including the reference block.
9. The moving image decoding apparatus according to claim 8, wherein the OOB determination unit determines whether the reference block is an OOB process target using coordinate values corrected using the same scale value as the reference picture resampling process when the reference picture resampling process is applied to the reference picture including the reference block.
10. A moving image encoding apparatus including an OOB determination unit that determines whether the reference block is an OOB process target by comparing the coordinates of the reference block with the coordinates of the picture, and an OOB mask derivation unit that derives mask data indicating whether each pixel can be used by comparing the coordinates of the pixels belonging to the reference block with the coordinates obtained by adding an offset to the coordinates of the picture. The coordinates of the reference block are coordinates obtained by adding a motion vector to the coordinates of the target block. The moving image encoding apparatus, wherein the OOB determination unit determines whether the reference block is an OOB process target using the presence or absence of a process applied to the reference picture including the reference block.