Image encoding device, image encoding method and program, image decoding device, image decoding method and program
By introducing motion-compensated pixel interpolation technology into the image encoding device, the problem of insufficient interpolation accuracy in inter-frame prediction in the prior art is solved, and higher encoding efficiency and image quality are achieved.
Patent Information
- Application Number
- CN202380072267.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-13
- Filing Date
- 2023-08-02
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art cannot effectively apply motion compensation pixel interpolation in inter-frame prediction, especially when the zoom window is set as a reference picture and the reference picture is enlarged or reduced, or blocks within the picture boundary of the reference picture do not include motion information.
An image encoding device is designed, including prediction, encoding, interpolation and transformation components, by using motion compensation pixel interpolation in inter-frame prediction, an out-of-screen pixel of a reference frame is generated, and the interpolation accuracy is improved in the case of resolution transformation.
The motion-compensated pixel interpolation accuracy in inter-frame prediction is achieved, the encoding efficiency is improved, and the image quality is improved.
Smart Images

Figure CN120077660A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image encoding and decoding technologies. Background Art
[0002] A known method for encoding compressed recordings of videos is the Versatile Video Coding (VVC) coding method (hereinafter simply referred to as VVC). To improve the encoding efficiency of VVC, a basic block called a Coding Tree Unit (CTU) is divided into rectangular sub-blocks instead of typical squares.
[0003] In addition, in the case of VVC, in order to achieve efficient inter-frame prediction for smooth scenes and the like, a technique for controlling the spatial resolution of a reference picture called a scaling window is applied. Further, in video coding such as represented by VVC, since pixels outside the picture boundary of the reference picture are used in inter-frame prediction, interpolation of pixels outside the picture boundary must be performed. Patent Document 1 describes an interpolation technique for pixels outside the picture boundary in encoding processing in units of tiles.
[0004] Recently, the Joint Video Exploration Team (JVET) that established the VVC standard has been researching novel coding technologies that can be superior to VVC in both encoding efficiency and image quality enhancement. To improve the encoding efficiency, one such technique being researched for introduction is a novel extrapolation method (hereinafter referred to as motion-compensated pixel interpolation), which is used to generate pixels outside the picture boundary of a reference picture used in inter-frame prediction using pixels within the picture boundary of a reference picture different from the first reference picture.
[0005] Citation List
[0006] Patent Document
[0007] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2018-050085 Summary of the Invention
[0008] Problems to be Solved by the Invention
[0009] In VVC, a technique called a scaling window is adopted, and this technique allows setting a rectangular area for each picture based on scaling processing. In inter-frame prediction, by comparing the sizes of the scaling windows set for each frame, inter-frame prediction considering the magnification or reduction of objects existing across frames can be performed.
[0010] In addition, in VVC, since a predicted image is generated using inter-frame prediction, information of pixels located outside the picture boundary of a reference picture can be used. Since pixels outside the picture boundary are not targets for encoding, interpolation that simply duplicates pixels within the picture boundary of the reference picture can be used as a common interpolation method for both encoding and decoding. Regarding this point, in JVET, research is being conducted on motion-compensated pixel interpolation, in which motion information included in blocks within the picture boundary is used to generate pixels outside the picture boundary of a reference frame from pixels within the picture boundary of a reference frame different from the first reference frame.
[0011] However, motion-compensated pixel interpolation cannot be applied to cases such as when a scaling window is set to a reference picture and the reference picture is enlarged or reduced, and cases such as when blocks within the picture boundary of the reference picture do not include motion information. Therefore, there is a problem that the prediction accuracy of inter-frame prediction cannot be improved.
[0012] In view of such problems, the present invention enables a technique for encoding in a case where the interpolation accuracy of motion-compensated pixels in inter-frame prediction along with resolution conversion is enhanced and more efficient than before.
[0013] Solution to the problem
[0014] To solve the above problems, an image encoding apparatus according to the present invention has the following configuration, for example. Provided are:
[0015] A prediction unit that generates a predicted image for a target block in a first frame to be encoded by referring to a second frame encoded before the first frame;
[0016] An encoding unit that encodes a prediction error of the target block with respect to the predicted image;
[0017] An interpolation unit that interpolates pixels outside the boundary of the second frame using pixels of a third frame encoded before the second frame in a case of referring to pixels outside the boundary of the second frame; and
[0018] A transformation unit that changes the resolution of a frame before the first frame.
[0019] Other features and advantages of the present invention will be apparent from the following description in conjunction with the accompanying drawings. Note that in all the drawings, the same reference numerals denote the same or similar components.
[0020] Effects of the invention
[0021] According to the present invention, pixels outside the picture of a reference frame for inter-frame prediction can be accurately generated, and the encoding efficiency can be improved. Description of the drawings
[0022] The accompanying drawings incorporated in and forming a part of the specification illustrate embodiments of the present invention and, together with the specification, serve to explain the principles of the present invention.
[0023] Figure 1 is a block configuration diagram of an image encoding device according to an embodiment.
[0024] Figure 2 is a block configuration diagram of an image decoding device according to an embodiment.
[0025] Figure 3 is a flowchart showing an encoding process according to an embodiment.
[0026] Figure 4 is a flowchart showing an image decoding process according to an embodiment.
[0027] Figure 5 is a diagram of the hardware configuration of a computer that can be applied to an image encoding device and a decoding device according to an embodiment.
[0028] Figure 6A is a diagram showing an example of a bitstream structure.
[0029] Figure 6B is a diagram showing another example of a bitstream structure.
[0030] Figure 7A is a diagram showing an example of sub-block division in which one sub-block has the same size as the basic block size used in this embodiment.
[0031] Figure 7B is a diagram showing an example of division into four square sub-blocks used in this embodiment.
[0032] Figure 7C Shows examples of the types of rectangular sub-blocks obtained by sub-block division.
[0033] Figure 7D Shows examples of the types of rectangular sub-blocks obtained by sub-block division.
[0034] Figure 7E Shows examples of the types of rectangular sub-blocks obtained by sub-block division.
[0035] Figure 7F Shows examples of the types of rectangular sub-blocks obtained by sub-block division.
[0036] Figure 8 is a diagram showing an example of simple copy pixel interpolation.
[0037] Figure 9 is a diagram showing an example of interpolating off-screen pixels via motion-compensated pixel interpolation.
[0038] Figure 10A It is a diagram showing an example of off - screen pixel interpolation by motion - compensated pixel interpolation along with resolution transformation of a reference picture.
[0039] Figure 10B It is a diagram showing an example of off - screen pixel interpolation by motion - compensated pixel interpolation along with resolution transformation of a reference picture.
[0040] Figure 11 It is a diagram showing an example of pixel generation of pixels outside a picture boundary that are not interpolated by motion - compensated pixel interpolation.
[0041] Figure 12 It is a diagram showing an example of off - screen pixel interpolation by motion - compensated pixel interpolation using information of blocks adjacent to a boundary block.
[0042] Figure 13A It is a diagram showing an example of high - speed processing of motion - compensated pixel interpolation according to this embodiment.
[0043] Figure 13B It is a diagram showing an example of high - speed processing of motion - compensated pixel interpolation according to this embodiment. Detailed Description of the Embodiment
[0044] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments are not intended to limit the scope of the claimed invention. In the embodiments, a plurality of features are described, but the invention is not limited to an invention that requires all such features, and a plurality of such features can be appropriately combined. In addition, in the drawings, the same reference numerals are given to the same or similar configurations, and redundant descriptions thereof are omitted.
[0045] First Embodiment
[0046] Figure 1 It is a block diagram showing an image encoding device according to this embodiment. In the same diagram, a control unit 100 that controls the entire device includes a CPU and a memory that stores programs for executing the CPU. A terminal 101 is an input terminal for inputting image data. A generation source of video data to be encoded is connected to the terminal 101. The type of the generation source of video data is not particularly limited, and a imaging unit or a storage device for storing video and image data to be encoded is generally used.
[0047] A block division unit 102 divides a frame of image received via the terminal 101 into a plurality of basic blocks and outputs block images to subsequent stages in units of basic blocks.
[0048] The generation unit 103 generates offset information and the like used in the resolution transformation of the reference picture, and outputs the offset information and the like to the resolution transformation unit 113 and the comprehensive encoding unit 111. In this embodiment, the offset information is information for setting a scaling window in each picture and information generated for calculating the magnification or reduction ratio of the picture (hereinafter referred to as resolution transformation control information). The method for generating the resolution transformation control information is not particularly limited. The user can input the resolution transformation control information, information derived from operations such as zoom-in / zoom-out operations detected in the video camera device can be input as the resolution transformation control information, or the resolution transformation control information specified in advance as an initial value can be used.
[0049] The prediction unit 104 generates sub-blocks by dividing the basic block. In addition, the prediction unit 104 determines whether to perform intra-frame prediction that is prediction within a frame in units of sub-blocks or inter-frame prediction that is prediction between frames. The prediction unit 104 appropriately refers to an image (hereinafter referred to as an interpolated image) obtained by interpolating pixels outside the picture boundary supplied from the interpolation unit 114, and generates prediction image data. In addition, the prediction unit 104 calculates a prediction error based on the image data to be encoded and the generated prediction image data, and outputs the prediction error to the transform / quantization unit 105. In addition, the prediction unit 104 outputs information required for prediction such as sub-block division, prediction mode, motion vector, and similar information together with the prediction error. Hereinafter, the information required for prediction is referred to as prediction information. In addition, the prediction unit 104 outputs the prediction information to the interpolation unit 114.
[0050] The transform / quantization unit 105 performs an orthogonal transform on the prediction error data for each sub-block, quantizes the obtained transform coefficients using the set quantization parameter, and obtains residual coefficients. Note that the quantization parameter is a parameter used in the quantization of the transform coefficients obtained through the orthogonal transform.
[0051] The inverse quantization / inverse transform unit 106 performs inverse quantization on the residual coefficients output from the transform / quantization unit 105, reconstructs the transform coefficients, performs an inverse orthogonal transform, and reconstructs the prediction error data.
[0052] The frame memory 108 is a memory that stores the reconstructed image data.
[0053] The resolution transformation unit 113 enlarges or reduces the image stored in the frame memory 108 based on the resolution transformation control information, and outputs the enlarged or reduced image as a resolution transformation image.
[0054] The interpolation unit 114 appropriately refers to the resolution-transformed image output by the resolution transformation unit 113, the prediction information output by the prediction unit 104, and the filtered image stored in the frame memory 108, and generates interpolation image data. Then, the interpolation unit 114 outputs the generated interpolation image data to the prediction unit 104 and the image reconstruction unit 107. Note that the interpolation unit 114 may store the image data including the generated out-of-picture pixel information as an interpolation image in the frame memory 108. In addition, the interpolation unit 114 may retrieve the generated interpolation image from the frame memory 108, and may output the generated interpolation image to the prediction unit 104 and the image reconstruction unit 107.
[0055] The image reconstruction unit 107 generates the reconstructed image data based on the prediction information output from the prediction unit 104, according to the interpolation image data output by the interpolation unit 114, and the prediction error data.
[0056] The loop filter unit 109 performs loop filter processing such as deblocking filtering and sample adaptive offset on the reconstructed image.
[0057] The encoding unit 110 encodes the residual coefficients output from the transform / quantization unit 105 and the prediction information output from the prediction unit 104, and generates encoded data.
[0058] The comprehensive encoding unit 111 encodes the resolution transformation control information that is the output of the generation unit 103, and generates header encoded data. In addition, the comprehensive encoding unit 111 forms a bitstream together with the encoded data output from the encoding unit 110.
[0059] The terminal 112 is an output terminal that outputs the bitstream generated by the comprehensive encoding unit 111 to an external unit. The types of output destinations include, for example, a network and a storage device (including a storage medium).
[0060] The configuration and basic operation of the image encoding device according to the embodiment have been described above. Next, the image encoding operation of the image encoding device will be described below. In this embodiment, video data of one frame is input at a time, but in other possible configurations, still image data of one frame is input.
[0061] Before image encoding, generation unit 103 generates resolution transformation control information. The resolution transformation control information according to the present embodiment includes the horizontal size and vertical size of the current picture, offset information representing the scaling window of the current picture, and the resolution transformation magnification in the horizontal direction and the vertical direction. The horizontal size and vertical size of the current picture are the number of pixels in the horizontal direction and the number of pixels in the vertical direction of the image input from terminal 101. The offset information representing the scaling window is information defining the position and size of the scaling window relative to the current picture. In the present embodiment, in the offset information, the distances from each side of the left, right, top, and bottom of the current picture to each side of the scaling window are used to define the position and size of the scaling window. The offset relative to each side is represented by the number of pixels. Further, if the offset is in the direction from the picture frame boundary to the center, the offset is represented by a positive value, and if the offset is in the direction from the picture frame boundary outward, the offset is represented by a negative value. In the following description, the offset relative to each side is referred to as the left offset, right offset, top offset, and bottom offset. Further, the offset of the current picture is referred to as offset C, and the offset of the reference picture is referred to as offset R to distinguish between the two.
[0062] Next, generation unit 103 uses the above-mentioned offsets relative to each side to calculate, for both the horizontal direction and the vertical direction, the resolution transformation magnification used to represent the magnification or reduction ratio of the reference picture relative to the current picture.
[0063] For example, the resolution transformation magnification in the horizontal direction is calculated using the following formula.
[0064] Resolution transformation magnification in the horizontal direction = (Reference picture horizontal size - Left offset R - Right offset R) / (Current picture horizontal size - Left offset C - Left offset C)
[0065] In addition, the resolution transformation magnification in the vertical direction is calculated using the following formula.
[0066] Resolution transformation magnification in the vertical direction = (Reference picture vertical size - Top offset R - Bottom offset R) / (Current picture vertical size - Top offset C - Bottom offset C)
[0067] Using these calculation formulas, for all reference pictures that can be referred to in the encoding of the current picture, the resolution transformation magnification relative to the current picture is calculated. Generation unit 103 stores the resolution transformation magnification calculated in this way in the resolution transformation control information and outputs it to resolution transformation unit 113 and comprehensive encoding unit 111.
[0068] The image data of one frame input from terminal 101 is supplied to block segmentation unit 102.
[0069] At the block segmentation unit 102, the input image data is segmented into a plurality of basic blocks, and the images in units of the basic blocks are output to the prediction unit 104.
[0070] The prediction unit 104 performs prediction processing on the image data (basic block images) input from the block segmentation unit 102. Specifically, the prediction unit 104 first performs a process (sub-block segmentation process) to further segment the input image of the basic block size into smaller sub-blocks.
[0071] Figures 7A to 7F An example of a basic block segmented into sub-blocks is shown. The reference numeral 700 indicates the basic block with a thick frame, and for simplicity of description, in this example, the size of the basic block is 32×32 pixels. The quadrilaterals within the thick frame represent sub-blocks. Figure 7A An example in which one sub-block has the same size as the basic block is shown.
[0072] Figure 7B An example of being segmented into four square sub-blocks is shown. Since the size of the basic block is 32×32 pixels, Figure 7B each of the four sub-blocks has a size of 16×16 pixels. In addition, Figures 7C to 7F Examples of different types of rectangular sub-blocks obtained through sub-segmentation are shown. Figure 7C An example of the basic block being segmented into two vertically long rectangular sub-blocks with a size of 16×32 pixels is shown. Figure 7D An example of being segmented into two horizontally long rectangular sub-blocks with a size of 32×16 pixels is shown. In addition, Figure 7E and Figure 7F An example of being segmented into rectangles in a 1∶2∶1 ratio is shown. In this way, in the present embodiment, encoding processing is performed using rectangular sub-blocks as well as square sub-blocks.
[0073] The prediction unit 104 determines a prediction mode for each sub-block to be processed. Specifically, the prediction unit 104 determines the prediction mode by judging whether to use intra prediction using the encoded pixels of the current picture including the sub-block to be processed or inter prediction using the pixels of an encoded picture different from the current picture. In addition, the prediction unit 104 generates prediction image data based on the determined prediction mode and the encoded pixels. In addition, the prediction unit 104 generates prediction error data from the input image data and the generated prediction image data, and outputs the generated prediction error data to the transform / quantization unit 105. Note that the interpolation unit 114 may generate pixels outside the picture boundary of the reference picture via pixel interpolation, or may generate pixels outside the picture boundary of the reference picture via resolution conversion of the encoded pixels. A method of generating pixels outside the picture boundary via motion compensation pixel interpolation in inter prediction together with resolution conversion will be described below. In addition, the prediction unit 104 outputs information such as sub-block division and prediction mode as prediction information to the encoding unit 110 and the image reconstruction unit 107.
[0074] The transform / quantization unit 105 performs frequency conversion on the prediction error data of the sub-block that has been subjected to prediction processing by the prediction unit 104, and further performs quantization. A method for determining the value of the quantization parameter used in quantization is not particularly limited, and the user may input the quantization parameter, or may calculate the value according to the characteristics of the input image, or may specify the value in advance as an initial value.
[0075] The inverse quantization / inverse transform unit 106 performs inverse quantization on the input residual coefficients to reconstruct the transform coefficients, and further performs inverse orthogonal transform on the reconstructed transform coefficients to reconstruct the prediction error data. Then, the inverse quantization / inverse transform unit 106 outputs the reconstructed prediction error data to the image reconstruction unit 107. Note that the quantization parameter used when the inverse quantization / inverse transform unit 106 performs inverse quantization of the sub-block is the same quantization parameter used when the transform / quantization unit 105 quantizes the sub-block.
[0076] The image reconstruction unit 107 appropriately refers to the interpolated image supplied from the interpolation unit 114 based on the prediction information input from the prediction unit 104, and generates a prediction image. Then, the image reconstruction unit 107 reconstructs the image data according to the generated prediction image and the prediction error data generated by the inverse quantization / inverse transform unit 106, and stores the reconstructed image data in the frame memory 108.
[0077] The loop filter unit 109 reads out the reconstructed image from the frame memory 108 and performs loop filter processing such as deblocking filtering. The loop filter processing is performed based on the prediction mode of the prediction unit 104 and the value of the quantization parameter used by the transform / quantization unit 105, and also based on whether there is a non-zero value in the sub-block or sub-block segmentation information of the post-quantization processing. The loop filter unit 109 stores the image data obtained through the filtering process back in the frame memory 108 again.
[0078] The resolution conversion unit 113 enlarges or reduces the image stored in the frame memory 108 based on the resolution conversion control information. In addition, the resolution conversion unit 113 outputs the enlarged or reduced image as a resolution-converted image.
[0079] Here, an example in the case where a scaling window is set for the reference picture and no scaling window is set for the current picture will be used to describe the enlargement or reduction process. In this example, the horizontal size of the current picture and the reference picture is set to 1920 pixels, and the vertical size is set to 1080 pixels. In addition, both the left offset and the right offset set for the reference picture are 240 pixels, and both the upper offset and the lower offset are 135 pixels. In addition, no offset is set for the current picture. In this case, using the above resolution conversion magnification calculation formula, the resolution conversion magnification in the horizontal and vertical directions is calculated as follows.
[0080] Resolution conversion magnification in the horizontal direction = (1920 - 240 - 240) / 1920 = 0.75
[0081] Resolution conversion magnification in the vertical direction = (1080 - 135 - 135) / 1080 = 0.75
[0082] The resolution transformation magnification represents the size of the rectangular area of the reference picture to which the offset is applied relative to the rectangular area of the current picture to which the offset is applied. In the above example, the reference picture is reduced by 0.75, i.e., 3 / 4, relative to the current picture. Therefore, in this case, the resolution transformation unit 113 generates a resolution transformation image by magnifying the reference picture by a magnification factor of 4 / 3, which is the reciprocal of the resolution transformation magnification. Then, the resolution transformation unit 113 outputs the generated resolution transformation image to the interpolation unit 114. The interpolation filter or decimation filter (hereinafter referred to as the resolution transformation filter) used in the resolution transformation is not particularly limited, and the user can input one or more resolution transformation filters, or can use a pre-specified value as the initial value. In addition, the resolution transformation unit 113 can switch between multiple resolution transformation filters according to the resolution transformation magnification and generate a resolution transformation image. In this way, by magnifying or reducing the reference picture based on the resolution transformation magnification, the inter-frame prediction can follow the magnification or reduction of the object in the video in the magnified or reduced scene.
[0083] In the above example, no offset is set for the current picture, but this embodiment is not limited thereto, and a scaling window may not be set for the current picture. In this example, the horizontal size of the current picture and the reference picture is set to 1920 pixels, and the vertical size is set to 1080 pixels. In addition, the left offset and the right offset set for the reference picture are both 510 pixels, and the upper offset and the lower offset are both 270 pixels. In addition, the left offset and the right offset set for the current picture are both 360 pixels, and the upper offset and the lower offset are both 180 pixels. In such a case, the resolution transformation magnifications in the horizontal and vertical directions are calculated as follows.
[0084] Resolution transformation magnification in the horizontal direction = (1920 - 510 - 510) / (1920 - 360 - 360) = 0.75 Resolution transformation magnification in the vertical direction = (1080 - 270 - 270) / (1080 - 180 - 180) = 0.75
[0085] In this way, the same resolution transformation magnification as when no offset is set for the current picture as described above can be specified.
[0086] In addition, the horizontal and vertical sizes of the current picture and the reference picture may be different in terms of the number of pixels. For example, the horizontal size of the current picture may be 1920 pixels, and the vertical size may be 1080 pixels, the horizontal size of the reference picture may be 960 pixels, and the vertical size may be 540 pixels, and a scaling window may not be set for the current picture and the reference picture. In this case, the resolution transformation magnifications in the horizontal and vertical directions are calculated as follows.
[0087] Resolution transformation magnification in the horizontal direction = (960 - 0 - 0) / (1920 - 0 - 0) = 0.5
[0088] Resolution transformation magnification in the vertical direction = (540 - 0 - 0) / (1080 - 0 - 0) = 0.5
[0089] In this way, even in the case of images with different numbers of pixels between the current picture and the reference picture, the resolution transformation magnification can be specified. As a result, even if a video is composed of images magnified or reduced in one section compared to other sections, the interpolation unit 114 can use motion compensation pixel interpolation to generate pixels outside the picture boundary of the reference picture.
[0090] The interpolation unit 114 appropriately refers to the resolution-transformed image output by the resolution transformation unit 113, the prediction information output by the prediction unit 104, and the filtered image stored in the frame memory 108, and generates image data obtained by interpolating pixels outside the picture boundary of the current picture. To interpolate pixels outside the picture boundary of the current picture, the interpolation unit 114 uses the prediction information output by the prediction unit 104 to identify the pixel positions required for interpolation of the pixels outside the picture, and generates the pixels outside the picture. Then, the interpolation unit 114 outputs the generated image data to the prediction unit 104 and the image reconstruction unit 107. The information on the pixels outside the picture can be generated each time the filtered image is referred to, or the information on the pixels outside the picture once calculated can be stored in the frame memory 108 as an interpolation image associated with the filtered image.
[0091] To generate the image data, the interpolation unit 114 first retrieves from the frame memory 108 the filtered image of the current picture that is the target of interpolation of the pixels outside the picture, and holds the filtered image together with the prediction information used in the encoding of the current picture input from the prediction unit 104. Next, the interpolation unit 114 reads out the prediction modes of all sub-blocks in contact with the inside of the picture boundary of the filtered image of the current picture from the prediction information. In addition, the interpolation unit 114 determines the interpolation method for the pixels outside the picture boundary of the filtered image of the current picture according to the states of the prediction modes of each block (hereinafter referred to as boundary blocks) in contact with the inside of the read picture boundary of the current picture. In this embodiment, the method for pixel interpolation of the pixels outside the picture can be interpolation (hereinafter referred to as simple copy pixel interpolation) generated by simply copying the pixels inside the picture boundary of the reference picture outside the picture boundary of the same reference picture, or motion compensation pixel interpolation that uses the pixels inside the picture boundary of a reference picture different from the first reference picture to generate the pixels inside the picture boundary of the reference picture. The simple copy pixel interpolation and the motion compensation pixel interpolation will be described in detail below.
[0092] Here, reference will be made to Figure 10A andFigure 10B A method of using motion compensation pixel interpolation to generate pixels outside a picture boundary in inter-frame prediction with resolution conversion together with reference pictures is described.
[0093] Figure 10A and Figure 10B The case where the prediction unit 104 encodes the sub-block 1002 via inter-frame prediction is shown. The current picture 1001 indicated by the thick frame in the right figure is the encoding target, and the rectangular area indicated by the thick frame in the central figure is the reference picture 1011. The thin frame 1010 is an area including pixels outside the picture boundary of the reference picture 1011, and the pixels outside the picture boundary are generated via simple copy pixel interpolation or motion compensation pixel interpolation. The predicted image 1012 indicated by the rectangular area in the reference picture 1011 is generated by the prediction unit 104 via inter-frame prediction and represents the predicted block of the sub-block 1002. The left side of the boundary block 1013 and the left side of the boundary block 1014 are in contact with the left side of the reference picture 1011 on the inner side, and the prediction modes of the boundary block 1013 and the boundary block 1014 are inter-frame prediction. In addition, the rectangular area 1015 and the rectangular area 1016 are filled with pixels generated via simple copy pixel interpolation or motion compensation pixel interpolation. Note that even when the upper side of the boundary block is in contact with the upper side of the reference picture on the inner side, when the right side of the boundary block is in contact with the right side of the reference picture on the inner side, and when the lower side of the boundary block is in contact with the lower side of the reference picture on the inner side, as in the case where the left side of the boundary block is in contact with the left side of the reference picture on the inner side, the interpolation unit 114 uses simple copy pixel interpolation or motion compensation pixel interpolation to generate pixels outside the picture boundary of the reference picture 1011. In addition, in the case of using motion compensation pixel interpolation, the interpolation unit 114 uses the pixels at the picture boundary of the reference picture to generate pixels outside the picture boundary of the reference picture that are not generated using motion compensation pixel interpolation via simple copy pixel interpolation. Then, in the motion compensation pixel interpolation of each boundary block, based on the motion information of each boundary block, for each boundary block among all boundary blocks in contact with the inner side within the boundary block of the reference picture, motion compensation pixel interpolation is used to generate pixels located outside the picture boundary of the reference picture. In addition, in the processing of each boundary block, pixels located outside the picture boundary of the reference picture are generated without referring to the motion information of other boundary blocks.
[0094] In Figure 10AIn the example shown, the interpolation image 1021 is a filtered image used in the inter-frame prediction of the boundary block of the reference picture 1011 and is a picture retrieved from the frame memory 108. In addition, the pixel values within the picture boundary of the picture correspond to the picture of the pixels outside the picture boundary for generating the reference picture via motion-compensated pixel interpolation. No scaling window is set for the interpolation image 1021 and the reference picture 1011, and all offsets are 0. Therefore, the resolution conversion ratio is 1 in both the horizontal and vertical directions, and no resolution conversion of the interpolation image 1021 is required in the inter-frame prediction of the boundary block 1013. As a result, the prediction block 1023 corresponding to the boundary block 1013 located within the prediction image 1012 is located within the picture boundary of the interpolation image 1021. In a similar manner, the prediction block 1024 corresponding to the boundary block 1014 located within the prediction image 1012 is located within the picture boundary of the interpolation image 1021. Therefore, the interpolation unit 114 uses the pixel group of the rectangular area 1025 of the interpolation image 1021 to generate the pixels of the rectangular area 1015 located outside the picture boundary of the reference picture. In a similar manner, the interpolation unit 114 uses the pixel group of the rectangular area 1026 of the interpolation image 1021 to generate the pixels of the rectangular area 1016 located outside the picture boundary of the reference picture. Therefore, the interpolation unit 114 uses the pixels within the picture boundary of the interpolation image 1021 and generates, via motion-compensated pixel interpolation, the pixels outside the picture boundary of the reference picture that constitutes the prediction image 1012 referred to in the inter-frame prediction of the sub-block 1002. In addition, for all boundary blocks of the reference picture 1011, the interpolation unit 114 stores the reference picture obtained by completing the interpolation outside the picture boundary of the reference picture generated by performing motion-compensated pixel interpolation in the frame memory 108.
[0095] In Figure 10B the example shown, the offsets of the scaling windows set in the interpolation image 1021 and the reference picture 1011 are not 0. In other words, in the inter-frame prediction of the boundary block 1013, resolution conversion of the interpolation image 1021 is required. Therefore, the prediction block 1023 corresponding to the boundary block 1013 located within the prediction image 1012 is located within the picture boundary of the resolution-converted image (hereinafter referred to as the resolution-converted interpolation image) obtained via resolution conversion based on the offset information of the interpolation image 1021 and the reference picture 1011 (instead of within the interpolation image 1021). In a similar manner, the prediction block 1024 corresponding to the boundary block 1014 located within the prediction image 1012 is located within the picture boundary of the resolution-converted interpolation image 1022 (instead of within the interpolation image 1021).
[0096] Now, the process of generating the interpolated image 1022 after resolution conversion from the interpolated image 1021 will be described. To simplify the description, in this example, the interpolated image 1021 is provided with a scaling window, and the current picture 1001 and the reference picture 1011 are not provided with a scaling window. However, this embodiment is not limited thereto. The current picture 1001 and the reference picture 1011 may be provided with a scaling window. In Figure 10B In the example shown, the horizontal size of the reference picture 1011 and the interpolated image 1021 is 1920 pixels, and the vertical size is 1080 pixels. The left offset and the right offset set for the interpolated image 1021 are both 240 pixels, and the upper offset and the lower offset are both 135 pixels. In addition, the offsets of the four sides of the reference picture are 0 pixels. In this case, using the above calculation formula, in the horizontal direction and the vertical direction, the resolution conversion magnification is calculated to be 0.75. Therefore, the interpolated image 1022 after resolution conversion is an image of the interpolated image 1021 enlarged by 4 / 3 times, which is the reciprocal of 0.75. In this way, based on the offset information set for the reference picture 1011 and the interpolated image 1021, the interpolated image 1022 after resolution conversion is generated from the interpolated image 1021.
[0097] In Figure 10BIn this case, the predicted image 1012 of the reference picture 1011 represents the predicted block of the sub-block 1002 generated by the prediction unit 104 via inter-frame prediction. In addition, the prediction modes of the boundary blocks 1013 and 1014 of the reference picture 1011 included in the predicted image 1012 are inter-frame prediction, and motion vectors are generated for the predicted blocks 1023 and 1024 within the picture boundary of the interpolated image 1022 after resolution conversion. In addition, the left sides of the boundary block 1013 and the boundary block 1014 are in contact with the left side of the reference picture 1011 from the inside. At this time, an offset is set only for the interpolated image 1021 that is normally referenced by the boundary blocks of the reference picture 1011 or the reference picture 1011. Therefore, the predicted image used in the inter-frame prediction of the boundary blocks 1013 and 1014 does not correspond to the interpolated image 1021, but corresponds to the interpolated image 1022 after resolution conversion obtained by resolution conversion of the interpolated image 1021 as described above. Therefore, the predicted block 1023 corresponding to the boundary block 1013 is within the picture boundary of the interpolated image 1022 after resolution conversion. Similarly, the predicted block 1024 corresponding to the boundary block 1014 is also within the picture boundary of the interpolated image 1022 after resolution conversion. In addition, one or more pixel groups of the predicted image 1012 correspond to the pixels of the rectangular region 1015 and the rectangular region 1016 formed by the pixels outside the picture boundary of the reference picture 1011. The pixels of the rectangular region 1015 and the rectangular region 1016 are generated using simple copy pixel interpolation or motion compensation pixel interpolation as described above. Specifically, the pixel group of the rectangular region 1015 located to the left of the boundary block 1013 is generated using the rectangular region 1025 within the picture boundary of the interpolated image 1022 after resolution conversion that is referenced by the boundary block 1013. Similarly, the pixel group of the rectangular region 1016 located to the left of the boundary block 1014 is generated using the rectangular region 1026 within the picture boundary of the interpolated image 1022 after resolution conversion that is referenced by the boundary block 1014. In this case, the interpolation unit 114 uses the pixels within the picture boundary of the interpolated image 1022 after resolution conversion and generates, via motion compensation pixel interpolation, the pixels outside the picture boundary of the reference picture that constitutes the predicted image 1012 of the reference picture 1011 referenced by the inter-frame prediction of the sub-block 1002. In addition, for all the boundary blocks of the reference picture 1011, the interpolation unit 114 stores the reference picture obtained by completing the interpolation outside the picture boundary of the reference picture generated by performing motion compensation pixel interpolation in the frame memory 108.
[0098] Return Figure 1 , the encoding unit 110 performs entropy encoding on the residual coefficients generated by the transform / quantization unit 105 and the prediction information input from the prediction unit 104 in units of blocks, and generates encoded data.
[0099] The entropy coding method is not particularly limited, but typical examples include Golomb coding, arithmetic coding, Huffman coding, etc. The generated coded data is output to the comprehensive coding unit 111. In the coding of the quantization parameters constituting the quantization information, an identifier indicating the difference between the quantization parameter of the sub-block to be coded and the predicted value calculated using the quantization parameters of the sub-blocks coded before the first sub-block is coded. In this embodiment, the quantization parameter coded immediately before the sub-block in the coding order is set as the predicted value, and the difference from the quantization parameter of the sub-block is calculated. However, the predicted value of the quantization parameter is not limited to this. The quantization parameter of the sub-block adjacent to the left or upper sub-block may be the predicted value, or a value such as an average value calculated from the quantization parameters of multiple sub-blocks may be used as the predicted value. In addition, in the case where the sub-block to be processed is the first sub-block in the coding order among the sub-blocks belonging to the first basic block in the basic block row, the quantization parameter of the sub-block in the basic block directly above may be used as the predicted value. Therefore, parallel processing can be performed in units of basic block rows. Note that the first basic block in the basic block row refers to the basic block with a picture boundary or a tile boundary on the left side.
[0100] The comprehensive coding unit 111 codes the resolution transformation control information. In this embodiment, the resolution transformation control information includes the horizontal pixel number and vertical pixel number of the current picture and offset information including left offset, right offset, upper offset, and lower offset. Each offset is a pixel number and includes a sign indicating positive or negative. In addition, the comprehensive coding unit 111 also codes a flag indicating the execution of motion compensation pixel interpolation. Specifically, in the case of executing motion compensation pixel interpolation, flag 1 is coded, and in the case of not executing motion compensation pixel interpolation, flag 0 is coded.
[0101] The method of coding the resolution transformation control information and the flag indicating the execution of motion compensation pixel interpolation is not particularly limited. However, Golomb coding, arithmetic coding, Huffman coding, etc. can be used. In addition, the comprehensive coding unit 111 forms a bitstream by multiplexing the coded data input from the coding and coding unit 110, etc. Finally, the bitstream is output from the terminal 112 to an external unit.
[0102] Figure 6A An example of the data structure of the bitstream including the coded resolution transformation control information is shown. The resolution transformation control information is included in any sequence, picture, or similar header. In this embodiment, as Figure 6A shown, the resolution transformation control information is included in the picture header. However, the position where the resolution transformation control information is stored is not limited to this, and as Figure 6B shown, the resolution transformation control information may be included in the sequence header. Here, reference will be made toFigure 8 Describe simple copy pixel interpolation. Figure 8 An example of simple copy pixel interpolation is shown in which pixels within the picture boundary of a reference picture are simply copied and generated outside the picture boundary of the same reference picture. The thick frame 801 represents the picture boundary of the reference picture, and the thin frame 802 represents the area outside the picture boundary of the reference picture. The circles marked with the letters of the alphabet represent pixels A to pixel T inside and outside the picture boundary. In addition, three consecutive black circles represent a thumbnail of the pixels outside the picture boundary. In simple copy pixel interpolation, pixels that are in contact with the picture boundary of the reference picture and are within the picture boundary of the reference picture are simply copied in the normal direction of the image boundary to generate pixels outside the picture boundary of the reference picture. For example, at the left picture boundary of the reference picture, the pixel located third from the top within the picture boundary is pixel C. In this case, pixels C with the same value are continuously copied in the left direction outside the picture boundary of the reference picture (i.e., the normal in the left of the reference picture) to generate pixel C outside the picture boundary of the reference picture. In a similar manner, pixels within the picture boundary are copied in the right direction at the right picture boundary of the reference picture, in the up direction at the upper picture boundary of the reference picture, and in the down direction at the lower picture boundary of the reference picture to generate pixels outside the picture boundary of the reference picture. Note that the pixels at the four corners within the picture boundary of the reference picture are used to generate the pixels in the upper left, upper right, lower left, and lower right regions of the reference picture that are not on the normals of the left, right, upper, and lower sides. Specifically, the pixels in the upper left region are generated by simply copying pixel A in the upper left within the picture boundary. In a similar manner, the pixels in the upper right region are generated by simply copying pixel K in the upper right within the picture boundary, the pixels in the lower left region are generated by simply copying pixel F in the lower left within the picture boundary, and the pixels in the lower right region are generated by simply copying pixel P in the lower right within the picture boundary. By using simple copy pixel interpolation in this way, the pixel values outside the picture boundary of each reference picture can be generated quickly.
[0103] Next, motion compensation pixel interpolation will be used Figure 9 to describe in detail. Figure 9An example of interpolation for generating pixels outside the picture boundary of the filtered image of the current picture using the pixels within the picture boundary of the filtered image retrieved from the frame memory 108 is shown. The retrieved filtered image is a filtered image obtained through loop filtering processing referenced by inter-frame prediction for each sub-block to be encoded of the current picture, or an interpolated image obtained by the interpolation unit 114 generating pixels outside the picture boundary. The thick frame 901 represents the picture boundary of the filtered image (hereinafter referred to as the current filtered image), and the area 902 indicated by the thin frame represents the area outside the picture boundary of the current filtered image. In addition, the boundary block 903 is the area that contacts the left side of the picture boundary of the current filtered image from the inside. To simplify the problem, Figure 9 An example of a 4×4 pixel square block is shown. However, this embodiment is not limited thereto. For example, an 8×8 pixel square block can be used, or an 8×16 pixel rectangular block can be used. Hereinafter, the size of the square block will be represented by N×N pixels. As Figure 9 shown, when the left side of the boundary block contacts the left side of the current filtered image from the inside, the rectangular area 904 adjacent to the outside of the boundary block of the current filtered image contacts the left side of the current filtered image from the outside. In this case, the right side of the rectangular area 904 is the same line segment as the left side of the boundary block.
[0104] Indicates Figure 9 The thick line surrounding the left figure indicates the picture boundary of the reference filtered image 905 referenced by the boundary block of the current filtered image different from the above-mentioned current filtered image retrieved from the frame memory 108 through inter-frame prediction, and the square block 906 is the area located within the reference filtered image. The interpolation unit 114 identifies the position of the square block 906 that is the reference target of the boundary block 903 based on the prediction information of the boundary block 903. The square block 906 is the block for calculating the prediction error of the boundary block 903. Therefore, the boundary block 903 and the square block 906 have the same block size. In addition, the rectangular area 907 indicated by the thick frame is the area whose left side contacts the left side of the square block 906 and the right side of the picture boundary of the reference filtered image 905, and the two sides form a distance M. Figure 9 The case where M = 5 is shown. In addition, the right side of the rectangular area 907 and the left side of the square block 906 are the same line segment. Therefore, the vertical size of the rectangular area 907 is N pixels, the same as the square block 906. Therefore, the rectangular area 907 is a rectangular area of M×N pixels.
[0105] The inside of the rectangular area 907 defined by the horizontal size M and the vertical size N is filled with the pixels within the picture boundary of the reference filtered image. In Figure 9 the example, each circle numbered from 0 to 19 is a pixel group filling the rectangular area 907 indicated by the thick line. In addition, in Figure 9In the example, the same pixel group consists of 5×4 (= 20) pixels. In this way, the pixel group of the rectangular area 907 of the recognized reference filtered image is used in the current filtered image to generate a rectangular area 904 having the same size as the rectangular area 907. The pixels included in the rectangular area 904 can be generated by simply copying the same values as the pixels of the rectangular area 907 of the reference filtered image or by applying a sharpening filter, a smoothing filter, etc.
[0106] In this way, pixels belonging to the area 902 outside the picture boundary of the current filtered image are generated by motion-compensated pixel interpolation. However, not all pixels of the area 902 are necessarily generated. Pixels having unknown values among the pixels outside the picture boundary can be generated using simple-copy pixel interpolation. For example, when the left side of the square block 906 is to the left of the left side of the reference filtered image 905, the boundary block 903 refers to outside the picture boundary of the reference filtered image 905, so the rectangular area 907 recognized within the picture boundary of the reference filtered image will not have any pixels. Since there will be no pixels corresponding to the rectangular area 904 of the current filtered image at that time, in this case, simple-copy pixel interpolation is used to generate the pixels outside the picture boundary to the left of the boundary block 903. In addition, for pixels further outside the pixel group of the rectangular area 904 generated by motion-compensated pixel interpolation, simple-copy pixel interpolation is used to generate the pixels. In addition, even when the upper side of the boundary block contacts the upper side of the current filtered image from the inside, when the right side of the boundary block contacts the right side of the current filtered image from the inside, and when the lower side of the boundary block contacts the lower side of the current filtered image from the inside, as in the case where the left side of the boundary block contacts the left side of the reference picture from the inside, pixels outside the pixel group generated by motion-compensated pixel interpolation are generated. Thus, pixels at the picture boundaries in the left, right, upper, and lower directions of the current filtered image are generated. Note that the pixels existing in the upper-left, upper-right, lower-left, and lower-right areas are used to generate the pixels outside the picture boundary by the above simple-copy pixel interpolation.
[0107] In this way, the interpolation unit 114 interpolates the pixels outside the picture boundary of the current picture using the above simple-copy pixel interpolation or motion-compensated pixel interpolation, and outputs the processing result as an interpolated image to the frame memory 108.
[0108] Figure 3 is a flowchart showing the encoding process in the image encoding device according to the present embodiment.
[0109] First, in step S301, the generation unit 103 determines the offset information for the resolution conversion unit 113 to perform resolution conversion of the reference picture. Then, the generation unit 103 sets the determined offset information as the resolution conversion control information. To encode the resolution conversion control information, the generation unit 103 outputs this information to the comprehensive encoding unit 111.
[0110] In step S302, the block segmentation unit 102 segments the input image in units of frames into basic block units.
[0111] In step S303, the prediction unit 104 performs a segmentation process on the image data in units of the basic blocks generated in step S301 and generates sub-blocks. Then, the prediction unit 104 performs a prediction process in units of the generated sub-blocks and generates prediction information including block segmentation, prediction mode, etc. and prediction image data. Then, the prediction unit 104 calculates prediction error data based on the input image data and the generated prediction image data. Specifically, the prediction unit 104 sets the prediction mode of the sub-block as inter-frame prediction. Then, the prediction unit 104 generates a prediction block, which includes pixels outside the picture boundary of the reference picture generated by the interpolation unit 114 using pixels within the picture boundary of the interpolation picture. In addition, the prediction unit 104 generates prediction error data based on the information of the prediction block and the sub-block. Alternatively, the generation unit 103 calculates the resolution conversion magnification of each reference picture that can be referenced relative to the current picture using the offset information of the current picture included in the resolution conversion control information and the offset information of each reference picture. Then, based on the above resolution conversion magnification, a resolution-converted interpolation picture obtained through the transformation of the interpolation picture is generated. In this case, the prediction unit 104 sets the prediction mode of the sub-block as inter-frame prediction. Then, the prediction unit 104 generates a prediction block, which includes pixels outside the picture boundary of the reference picture generated by the interpolation unit 114 using pixels within the picture boundary of the resolution-converted interpolation picture. In addition, the prediction unit 104 generates prediction error data based on the information of the prediction block and the sub-block.
[0112] In step S304, the transform / quantization unit 105 generates transform coefficients through orthogonal transformation of the prediction error data calculated in step S303. Then, the transform / quantization unit 105 quantizes the transform coefficients using quantization parameters and generates residual coefficients.
[0113] In step S305, the inverse quantization / inverse transform unit 106 performs inverse quantization / inverse orthogonal transformation on the residual coefficients generated in step S304 and reconstructs the prediction error. In the inverse quantization process of this step, the same parameters as those used in step S304 are used.
[0114] In step S306, the image reconstruction unit 107 reconstructs a predicted image based on the prediction information generated in step S303. The image reconstruction unit 107 also reconstructs image data from the reconstructed predicted image and the prediction error generated in step S305.
[0115] In step S307, the encoding unit 110 encodes the prediction information generated in step S303, the residual coefficients generated in step S304, and the block segmentation information together, and generates encoded data. The encoding unit 110 generates a bitstream that further includes other encoded data such as quantization parameters.
[0116] In step S308, the control unit 100 determines whether the encoding of all basic blocks within the target frame (current picture) has ended. When the control unit 100 determines that this has ended, the process proceeds to step S309. In addition, when the control unit 100 determines that there are still unprocessed basic blocks, the process returns to step S303 to encode the next basic block.
[0117] In step S309, the loop filter unit 109 performs loop filter processing on the image data reconstructed in step S306, and generates a filtered image (filtered picture).
[0118] In step S310, the resolution conversion unit 113 retrieves the filtered image of the current picture and the reference pictures referred to for inter-frame prediction of the boundary blocks of the current picture from the frame memory 108. In addition, the resolution conversion unit 113 magnifies or reduces the reference pictures using the resolution conversion control information supplied from the generation unit 103, and generates a resolution-converted image. Next, the interpolation unit 114 generates and interpolates a pixel group adjacent to the boundary block of the filtered image of the current picture and located outside the picture boundary of the filtered image of the current picture using simple copy pixel interpolation or motion compensation pixel interpolation. For each boundary block, in the case of using motion compensation pixel interpolation, pixels located within the picture boundary of the resolution-converted image are used to generate pixels outside the picture boundary. Then, the interpolation unit 114 stores the interpolated image obtained by generating and interpolating pixels outside the picture boundary in the frame memory 108, and the process ends.
[0119] The above configuration and operations can improve the generation accuracy of the predicted image in motion compensation pixel interpolation by specifically generating resolution conversion control information in step S301 and magnifying or reducing the reference pictures based on the resolution conversion control information in step S310. In addition, by using such a predicted image, the prediction error can be reduced. As a result, the data amount of the entire generated bitstream can be reduced, and the image quality of the encoded image can be improved.
[0120] Note that, in the present embodiment, pixels outside the picture boundary of the reference picture generated without using motion compensation pixel interpolation are generated using pixels within the picture boundary of the reference picture. However, such a limitation is not intended. Pixels calculated via motion compensation pixel interpolation outside the picture boundary at the position farthest from the picture boundary of the reference picture can be set as end pixels, and the end pixels can be further replicated, or an average value between the end pixels and the pixels within the picture boundary used in simple copy pixel interpolation can be replicated. Now, reference will be made to Figure 11 describe detailed examples.
[0121] Figure 11The thick frame in the center of the right figure in [Figure 0] is the reference picture 1101, and the thin frame 1102 represents the area outside the picture boundary of the reference picture 1101. The area 1103 is a boundary block that contacts the left side of the reference picture 1101 from the inside on the left side. The area 1104 is generated by motion-compensated pixel interpolation using the pixels of the picture encoded before the reference picture, and contacts the left side of the reference picture 1101 from the outside on the right side. Ten pixels numbered from 0 to 9 fill the inside of the area 1104. Among them, pixels 0 and 5 are end pixels. In this case, the pixel Y located outside the area 1104 can have the same value as the pixel V, or can have the same value as the end pixel 0. Alternatively, the average value of the pixel V and the end pixel 0 can be used. Alternatively, the average value, median value, maximum value, or minimum value of the pixel group on the normal line of the left side of the boundary block where the pixel V exists, the pixel group on the normal line, or the pixel group including the pixel V included in the area 1103 can be used to calculate the pixel Y. Alternatively, the average value, median value, maximum value, or minimum value of the pixel group included in the area 1104 (i.e., pixels 0 to 4) can be used to calculate the pixel Y. Alternatively, the average value, median value, maximum value, or minimum value of the pixel group including the area 1103 and the area 1104 in the pixel group on the normal line can be used to calculate the pixel Y. In addition, the pixel Y can be calculated using the average value, median value, maximum value, or minimum value of the value calculated using the pixel group of the area 1104 and the value calculated using the pixel group of the area 1103. Similarly, in a similar manner, the pixel Z can be calculated using the pixels on the normal line of the left side of the boundary block where the pixel W exists. As a result, by selecting the average value, median value, maximum value, or minimum value according to the characteristics of the pixel group used in the calculation, the interpolation pixels can be generated with good accuracy. In other words, in the case where the pixel values of the pixel group used in the calculation change, by using the average value, high-precision pixels that are not affected by the change can be generated. In the case of using the median value, high-precision pixels that are not affected by outliers can be generated. In addition, in the pixel group used in the calculation, when the value of the pixel inside the picture boundary or the end pixel used in the simple copy pixel interpolation is significantly smaller than other values, the maximum value is used. When the value of the pixel inside the picture boundary or the end pixel used in the simple copy pixel interpolation is significantly larger than other values, the minimum value is used. In this way, the influence of local noise can be avoided. Note that even when the upper side of the boundary block contacts the upper side of the reference picture from the inside, when the right side of the boundary block contacts the right side of the reference picture from the inside, and when the lower side of the boundary block contacts the lower side of the reference picture from the inside, pixels located outside the pixel group generated by motion-compensated pixel interpolation can be generated in the same way as when the left side of the boundary block contacts the left side of the reference picture from the inside as described above.In this way, considering the values of the pixel groups generated by motion compensation interpolation, pixels located outside the pixel groups generated by motion compensation pixel interpolation can be generated. This can reduce the discontinuity between the interpolated pixels outside the picture boundary and allow for efficient encoding of the prediction error signal of the sub-blocks to be encoded.
[0122] In addition, in the present embodiment, the pixels in the upper left, upper right, lower left, and lower right regions are generated using the pixels at the four corners within the picture boundary of the reference picture. However, such a limitation is not intended. These can be calculated using the pixels outside the pixels at the four corners of the reference picture. A detailed example will now be described using Figure 11 a description of a detailed example. Figure 11 The rectangular region as described above will not be described again. However, the interpolation image 1111 is the image referenced by the inter-frame prediction from the boundary block 1105 at the upper left corner of the reference picture 1101, and the region 1115 is the prediction block corresponding to the boundary block 1105.
[0123] Pixels 10, 12, 20, and 24 are generated by motion compensation pixel interpolation for the pixel A at the upper left corner of the reference picture 1101. In this case, the pixels 11, 21, 22, and 23 of the reference picture 1101 can be generated from the pixels 11, 21, 22, and 23 of the interpolation image 1111. Therefore, the pixels located at the upper left of the reference picture can be generated more precisely compared to the simple copying of the pixel A. Alternatively, the average value of pixels 10 and 12 can be used to generate pixel 11. Alternatively, the average value or median value of pixels 10, 12, and A can be used to generate pixel 11. In addition, the average value of pixels 20 and 24 can be used to generate pixels 21 to 23. Alternatively, the average value or median value of pixels 20, 24, and A can be used to generate pixels 21 to 23. By using the average value, pixels that reflect the state of the signals in the left and upper directions of the reference picture can be generated. In the case of using the median value, pixels that are not affected by outliers can be generated.
[0124] Pixels 50, 52, 60, and 64 are generated by motion-compensated pixel interpolation for pixel K in the upper-right corner of reference picture 1101. In this case, pixels 51, 61, 62, and 63 in reference picture 1101 can be generated from pixels 51, 61, 62, and 63 in interpolation picture 1111. Therefore, pixels in the upper-right of the reference picture can be generated more precisely than by simply copying pixel K. Alternatively, the average of pixels 50 and 52 can be used to generate pixel 51. Alternatively, the average or median of pixels 50, 52, and K can be used to generate pixel 51. Further, the average of pixels 60 and 64 can be used to generate pixels 61 to 63. Alternatively, the average or median of pixels 60, 64, and K can be used to generate pixels 61 to 63. By using the average, pixels can be generated that reflect the state of the signals in the upward and rightward directions of the reference picture. In the case of using the median, pixels that are not affected by outliers can be generated.
[0125] Pixels 30, 32, 40, and 44 are generated by motion-compensated pixel interpolation for pixel F in the lower-left corner of reference picture 1101. In this case, pixels 31, 41, 42, and 43 in reference picture 1101 can be generated from pixels 31, 41, 42, and 43 in interpolation picture 1111. Therefore, pixels in the upper-left of the reference picture can be generated more precisely than by simply copying pixel F. Alternatively, the average of pixels 30 and 32 can be used to generate pixel 31. Alternatively, the average or median of pixels 30, 32, and F can be used to generate pixel 31. Further, the average of pixels 40 and 44 can be used to generate pixels 41 to 43. Alternatively, the average or median of pixels 40, 44, and F can be used to generate pixels 41 to 43. By using the average, pixels can be generated that reflect the state of the signals in the leftward and downward directions of the reference picture. In the case of using the median, pixels that are not affected by outliers can be generated.
[0126] Pixels 70, 72, 80, and 84 are generated by motion-compensated pixel interpolation for pixel P at the lower right corner of reference picture 1101. In this case, pixels 71, 81, 82, and 83 of reference picture 1101 can be generated from pixels 71, 81, 82, and 83 of interpolation picture 1111. Therefore, compared with simply copying pixel P, pixels located at the upper left of the reference picture can be generated more accurately. Alternatively, the average value of pixels 70 and 72 can be used to generate pixel 71. Alternatively, the average value or median of pixels 70, 72, and P can be used to generate pixel 71. In addition, the average value of pixels 80 and 84 can be used to generate pixels 81 to 83. Alternatively, the average value or median of pixels 80, 84, and P can be used to generate pixels 81 to 83. By using the average value, pixels that reflect the state of the signals in the right and downward directions of the reference picture can be generated. In the case of using the median, pixels that are not affected by outliers can be generated.
[0127] In this way, considering the values of the pixel group generated by motion-compensated interpolation, pixels at the upper left, upper right, lower left, and lower right of the reference picture surrounding the pixel group generated by motion-compensated pixel interpolation can be generated. This can reduce the discontinuity between the interpolation pixels outside the picture boundary and allow efficient coding of the prediction error signal of the sub-block to be encoded.
[0128] Note that in the above-described embodiment, based on the motion information of each boundary block, pixels located outside the picture boundary of the reference picture are generated without referring to the motion information of other boundary blocks. However, this embodiment is not limited thereto. Motion-compensated pixel interpolation can be performed with reference to the motion information of the boundary blocks above / below or to the left / right of the boundary block to be processed. Now, a detailed example will be used Figure 12 to describe. Figure 12 The thick frame in the right figure of is the picture boundary of reference picture 1211, and the thin frame 1210 represents the area outside the boundary of reference picture 1211. In addition, the left sides of boundary blocks 1212, 1213, and 1214 contact the left side of the reference picture from the inside; and regions 1215, 1216, and 1217 corresponding to the respective boundary blocks are generated by motion-compensated pixel interpolation using pixels within the picture boundary of the interpolation picture 1221 after resolution conversion, which is obtained by performing resolution conversion on the picture encoded before the reference picture. The right sides of these regions contact the left side of the reference picture from the outside. For each boundary block located within reference picture 1211, rectangular regions of the interpolation picture 1221 after resolution conversion obtained by using inter-frame prediction within the picture boundary of the interpolation picture 1221 after resolution conversion are sequentially identified from the upper left to the lower right based on the prediction information of each boundary block. In Figure 12In [the above situation], the regions used in the inter-frame prediction corresponding to the boundary blocks 1212, 1213, and 1214 are represented by the boundary blocks 1222, 1223, and 1224, and the rectangular regions 1225, 1226, and 1227 adjacent to the respective blocks are used to generate the regions 1215, 1216, and 1217 which are pixel groups outside the picture region serving as the reference picture. In addition, the motion vectors of the boundary block 1212 and the boundary block 1214 are the same. In this case, the region 1223 which is the region used in the inter-frame prediction indicated by the motion information of the boundary block 1213 is not adjacent to the boundary block 1222 and the boundary block 1224. However, the motion vectors of the boundary block 1212 and the boundary block 1214 are the same, and the relative positional relationship between the boundary block 1212 and the boundary block 1214 is the same as the relative positional relationship between the boundary block 1222 and the boundary block 1224. In a scaling scenario with a set scaling window, since the objects in the scene are enlarged or reduced to the same size via resolution transformation, it can be considered that the motion vectors generated by the inter-frame prediction reflect the global motion of the entire picture. Therefore, in such a case, the region 1228 between the rectangular region 1225 and the rectangular region 1227 can be used, without using the pixels of the region 1226, to generate the pixels of the region 1216. In this way, for the left side of the reference picture, in the case of performing a correction process on the motion compensation pixel interpolation of the boundary block to be processed using the motion information of the adjacent boundary block, 1 can be encoded as a flag indicating the execution of the correction process for the left side of the reference picture, and in the case of not performing the correction process, 0 can be encoded as a flag indicating the non-execution of the correction process. Note that even in the case where the right side of the boundary block contacts the right side of the reference picture from the inside, in the case where the upper side of the boundary block contacts the upper side of the reference picture from the inside, and in the case where the lower side of the boundary block contacts the lower side of the reference picture from the inside, pixels outside the pixel group generated using motion compensation pixel interpolation can be generated in the same way as in the case where the left side of the boundary block contacts the left side of the reference picture from the inside. In this case, 1 is encoded as a flag indicating the execution of the correction process for the right side of the reference picture, and in the case of not performing the correction process, 0 is encoded as a flag indicating the non-execution of the correction process. Alternatively, 1 is encoded as a flag indicating the execution of the correction process for the upper side of the reference picture, and in the case of not performing the correction process, 0 can be encoded as a flag indicating the non-execution of the correction process. 1 can be encoded as a flag indicating the execution of the correction process for the lower side of the reference picture, and in the case of not performing the correction process, 0 can be encoded as a flag indicating the non-execution of the correction process. In this way, by controlling whether to perform the correction process on each side of the reference picture, the execution of the correction process can be limited to the sides where the interpolation accuracy needs to be improved via the correction process.By performing the processing in this manner, the interpolation accuracy of pixels interpolated outside the picture boundary is improved, and the prediction error signal of the sub-block to be encoded can be efficiently encoded.
[0129] Note that, in the present embodiment, based on the motion information of each boundary block, for all boundary blocks in contact with the inside of the picture boundary of the reference picture, motion-compensated pixel interpolation is used to generate pixels located outside the picture boundary of the reference picture. However, the present embodiment is not limited thereto. For each side around the reference picture, pixels outside the picture boundary can be generated collectively for the rectangular area between each rectangular area adjacent to the boundary blocks at the four corners. Now, reference will be made to Figure 13A and Figure 13B to describe a detailed example.
[0130] Figure 13A and Figure 13B The thick frames in Figure 13A and Figure 13B are the picture boundaries of the reference picture 1311, and the thin frames 1310 represent the areas outside the boundaries of the reference picture 1311. In addition, the left sides of the boundary blocks 1312 and 1313 are in contact with the left side of the reference picture from the inside. Regions 1314 and 1315 corresponding to each boundary block are generated by motion-compensated pixel interpolation using pixels within the picture boundary of the interpolation picture 1321 after resolution conversion, and the interpolation picture 1321 after resolution conversion is obtained by performing resolution conversion on the picture encoded before the reference picture. Regions 1314 and 1315 are in contact with the left side of the reference picture from the outside on the right. First, motion-compensated pixel interpolation is performed on the boundary blocks located at the upper left, upper right, lower left, and lower right corners. For each boundary block among the boundary blocks within the reference picture 1311 other than the boundary blocks located at the upper left, upper right, lower left, and lower right corners, based on the prediction information of each boundary block in sequence from the upper left to the lower right, a rectangular area of the interpolation picture 1321 after resolution conversion obtained by using inter-frame prediction within the picture boundary of the interpolation picture 1321 after resolution conversion is identified. In
[0131] In Figure 13AIn the case shown, the motion vectors of boundary block 1312 and boundary block 1313 are the same, and the relative positional relationship between region 1314 and region 1315 is the same as the relative positional relationship between region 1324 and region 1325. In a scaling scene with a set scaling window, since the objects in the scene are enlarged or reduced to the same size via resolution transformation, it can be considered that the motion vectors generated by inter-frame prediction reflect the global motion of the entire picture. Therefore, in such a case, region 1326 between region 1324 and region 1325 can be used to generate region 1316. In this way, in the case where an integrated interpolation process (instead of motion compensation pixel interpolation) is performed for each boundary block between boundary block 1312 and boundary block 1313 to interpolate the pixels located to the left of the picture boundary of the reference picture, 1 can be encoded as a flag indicating the execution of the integrated interpolation process, and 0 can be encoded as a flag indicating the non-execution of the integrated interpolation process in the case where the integrated interpolation process is not performed.
[0132] In Figure 13B the case shown, the motion vectors of boundary block 1312 and boundary block 1313 are the same, and the relative positional relationship between region 1314 and region 1315 is the same as the relative positional relationship between region 1324 and region 1325. In a scaling scene with a set scaling window, since the objects in the scene are enlarged or reduced to the same size via resolution transformation, it can be considered that the motion vectors generated by inter-frame prediction reflect the global motion of the entire picture. Therefore, in such a case, region 1326 between region 1324 and region 1325 can be used to generate region 1316. However, in Figure 13B the case where one or more pixel groups included in region 1326 are pixels outside the picture boundary of the interpolation image 1321 after resolution transformation. Therefore, in this case, only the pixels included in the region where region 1326 overlaps with the region of the interpolation image 1321 after resolution transformation can be copied in region 1316. In addition, for the pixels in region 1316 other than the copied pixels, the above single copied pixel interpolation can be used to interpolate the pixels. In this way, in the case where an integrated interpolation process (instead of motion compensation pixel interpolation) is performed for each boundary block between boundary block 1312 and boundary block 1313 to interpolate the pixels located to the left of the picture boundary of the reference picture, 1 can be encoded as a flag indicating the execution of the integrated interpolation process, and 0 can be encoded as a flag indicating the non-execution of the integrated interpolation process in the case where the integrated interpolation process is not performed.
[0133] Note that even when the right side of the boundary block touches the right side of the reference picture from the inside, when the upper side of the boundary block touches the upper side of the reference picture from the inside, and when the lower side of the boundary block touches the lower side of the reference picture from the inside, pixels outside the pixel group generated using motion compensation pixel interpolation can be generated in the same way as when the left side of the boundary block touches the left side of the reference picture from the inside. In this case, 1 is encoded as a flag indicating the execution of the correction process for the right side of the reference picture, and 0 is encoded as a flag indicating the non-execution of the correction process when the correction process is not executed. Alternatively, 1 can be encoded as a flag indicating the execution of the correction process for the upper side of the reference picture, and 0 can be encoded as a flag indicating the non-execution of the correction process when the correction process is not executed. 1 can be encoded as a flag indicating the execution of the correction process for the lower side of the reference picture, and 0 can be encoded as a flag indicating the non-execution of the correction process when the correction process is not executed. In this way, by controlling whether to perform the correction process on each side of the reference picture, the execution of the correction process can be limited to the sides for which the interpolation accuracy is to be improved through the correction process.
[0134] By performing the process in this way, based on the processing results of the boundary blocks (not all boundary blocks) located at the four corners of the reference picture, the execution of motion compensation pixel interpolation for other boundary blocks can be controlled, and compared with the case of performing pixel interpolation all at once, pixels outside the picture boundary of the reference picture can be generated with a smaller amount of processing.
[0135] Note that in this embodiment, image data is input in units of frames, and encoding processing is performed to generate and output a bitstream. However, the target of the encoding processing is not limited to image data. For example, feature amounts used in machine learning for object identification etc. can be input in a two-dimensional form, and encoding processing can be performed to encode the bitstream. In this way, the feature amount data used in machine learning can be encoded efficiently.
[0136] Second Embodiment
[0137] The following second embodiment is an image decoding device that decodes the encoded data (bitstream) output by the image encoding device of the above first embodiment.
[0138] Figure 2 FIG. is a block configuration diagram showing the configuration of the image decoding device according to the second embodiment. A control unit 200 that controls the entire device includes a CPU and a memory that stores programs for executing the CPU.
[0139] Terminal 201 is an input terminal for inputting an encoded bitstream. The demultiplexer decoding unit 202 demultiplexes the bitstream input via terminal 201 into information related to decoding processing, encoded data related to residual coefficients, etc. In addition, the demultiplexer decoding unit 202 decodes the encoded data present in the header of the bitstream. The demultiplexer decoding unit 202 according to the present embodiment decodes the resolution conversion control information and outputs it to the subsequent stage. The demultiplexer decoding unit 202 can be considered to perform operations opposite to those of Figure 1 the integrated encoding unit 111.
[0140] The decoding unit 203 decodes the encoded data output from the demultiplexer decoding unit 202 to obtain residual coefficients and prediction information.
[0141] The inverse quantization / inverse transform unit 204 inverse-quantizes the residual coefficients input in units of blocks and obtains a prediction error by further performing an inverse orthogonal transform.
[0142] The frame memory 206 is a memory for storing reconstructed picture image data.
[0143] The image reconstruction unit 205 reconstructs the predicted image data using the prediction information input from the decoding unit 203 and the interpolated image data input from the interpolation unit 210. Then, the image reconstruction unit 205 generates the reconstructed image data from the predicted image data and the prediction error data reconstructed by the inverse quantization / inverse transform unit 204 and outputs the reconstructed image data.
[0144] In a manner similar to Figure 1 the loop filter unit 109, the loop filter unit 207 performs a loop filter process such as deblocking filtering on the reconstructed image and outputs the filtered image.
[0145] The resolution conversion unit 209 enlarges or reduces the filtered image stored in the frame memory 206 based on the resolution conversion control information, generates a resolution-converted image, and outputs it as the resolution-converted image.
[0146] The interpolation unit 210 appropriately refers to the resolution-converted image output from the resolution conversion unit 209, the prediction information output from the decoding unit 203, and the filtered image stored in the frame memory 206, and generates interpolated image data. Then, the interpolation unit 210 outputs the generated interpolated image data to the image reconstruction unit 205. Note that the interpolation unit 210 may store the image data including the generated out-of-picture pixel information as the interpolated image in the frame memory 206. In addition, the interpolation unit 210 may retrieve the generated interpolated image from the frame memory 206 and may output the generated interpolated image to the image reconstruction unit 205.
[0147] The image data of the current frame stored in the frame memory 206 is output to an external unit via the terminal 208.
[0148] The image decoding operation in the above image decoding device will be described below. The image decoding device according to the present embodiment decodes the bitstream generated at the image encoding device according to the first embodiment. As Figure 2 shown, the control unit 200 is a processor that controls the entire image decoding device, and the bitstream input from the terminal 201 is input to the demultiplexer decoding unit 202.
[0149] The demultiplexer decoding unit 202 demultiplexes the bitstream into information related to decoding processing and encoded data related to coefficients, and decodes the encoded data present in the header of the bitstream. Specifically, the demultiplexer decoding unit 202 first decodes the resolution transformation control information from the picture header of the bitstream as Figure 6A shown. Then, the demultiplexer decoding unit 202 outputs the decoded resolution transformation control information to the resolution transformation unit 209. In addition, the demultiplexer decoding unit 202 outputs the encoded data to the decoding unit 203 in units of picture data blocks. Note that the decoded resolution transformation control information includes offset information indicating the horizontal size and vertical size of the current picture and the offset of the scaling window of the current picture. The offset information is offset information corresponding to the distances from the left, right, top, and bottom sides of the current picture. The offset with respect to each side is represented by the number of pixels. Similarly, if the offset is in the direction from the picture frame boundary to the center, the offset is represented by a positive value, and if the offset is in the direction from the frame boundary outward, the offset is represented by a negative value. For both the horizontal and vertical directions, the resolution magnification representing the magnification or reduction ratio of the reference picture with respect to the current picture is calculated. The resolution transformation magnification in the horizontal and vertical directions is obtained using the following formula.
[0150] Resolution transformation magnification in the horizontal direction = (Reference picture horizontal size - Left offset R - Right offset R) / (Current picture horizontal size - Left offset C - Left offset C)
[0151] Resolution transformation magnification in the vertical direction = (Reference picture vertical size - Top offset R - Bottom offset R) / (Current picture vertical size - Top offset C - Bottom offset C)
[0152] According to these formulas, the demultiplexer decoding unit 202 calculates the resolution transformation magnification with respect to the current picture for all reference pictures that can be referred to in the decoding of the current picture.
[0153] The decoding unit 203 decodes the encoded data and obtains residual coefficients, prediction information, and quantization parameters. Then, the decoding unit 203 outputs the residual coefficients and quantization parameters to the inverse quantization / inverse transformation unit 204, and outputs the obtained prediction information to the image reconstruction unit 205.
[0154] The inverse quantization / inverse transformation unit 204 inverse-quantizes the input residual coefficients and generates orthogonal transformation coefficients. In addition, the inverse quantization / inverse transformation unit 204 performs an inverse orthogonal transformation on the generated orthogonal transformation coefficients and generates a prediction error. Note that the quantization parameters used by the inverse quantization / inverse transformation unit 204 in the inverse quantization of each sub-block are the same as those used on the encoding side. The inverse quantization / inverse transformation unit 204 outputs the obtained prediction information to the image reconstruction unit 205.
[0155] The image reconstruction unit 205 generates a predicted image based on the interpolated image after generation of pixels outside the picture boundary input from the interpolation unit 210 and the prediction information input from the decoding unit 203. Then, the image reconstruction unit 205 reconstructs the image data according to the predicted image and the prediction error input from the inverse quantization / inverse transformation unit 204, and stores the reconstructed image data in the frame memory 206. When making a prediction in subsequent decoding of a sub-block to be decoded, the image data stored in the frame memory 206 is used by being referenced.
[0156] In this embodiment, when generating the predicted image, simple copy pixel interpolation or motion compensation pixel interpolation in inter-frame prediction together with resolution transformation of a reference picture is used to generate pixels outside the picture boundary of the reference picture. Now, reference will be made to Figure 10A and Figure 10B to describe the detailed generation method. Figure 10A and Figure 10BThe case where the decoding unit 203 decodes the sub-block 1002 using inter-frame prediction is shown. The current picture 1001 indicated by the thick frame in the right figure is the decoding target, and the reference picture 1011 indicated by the thick frame in the central figure is the picture referred to by the sub-block of the current picture. The thin frame 1010 is an area including pixels outside the picture boundary of the reference picture 1011, and the pixels outside the picture boundary are generated by simple copy pixel interpolation or motion compensation pixel interpolation. The predicted image 1012 of the reference picture 1011 represents the predicted block of the sub-block 1002 identified by the image reconstruction unit 205 based on the prediction information. The left sides of the boundary block 1013 and the boundary block 1014 contact the left side of the reference picture 1011 from the inside. In addition, the prediction modes of the boundary block 1013 and the boundary block 1014 are inter-frame prediction. In addition, the rectangular areas 1015 and 1016 are filled with pixels generated by simple copy pixel interpolation or motion compensation pixel interpolation. Note that the upper side of the boundary block may contact the upper side of the reference picture from the inside. In addition, the right side of the boundary block may contact the right side of the reference picture from the inside. In addition, the lower side of the boundary block may contact the lower side of the reference picture from the inside. In any case, similar to the case where the left side of the boundary block contacts the left side of the reference picture from the inside, the interpolation unit 210 generates the pixels outside the picture boundary of the reference picture 1011 using simple copy pixel interpolation or motion compensation pixel interpolation. In addition, in the case of using motion compensation pixel interpolation, the interpolation unit 210 uses the pixels at the picture boundary of the reference picture and generates the pixels outside the picture boundary of the reference picture that are not generated by motion compensation pixel interpolation by simple copy pixel interpolation. Then, in the motion compensation pixel interpolation of each boundary block, based on the motion information of each boundary block, for each boundary block among all the boundary blocks that contact the inside of the boundary block in the reference picture, the interpolation unit 210 generates the pixels located outside the picture boundary of the reference picture using motion compensation pixel interpolation. In addition, in the processing of each boundary block, without referring to the motion information of other boundary blocks, the interpolation unit 210 generates the pixels located outside the picture boundary of the reference picture.
[0157] In Figure 10AIn the example shown, the interpolation image 1021 is a filtered image used in the inter-frame prediction of the boundary block of the reference picture 1011 and is a picture retrieved from the frame memory 206. In addition, the pixel values within the picture boundary of the interpolation image 1021 correspond to the picture of the pixels outside the picture boundary of the reference picture 1011 generated by motion-compensated pixel interpolation. No scaling window is set for the interpolation image 1021 and the reference picture 1011, and all offsets are 0. Therefore, the resolution conversion ratio is 1 in both the horizontal and vertical directions, and resolution conversion of the interpolation image 1021 is not required in the inter-frame prediction of the boundary block 1013. As a result, the prediction block 1023 corresponding to the boundary block 1013 located within the prediction image 1012 is located within the picture boundary of the interpolation image 1021. In a similar manner, the prediction block 1024 corresponding to the boundary block 1014 located within the prediction image 1012 is located within the picture boundary of the interpolation image 1021. Therefore, the interpolation unit 210 generates the pixels of the rectangular region 1015 located outside the picture boundary of the reference picture using the pixel group of the rectangular region 1025 of the interpolation image 1021. In a similar manner, the interpolation unit 210 generates the pixels of the rectangular region 1016 located outside the picture boundary of the reference picture using the pixel group of the rectangular region 1026 of the interpolation image 1021. Therefore, the interpolation unit 210 generates the pixels outside the picture boundary of the reference picture that constitutes the prediction image 1012 of the reference picture referred to by the inter-frame prediction of the sub-block 1002 using the pixels within the picture boundary of the interpolation image 1021 via motion-compensated pixel interpolation. In addition, for all boundary blocks of the reference picture 1011, the interpolation unit 210 stores the reference picture obtained by completing the interpolation outside the picture boundary of the reference picture generated by performing motion-compensated pixel interpolation in the frame memory 206.
[0158] In Figure 10B the example shown, the offsets of the scaling windows set in the interpolation image 1021 and the reference picture 1011 are not 0. In other words, in the inter-frame prediction of the boundary block 1013, resolution conversion of the interpolation image 1021 is required. Therefore, the prediction block 1023 corresponding to the boundary block 1013 located within the prediction image 1012 is located within the picture boundary of the resolution-converted interpolation image (instead of the interpolation image 1021) obtained via resolution conversion based on the offset information of the interpolation image 1021 and the reference picture 1011. In a similar manner, the prediction block 1024 corresponding to the boundary block 1014 located within the prediction image 1012 is located within the picture boundary of the resolution-converted interpolation image 1022 (instead of the interpolation image 1021).
[0159] Now, the process of generating the interpolated image 1022 after resolution conversion from the interpolated image 1021 will be described. To simplify the description, in this example, the interpolated image 1021 is set with a scaling window, and the current picture 1001 and the reference picture 1011 are not set with a scaling window. However, it is not intended to impose such a restriction, and the current picture 1001 and the reference picture 1011 may be set with a scaling window. In Figure 10B the example of Figure 10B , the horizontal size of the reference picture 1011 and the interpolated image 1021 is set to 1920 pixels, and the vertical size is set to 1080 pixels. In addition, both the left offset and the right offset set for the interpolated image 1021 are 240 pixels, and both the upper offset and the lower offset are 135 pixels. In addition, the offsets of the four sides of the reference picture 1011 are 0 pixels. In this case, using the above calculation formula, the resolution conversion magnification is calculated to be 0.75 in both the horizontal and vertical directions. Therefore, the interpolated image 1022 after resolution conversion is an image of the interpolated image 1021 magnified by 4 / 3 times, which is the reciprocal of 0.75. In this way, based on the offset information set for the reference picture 1011 and the interpolated image 1021, the interpolated image 1022 after resolution conversion is generated from the interpolated image 1021.
[0160] In Figure 10BIn this case, the predicted image 1012 indicated by the rectangular area of the reference picture 1011 represents the image of the sub-block 1002 generated by the image reconstruction unit 205 via inter-frame prediction. In addition, the prediction modes of the boundary blocks 1013 and 1014 of the reference picture 1011 included in the predicted image 1012 are inter-frame prediction, and motion vectors are generated for the prediction blocks 1023 and 1024 within the picture boundary of the interpolated image 1022 after resolution conversion. In addition, the left sides of the boundary block 1013 and the boundary block 1014 are in contact with the left side of the reference picture 1011 from the inside. At this time, an offset is set for the interpolated image 1021 that is normally referenced by the boundary blocks of the reference picture 1011, or for the reference picture 1011. Therefore, the predicted image used in the inter-frame prediction of the boundary blocks 1013 and 1014 does not correspond to the interpolated image 1021, but corresponds to the interpolated image 1022 after resolution conversion obtained via the resolution conversion of the interpolated image 1021 as described above. Therefore, the prediction block 1023 corresponding to the boundary block 1013 is within the picture boundary of the interpolated image 1022 after resolution conversion. In a similar manner, the prediction block 1024 corresponding to the boundary block 1014 is also within the picture boundary of the interpolated image 1022 after resolution conversion. In addition, one or more pixel groups of the predicted image 1012 correspond to the pixels of the rectangular area 1015 and the rectangular area 1016 formed by the pixels outside the picture boundary of the reference picture 1011. The pixels of the rectangular area 1015 and the rectangular area 1016 are generated using simple copy pixel interpolation or motion compensation pixel interpolation as described above. Specifically, the rectangular area 1025 within the picture boundary of the interpolated image 1022 after resolution conversion, which is referenced by the boundary block 1013, is used to generate the pixel group of the rectangular area 1015 located to the left of the boundary block 1013. In a similar manner, the rectangular area 1026 within the picture boundary of the interpolated image 1022 after resolution conversion, which is referenced by the boundary block 1014, is used to generate the pixel group of the rectangular area 1016 located to the left of the boundary block 1014. The interpolation unit 210 uses the pixels within the picture boundary of the interpolated image 1022 after resolution conversion to generate, via motion compensation pixel interpolation, the pixels outside the picture boundary of the reference picture that constitutes the predicted image 1012 of the reference picture 1011 referenced by the inter-frame prediction of the sub-block 1002. In addition, for all the boundary blocks of the reference picture 1011, the interpolation unit 210 stores the reference picture obtained by completing the interpolation outside the picture boundary of the reference picture generated by performing motion compensation pixel interpolation in the frame memory 206.
[0161] With Figure 1Similar to the loop filter unit 109 in the encoding device, the loop filter unit 207 reads the reconstructed image from the frame memory 206 and performs loop filter processing such as deblocking and sample adaptive offset. The loop filter unit 207 stores the updated filtered image back in the frame memory 206.
[0162] The resolution conversion unit 209 enlarges or reduces the image stored in the frame memory 206 according to the resolution conversion control information. In addition, the resolution conversion unit 209 outputs the enlarged or reduced image as the resolution-converted image. The resolution conversion filter used in the resolution conversion is not particularly limited, and the user can input one or more resolution conversion filters, or can use a pre-specified value as the initial value. In addition, the resolution conversion unit 209 can switch between multiple resolution conversion filters according to the resolution conversion magnification calculated using the offset information and generate the resolution-converted image.
[0163] The interpolation unit 210 appropriately refers to the resolution-converted image output by the resolution conversion unit 209, the prediction information output by the decoding unit 203, and the filtered image stored in the frame memory 206, and generates image data in which pixels outside the picture boundary of the current picture are interpolated by using simple copy pixel interpolation or motion compensation pixel interpolation according to the value of the flag indicating motion compensation pixel interpolation decoded by the demultiplexer decoding unit 202. Specifically, when the decoded flag is 1, the interpolation unit 210 generates image data by performing motion compensation pixel interpolation on the pixels outside the picture boundary. When the decoded flag is 0, the interpolation unit 210 generates image data by performing simple copy pixel interpolation on the pixels outside the picture boundary. Then, the interpolation unit 210 outputs the generated image data as the interpolated image to the image reconstruction unit 205. In addition, when the generated image data includes out-of-picture pixel information of the filtered image, the interpolation unit 210 can output the image data including the generated out-of-picture pixel information as the interpolated image to the frame memory 206. The processing details of simple copy pixel interpolation and motion compensation pixel interpolation are the same as those in the image encoding device of the first embodiment, and thus will not be described.
[0164] The reconstructed image stored in the frame memory 206 is finally output from the terminal 208 to an external unit.
[0165] Figure 4 is a flowchart showing the image decoding process of the image decoding device according to the present embodiment.
[0166] In step S401, the demultiplexer decoding unit 202 decodes the encoded data of the header from the input bitstream and obtains resolution transformation control information. In addition, the demultiplexer decoding unit 202 demultiplexes the bitstream into information related to the decoding process, encoded data related to coefficients, and the like. In this embodiment, as the control information related to the resolution transformation, the horizontal size and vertical size of the current picture to be decoded and the offset information of the current picture are decoded. Specifically, the demultiplexer unit 202 decodes the horizontal pixel count and vertical pixel count of the current picture and the offset information including the left offset, right offset, top offset, and bottom offset as the resolution transformation control information. Each offset is a pixel count and includes a sign indicating positive or negative. In addition, the demultiplexer decoding unit 202 also decodes a flag indicating the execution of motion compensation pixel interpolation. In other words, in the case where motion compensation pixel interpolation is to be performed, the flag is decoded as 1, and in the case where motion compensation pixel interpolation is not performed, the flag is decoded as 0.
[0167] In step S402, the decoding unit 203 decodes the encoded data demultiplexed in step S401 and obtains block segmentation information, residual coefficients, prediction information, and quantization parameters.
[0168] In step S403, the inverse quantization / inverse transformation unit 204 performs inverse quantization on the residual coefficients in units of sub-blocks and obtains a prediction error by further performing inverse orthogonal transformation.
[0169] In step S404, the image reconstruction unit 205 generates a prediction image based on the prediction information obtained in step S402. In addition, the image reconstruction unit 205 reconstructs the image data from the generated prediction image and the prediction error generated in step S403 and stores the image data in the frame memory 206.
[0170] In step S405, the control unit 200 of the image decoding device determines whether the decoding of all blocks in the frame has been completed. If the control unit 200 determines that the decoding of all blocks has been completed, the process proceeds to step S406. If the control unit 200 determines that there are still undecoded blocks, the process returns to step S402 to perform decoding processing on that block.
[0171] In step S406, the loop filter unit 207 performs loop filter processing on the image data reconstructed in step S404, generates a filtered image, and stores the filtered image back in the frame memory 206.
[0172] In step S407, the resolution conversion unit 209 retrieves the filtered image of the current picture and the reference pictures referred to for the inter-frame prediction of the boundary blocks of the current picture from the frame memory 206. Further, the resolution conversion unit 209 enlarges or reduces the reference pictures using the resolution conversion control information supplied from the demultiplexer decoding unit 202, and generates a resolution-converted image. Next, the interpolation unit 210 generates and interpolates a pixel group adjacent to the boundary block of the filtered image of the current picture and located outside the picture boundary of the filtered image of the current picture using simple copy pixel interpolation or motion-compensated pixel interpolation. For each boundary block, in the case of using motion-compensated pixel interpolation, the interpolation unit 210 uses the pixels located within the picture boundary of the resolution-converted image to generate the pixels outside the picture boundary. Then, the interpolation unit 210 stores the interpolated image obtained by generating and interpolating the pixels outside the picture boundary in the frame memory 206, and the processing ends.
[0173] The above configuration and operation can improve the generation accuracy of the predicted image in the motion-compensated pixel interpolation by enlarging or reducing the reference pictures based on the resolution conversion control information. Further, using such a predicted image, a bitstream that can be represented with an encoding amount smaller than the prediction error signal can be decoded.
[0174] Note that, in the present embodiment, the pixels outside the picture boundary of the reference picture that are not generated using motion-compensated pixel interpolation are generated using the pixels within the picture boundary of the reference picture. However, such a limitation is not intended. The pixels calculated by the motion-compensated pixel interpolation outside the picture boundary located at the position farthest from the picture boundary of the reference picture can be set as the end pixels, and the end pixels can be further copied, or the average value between the end pixels and the pixels within the picture boundary used in the simple copy pixel interpolation can be copied. Now, a detailed example will be used Figure 11 to describe.
[0175] Figure 11The thick frame shown in the upper right figure is the reference picture 1101, and the thin frame 1102 represents the area outside the picture boundary of the reference picture 1101. The area 1103 is a boundary block that contacts the left side of the reference picture 1101 from the inside on the left. The area 1104 is generated by motion-compensated pixel interpolation using the pixels of the picture decoded before the reference picture, and contacts the left side of the reference picture 1101 from the outside on the right. Ten pixels numbered from 0 to 9 fill the inside of the area 1104. Among them, pixels 0 and 5 are end pixels. In this case, the pixel Y located outside the area 1104 may have the same value as the pixel V, or may have the same value as the end pixel 0. Alternatively, the average value of the pixel V and the end pixel 0 can be used. Alternatively, the average value, median value, maximum value, or minimum value of the pixel group on the normal line of the left side of the boundary block where the pixel V exists, the pixel group on the normal line, or the pixel group including the pixel V included in the area 1103 can be used to calculate the pixel Y. Alternatively, the average value, median value, maximum value, or minimum value of the pixel group included in the area 1104 (i.e., pixels 0 to 4) can be used to calculate the pixel Y. Alternatively, the average value, median value, maximum value, or minimum value of the pixel group including the area 1103 and the area 1104 in the pixel group on the normal line can be used to calculate the pixel Y. In addition, the pixel Y can be calculated using the average value, median value, maximum value, or minimum value of the value calculated using the pixel group of the area 1104 and the value calculated using the pixel group of the area 1103. Similarly, in a similar manner, the pixel Z can be calculated using the pixels on the normal line of the left side of the boundary block where the pixel W exists. As a result, by selecting the average value, median value, maximum value, or minimum value according to the characteristics of the pixel group used in the calculation, interpolation pixels can be generated with good accuracy. In other words, when the pixel values of the pixel group used in the calculation change, high-precision pixels that are not affected by the change can be generated by using the average value. When the median value is used, pixels that are not affected by outliers can be generated. In addition, in the pixel group used in the calculation, when the value of the pixel inside the picture boundary or the end pixel used in simple copy pixel interpolation is significantly smaller than other values, the maximum value is used. When the value of the pixel inside the picture boundary or the end pixel used in simple copy pixel interpolation is significantly larger than other values, the minimum value is used. In this way, the influence of local noise can be avoided. Note that even when the upper side of the boundary block contacts the upper side of the reference picture from the inside, when the right side of the boundary block contacts the right side of the reference picture from the inside, and when the lower side of the boundary block contacts the lower side of the reference picture from the inside, pixels located outside the pixel group generated by motion-compensated pixel interpolation can be generated in the same way as when the left side of the boundary block contacts the left side of the reference picture from the inside as described above.In this manner, considering the values of the pixel groups generated by motion compensation interpolation, pixels located outside the pixel groups generated by motion compensation pixel interpolation can be generated. This can reduce the discontinuity between the interpolated pixels outside the picture boundary and allows decoding of a bitstream that is represented with an amount of coding smaller than the prediction error signal of the sub-blocks to be coded.
[0176] Furthermore, in the present embodiment, the pixels in the upper left, upper right, lower left, and lower right regions are generated using the pixels at four corners within the picture boundary of the reference picture. However, such a limitation is not intended. These can be calculated using the pixels outside the pixels at the four corners of the reference picture. A detailed example will now be described using Figure 11 a description. Figure 11 The rectangular region of
[0177] is as described above and will not be described again. However, the interpolation image 1111 is an interpolation image referenced by the boundary block 1105 located at the upper left corner of the reference picture 1101 via inter-frame prediction, and the region 1115 is a prediction block corresponding to the boundary block 1105.
[0178] Pixels 50, 52, 60, and 64 are generated by motion-compensated pixel interpolation for pixel K in the upper-right corner of reference picture 1101. In this case, pixels 51, 61, 62, and 63 in reference picture 1101 can be generated from pixels 51, 61, 62, and 63 in interpolation picture 1111. Therefore, pixels in the upper-right of the reference picture can be generated more precisely than by simply copying pixel K. Alternatively, the average value of pixels 50 and 52 can be used to generate pixel 51. Alternatively, the average value or median value of pixels 50, 52, and K can be used to generate pixel 51. In addition, the average value of pixels 60 and 64 can be used to generate pixels 61 to 63. Alternatively, the average value or median value of pixels 60, 64, and K can be used to generate pixels 61 to 63. By using the average value, pixels can be generated that reflect the state of the signals in the upward and rightward directions of the reference picture. In the case of using the median value, pixels that are not affected by outliers can be generated.
[0179] Pixels 30, 32, 40, and 44 are generated by motion-compensated pixel interpolation for pixel F in the lower-left corner of reference picture 1101. In this case, pixels 31, 41, 42, and 43 in reference picture 1101 can be generated from pixels 31, 41, 42, and 43 in interpolation picture 1111. Therefore, pixels in the upper-left of the reference picture can be generated more precisely than by simply copying pixel F. Alternatively, the average value of pixels 30 and 32 can be used to generate pixel 31. Alternatively, the average value or median value of pixels 30, 32, and F can be used to generate pixel 31. In addition, the average value of pixels 40 and 44 can be used to generate pixels 41 to 43. Alternatively, the average value or median value of pixels 40, 44, and F can be used to generate pixels 41 to 43. By using the average value, pixels can be generated that reflect the state of the signals in the leftward and downward directions of the reference picture. In the case of using the median value, pixels that are not affected by outliers can be generated.
[0180] Pixels 70, 72, 80, and 84 are generated by motion-compensated pixel interpolation for pixel P at the lower right corner of reference picture 1101. In this case, pixels 71, 81, 82, and 83 of reference picture 1101 can be generated from pixels 71, 81, 82, and 83 of interpolation picture 1111. Therefore, pixels located at the upper left of the reference picture can be generated more precisely than by simply copying pixel P. Alternatively, the average value of pixels 70 and 72 can be used to generate pixel 71. Alternatively, the average value or median value of pixels 70, 72, and P can be used to generate pixel 71. In addition, the average value of pixels 80 and 84 can be used to generate pixels 81 to 83. Alternatively, the average value or median value of pixels 80, 84, and P can be used to generate pixels 81 to 83. By using the average value, pixels reflecting the state of signals in the right and downward directions of the reference picture can be generated. In the case of using the median value, pixels not affected by outliers can be generated.
[0181] In this way, considering the values of the pixel group generated by motion-compensated interpolation, pixels at the upper left, upper right, lower left, and lower right of the reference picture surrounding the pixel group generated by motion-compensated pixel interpolation can be generated. This can reduce the discontinuity between interpolation pixels outside the picture boundary and allow decoding of a bitstream represented by an encoding amount smaller than the prediction error signal of the sub-block to be encoded.
[0182] Note that in the above-described embodiment, based on the motion information of each boundary block, pixels outside the picture boundary of the reference picture are generated without referring to the motion information of other boundary blocks. However, the present embodiment is not limited to this. Motion-compensated pixel interpolation can be performed with reference to the motion information of boundary blocks above / below or to the left / right of the boundary block to be processed. Now, a detailed example will be used Figure 12 to describe. The thick frame surrounding reference picture 1211 is the picture boundary of the picture, and the thin frame 1210 represents the area outside the boundary of reference picture 1211. In addition, the left sides of boundary blocks 1212, 1213, and 1214 contact the left side of the reference picture from the inside. Regions 1215, 1216, and 1217 corresponding to boundary blocks 1212, 1213, and 1214 are generated by motion-compensated pixel interpolation using pixels within the picture boundary of interpolation picture 1221 after resolution conversion, which is obtained by performing resolution conversion on a picture decoded before the reference picture. The right sides of these regions contact the left side of the reference picture from the outside. For each boundary block within reference picture 1211, rectangular regions of interpolation picture 1221 after resolution conversion obtained by using inter-frame prediction within the picture boundary of interpolation picture 1221 after resolution conversion are sequentially identified from the upper left to the lower right based on the prediction information of each boundary block. In Figure 12In [the figure], the regions used in the inter-frame prediction corresponding to the boundary blocks 1212, 1213, and 1214 are represented by the boundary blocks 1222, 1223, and 1224, and the rectangular regions 1225, 1226, and 1227 adjacent to the respective blocks are used to generate the regions 1215, 1216, and 1217 which are pixel groups outside the picture region serving as the reference picture. In addition, the motion vectors of the boundary block 1212 and the boundary block 1214 are the same. In this case, the region 1223 which is the region used in the inter-frame prediction indicated by the motion information of the boundary block 1213 is not adjacent to the boundary block 1222 and the boundary block 1224. However, the motion vectors of the boundary block 1212 and the boundary block 1214 are the same, and the relative positional relationship between the boundary block 1212 and the boundary block 1214 is the same as the relative positional relationship between the boundary block 1222 and the boundary block 1224. In a scaling scenario with a set scaling window, since the objects in the scene are enlarged or reduced to the same size via resolution transformation, it can be considered that the motion vectors generated by the inter-frame prediction reflect the global motion of the entire picture. Therefore, in this case, the region 1228 between the rectangular region 1225 and the rectangular region 1227 can be used, without using the pixels of the region 1226, to generate the pixels of the region 1216. In this way, for the left side of the reference picture, in the case of performing a correction process on the motion compensation pixel interpolation of the boundary block to be processed using the motion information of the adjacent boundary blocks, 1 can be decoded as a flag indicating the execution of the correction process for the left side of the reference picture, and in the case of not performing the correction process, 0 can be decoded as a flag indicating the non-execution of the correction process.
[0183] Note that even when the upper side of the boundary block touches the upper side of the reference picture from the inside, when the right side of the boundary block touches the right side of the reference picture from the inside, and when the lower side of the boundary block touches the lower side of the reference picture from the inside, pixels located outside the pixel group generated using motion-compensated pixel interpolation can be generated in the same way as when the left side of the boundary block touches the left side of the reference picture from the inside. 1 can be decoded as a flag indicating the execution of the correction process for the right side of the reference picture, and 0 can be decoded as a flag indicating the non-execution of the correction process when the correction process is not performed. Alternatively, 1 can be decoded as a flag indicating the execution of the correction process for the upper side of the reference picture, and 0 can be decoded as a flag indicating the non-execution of the correction process when the correction process is not performed. 1 can be decoded as a flag indicating the execution of the correction process for the lower side of the reference picture, and 0 can be decoded as a flag indicating the non-execution of the correction process when the correction process is not performed. In this way, by controlling whether to perform the correction process on each side of the reference picture, the execution of the correction process can be limited to the sides for which the interpolation accuracy is to be improved via the correction process. By performing the process in this way, the interpolation accuracy of the pixels interpolated outside the picture boundary is improved, and a bitstream that can be represented with a coding amount smaller than the prediction error signal of the sub-block to be coded can be decoded.
[0184] Note that in this embodiment, based on the motion information of each boundary block, for all boundary blocks that touch the inside of the picture boundary of the reference picture, pixels located outside the picture boundary of the reference picture are generated using motion-compensated pixel interpolation. However, such a limitation is not intended. For each side around the reference picture, pixels outside the picture boundary can be generated collectively for the rectangular regions between the rectangular regions adjacent to the boundary blocks at the four corners. Now, Figure 13A and Figure 13B a detailed example will be described. Figure 13A and Figure 13BThe thick frame around the periphery of the reference picture 1311 in is the picture boundary, and the thin frame 1310 represents the area outside the boundary of the reference picture 1311. In addition, the left sides of the boundary blocks 1312 and 1313 are in contact with the left side of the reference picture from the inside; and regions 1314 and 1315 corresponding to the respective boundary blocks are generated by motion-compensated pixel interpolation using pixels within the picture boundary of the interpolated picture 1321 after resolution conversion, which is obtained by performing resolution conversion on a picture decoded before the reference picture. The right sides of these regions are in contact with the left side of the reference picture from the outside. First, motion-compensated pixel interpolation is performed on the boundary blocks located at the upper left, upper right, lower left, and lower right corners. For each of the boundary blocks within the reference picture 1311 other than the boundary blocks located at the upper left, upper right, lower left, and lower right corners, rectangular regions of the interpolated picture 1321 after resolution conversion obtained by using inter-frame prediction within the picture boundary of the interpolated picture 1321 after resolution conversion are sequentially identified from the upper left to the lower right based on the prediction information of each boundary block. In Figure 13A and Figure 13B In and , the regions used in the inter-frame prediction corresponding to the boundary blocks 1312 and 1313 are represented by the prediction blocks 1322 and 1323, and the rectangular regions 1324 and 1325 adjacent to the respective prediction blocks are used to generate the regions 1314 and 1315, which are pixel groups outside the picture area of the reference picture. In addition, the motion vectors of the boundary block 1312 and the region 1314 are the same.
[0185] In Figure 13A In the case shown in , the motion vectors of the boundary block 1312 and the boundary block 1313 are the same, and the relative positional relationship between the region 1314 and the region 1315 is the same as the relative positional relationship between the region 1324 and the region 1325. In a zooming scenario with a set zoom window, since the objects in the scene are enlarged or reduced to the same size via resolution conversion, it can be considered that the motion vectors generated by inter-frame prediction reflect the global motion of the entire picture. Therefore, in such a case, the region 1326 between the region 1324 and the region 1325 can be used to generate the region 1316. In this way, in the case where an interpolation process (instead of motion-compensated pixel interpolation) is performed on each boundary block between the boundary block 1312 and the boundary block 1313 to interpolate the pixels on the left side of the picture boundary of the reference picture, 1 can be decoded as a flag indicating the execution of the interpolation process, and 0 can be decoded as a flag indicating the non-execution of the interpolation process in the case where the interpolation process is not performed.
[0186] In Figure 13BIn the case shown, the motion vectors of boundary block 1312 and boundary block 1313 are the same, and the relative positional relationship between region 1314 and region 1315 is the same as the relative positional relationship between region 1324 and region 1325. In a scaling scenario with a set scaling window, since objects in the scene are magnified or reduced to the same size via resolution transformation, the motion vectors generated by inter-frame prediction can be considered to reflect the global motion of the entire picture. Therefore, in such a case, region 1326 between region 1324 and region 1325 can be used to generate region 1316. However, in Figure 13B the case where one or more pixel groups included in region 1326 are pixels outside the picture boundary of the interpolation image 1321 after resolution transformation. Therefore, in this case, only the pixels included in the region where region 1326 and the region of the interpolation image 1321 after resolution transformation overlap can be copied in region 1316. In addition, for the pixels in region 1316 other than the copied pixels, the above single-copy pixel interpolation can be used to interpolate the pixels. In this way, in the case where an integrated interpolation process (instead of motion compensation pixel interpolation) is performed for each boundary block between boundary block 1312 and boundary block 1313 to interpolate the pixels located to the left of the picture boundary of the reference picture, 1 can be decoded as a flag indicating the execution of the integrated interpolation process, and 0 can be decoded as a flag indicating the non-execution of the integrated interpolation process in the case where the integrated interpolation process is not performed.
[0187] Note that even in the case where the upper side of the boundary block contacts the upper side of the reference picture from the inside, in the case where the right side of the boundary block contacts the right side of the reference picture from the inside, and in the case where the lower side of the boundary block contacts the lower side of the reference picture from the inside, pixels located outside the pixel group generated using motion compensation pixel interpolation can be generated in the same way as in the case where the left side of the boundary block contacts the left side of the reference picture from the inside. In addition, in this case, 1 can be decoded as a flag indicating the execution of the correction process for the right side of the reference picture, and 0 can be decoded as a flag indicating the non-execution of the correction process in the case where the correction process is not performed. Alternatively, 1 can be decoded as a flag indicating the execution of the correction process for the upper side of the reference picture, and 0 can be decoded as a flag indicating the non-execution of the correction process in the case where the correction process is not performed. 1 can be decoded as a flag indicating the execution of the correction process for the lower side of the reference picture, and 0 can be decoded as a flag indicating the non-execution of the correction process in the case where the correction process is not performed. In this way, by controlling whether to perform the correction process for each side of the reference picture, the execution of the correction process can be limited to the sides for which the interpolation accuracy is to be improved via the correction process.
[0188] By performing the processing in this manner, the execution of motion compensation pixel interpolation for other boundary blocks can be controlled based on the processing results of the boundary blocks located at the four corners of the reference picture (instead of all boundary blocks), and compared with the case of performing pixel interpolation all at once, pixels outside the picture boundary of the reference picture can be generated with a smaller amount of processing.
[0189] In addition, in the present embodiment, image data is input in units of frames, and the bitstream generated by performing encoding processing is decoded. However, the target of the decoding processing is not limited to the bitstream obtained by encoding image data. For example, feature amounts used in machine learning such as for object recognition can be input in a two-dimensional form, and decoding processing can be performed to decode the generated bitstream. In this way, the bitstream in which the feature amount data used in machine learning is efficiently encoded can be decoded.
[0190] Third Embodiment
[0191] In the first embodiment described above, the image encoding device includes Figure 1 the hardware shown. In addition, in the second embodiment, the image decoding device includes Figure 2 the hardware shown. However, the processing performed by Figure 1 and Figure 2 the processing units shown can be implemented via a computer program. The third embodiment described below is an example implemented using a computer program.
[0192] Figure 5 is a block diagram showing an example of the hardware configuration of a computer that can be applied to the devices in the first embodiment and the second embodiment described above.
[0193] The CPU 501 controls the entire computer using the computer programs and data stored in the RAM 502 and the ROM 503, and performs the respective processes described above as an image processing device according to the above embodiments. In other words, the CPU 501 serves as Figure 1 and Figure 2 the processing units shown.
[0194] The RAM 502 includes areas for temporarily storing computer programs and data loaded from the external storage device 506, data obtained from external units via the I / F (interface) 507, and the like. In addition, the RAM 502 includes work areas used when the CPU 501 performs various types of processing. In other words, the RAM 502 can be allocated functions as a frame memory, and other types of areas and the like can be provided appropriately.
[0195] The ROM 503 stores the setting data, boot program, etc. of this computer. The operation unit 504 includes a keyboard, a mouse, etc., and can input various types of instructions to the CPU 501 when operated by the user of this computer. The display unit 505 displays the processing result from the CPU 501. In addition, the display unit 505 includes, for example, a liquid crystal display.
[0196] The external storage device 506 is a large-capacity information storage device typified by a hard disk drive device. The external storage device 506 stores an OS (operating system) and computer programs for enabling the CPU 501 to implement Figure 1 and Figure 2 the functions of the units shown. In addition, the external storage device 506 can store the image data to be processed.
[0197] The computer programs and data stored in the external storage device 506 are appropriately loaded onto the RAM 502 under the control of the CPU 501 and become the processing targets of the CPU 501. At the I / F 507, networks such as a LAN and the Internet and other devices such as a projection device or a display device can be connected, and this computer can obtain and transfer various information via the I / F 507. 508 represents a bus connecting the above units.
[0198] The operations performed through the above configuration are controlled and executed with the CPU 501 as the center using the above flowchart for the above operations.
[0199] According to the present invention, a storage medium storing computer program codes for implementing the above functions is supplied to the system, and the system can read out and execute the computer program codes. In this case, the codes of the computer program read out from the storage medium implement the functions of the above embodiments, and the storage medium storing the codes of the computer program forms a part of the present invention. This also includes the case where, based on the instructions in the program codes, an operating system (OS), etc. running on a computer executes part or all of the actual processing, and the above functions are implemented via this processing.
[0200] The following mode can also be implemented. This includes an example of writing the computer program codes read out from the storage medium into the memory of a function expansion card inserted into the computer or a function expansion unit connected to the computer. Then, based on the instructions in the computer program codes, a CPU, etc. in the function expansion card or the function expansion unit executes part or all of the actual processing to implement the above functions.
[0201] When the present invention is applied to the above storage medium, the storage medium stores computer program codes corresponding to the above flowchart.
[0202] The present invention can be used in an encoding device / decoding device for encoding and decoding still images and videos. Specifically, the present invention can be applied to an encoding method and a decoding method for generating pixels outside the picture boundary for use in inter-frame prediction.
[0203] Other embodiments
[0204] The present invention can be implemented by the following processing: supplying a program for implementing one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and causing one or more processors in a computer of the system or device to read and execute the program. The present invention can also be implemented by a circuit (e.g., ASIC) for implementing one or more functions.
[0205] The present invention is not limited to the above-described embodiments, and various changes and modifications can be made within the spirit and scope of the present invention. Therefore, the appended claims are presented to inform the public of the scope of the present invention.
[0206] This application claims the priority of Japanese Patent Application No. 2022-165024 filed on Oct. 13, 2022, which is incorporated herein by reference.
Claims
1. An image encoding device, characterized in that it comprises: a prediction component for generating a predicted image for a target block in a first frame to be encoded by referring to a second frame encoded before the first frame; an encoding component for encoding a prediction error of the target block with respect to the predicted image; an interpolation component for interpolating pixels outside the boundary of the second frame using pixels of a third frame encoded before the second frame when referring to pixels outside the boundary of the second frame; and a transformation component for changing the resolution of a frame before the first frame.
2. The image encoding device according to claim 1, characterized in that the frame before the first frame is the second frame or the third frame.
3. The image encoding device according to claim 1, characterized in that the frame before the first frame includes the second frame and the third frame.
4. An image encoding method, characterized in that it comprises: generating a predicted image for a target block in a first frame to be encoded by referring to a second frame encoded before the first frame; encoding a prediction error of the target block with respect to the predicted image; interpolating pixels outside the boundary of the second frame using pixels of a third frame encoded before the second frame when referring to pixels outside the boundary of the second frame; and changing the resolution of a frame before the first frame.
5. A program which is read and executed by a computer to cause the computer to execute the image encoding method according to claim 4.
6. An image decoding device, characterized in that it comprises: a decoding component for obtaining prediction error data of a target block in a first frame by decoding encoded data; a generation component for generating a predicted image by referring to a second frame decoded before the first frame; a reconstruction component for reconstructing an image of the target block based on the prediction error data and the predicted image; an interpolation component for interpolating pixels outside the boundary of the second frame using pixels of a third frame decoded before the second frame when referring to pixels outside the boundary of the second frame; and a transformation component for changing the resolution of a frame before the first frame.
7. The image decoding device according to claim 6, characterized in that the frame before the first frame is the second frame or the third frame.
8. The image decoding device according to claim 6, characterized in that the frame before the first frame includes the second frame and the third frame.
9. An image decoding method, characterized in that it comprises: obtaining prediction error data of a target block in a first frame by decoding encoded data; generating a predicted image by referring to a second frame decoded before the first frame; reconstructing an image of the target block based on the prediction error data and the predicted image; interpolating pixels outside the boundary of the second frame using pixels of a third frame decoded before the second frame when referring to pixels outside the boundary of the second frame; and changing the resolution of a frame before the first frame.
10. A program, which is read and executed by a computer to cause the computer to execute the image decoding method according to claim 9.
Citation Information
Patent Citations
Image encoding method, image decoding method, image encoder and image decoder
JP2018050085A
Polishing method and polishing device
JP2022165024A