Image decoding device, image decoding method and program

By performing detailed processing before the BDOF processing and controlling the application of the BDOF processing using the calculated information, the problem that the BDOF processing time cannot be shortened in both software and hardware implementation in the prior art, and the processing time in both is reduced.

CN116193145BActive Publication Date: 2025-05-23KDDI CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310232541.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-11
Filing Date
2020-03-04
Publication Date
2025-05-23
Estimated Expiration
2040-03-04

Smart Images

  • Figure CN116193145B_ABST
    Figure CN116193145B_ABST
Patent Text Reader

Abstract

An image decoding device (200) comprises: a motion vector decoding unit (241B) configured to decode a motion vector from encoded data; a thinning unit (241C) configured to perform a thinning process to correct the decoded motion vector; and a prediction signal generating unit (241D) configured to generate a prediction signal based on the corrected motion vector output from the thinning unit (241C), the prediction signal generating unit (241D) configured to determine whether to apply BDOF processing to each block based on information calculated during the thinning process.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application with application number 202080019761.5, application date March 4, 2020, and invention name “Image decoding device, image decoding method and program”. Technical Field

[0002] The present invention relates to an image decoding device, an image decoding method and a program. Background Art

[0003] In the past, regarding BODF (Bi-Directional Optical Flow) technology, the following technology has been disclosed: in order to shorten the execution time of the software, when the absolute difference sum of the pixel values ​​between two reference images used in the BDOF processing is calculated, and if such absolute difference sum is less than a predetermined threshold, the BDOF processing in the block is skipped (for example, refer to non-patent document 1).

[0004] On the other hand, from the perspective of reducing processing delays during hardware implementation, a technology for deleting skip processing of BDOF based on the above-mentioned calculation of the absolute value difference sum is also disclosed (for example, refer to Non-Patent Document 2).

[0005] Prior art literature

[0006] Non-patent literature

[0007] Non-patent document 1: Versatile Video Coding (Draft 4), JVET-M1001

[0008] Non-patent literature 2: CE9-related: BDOF buffer reduction and enabling VPDU based application, JVET-M0890 Summary of the invention

[0009] Problems to be solved by the invention

[0010] However, for example, in the technology disclosed in Non-Patent Document 1, there is a problem that, while the execution time can be shortened when such technology is implemented by software, the execution time increases when it is implemented by hardware.

[0011] On the other hand, the technology disclosed in Non-Patent Document 2 has a problem that the execution time in hardware can be shortened, but the execution time in software increases.

[0012] Therefore, in the above-mentioned conventional technology, there is a problem that the processing time cannot be shortened both when implementing the software and when implementing the hardware.

[0013] Therefore, the present invention is completed in view of the above-mentioned problems, and its purpose is to provide the following image decoding device, image decoding method and program: use the information calculated during the refinement processing performed before the BDOF processing to control whether the BDOF processing is applied or not, thereby, from the perspective of hardware implementation, the processing amount can be reduced by continuing to use the values ​​that have been calculated, and from the perspective of software implementation, the processing time can be shortened by reducing the number of blocks to which the BDOF processing is applied.

[0014] Means for solving problems

[0015] The gist of a first feature of the present invention is an image decoding device comprising: a motion vector decoding unit configured to decode a motion vector from encoded data; a thinning unit configured to perform a thinning process for correcting the decoded motion vector; and a prediction signal generating unit configured to generate a prediction signal based on the corrected motion vector output from the thinning unit, the prediction signal generating unit configured to determine whether to apply BDOF processing to each block based on information calculated during the thinning process.

[0016] The gist of a second feature of the present invention is an image decoding device comprising: a motion vector decoding unit configured to decode a motion vector from encoded data; a refinement unit configured to perform refinement processing of correcting the decoded motion vector; and a prediction signal generation unit configured to generate a prediction signal based on the corrected motion vector output from the refinement unit, the prediction signal generation unit configured to apply BDOF processing when an application condition is satisfied, the application condition being the following condition: the motion vector is encoded in Symmetric MVD mode, and the size of a differential motion vector transmitted in the Symmetric MVD mode is within a preset threshold.

[0017] The gist of the third feature of the present invention is that it comprises: step A of decoding a motion vector from encoded data; step B of performing a refinement process of correcting the decoded motion vector; and step C of generating a prediction signal based on the corrected motion vector output from the refinement unit, wherein in step C, it is determined whether to apply BDOF processing to each block based on information calculated during the refinement process.

[0018] The gist of the fourth feature of the present invention is a program used in an image decoding device, which causes a computer to execute the following steps: step A, decoding a motion vector from encoded data; step B, performing a thinning process to correct the decoded motion vector; and step C, generating a prediction signal based on the corrected motion vector output from the thinning unit, in which step C, based on information calculated during the thinning process, it is determined whether to apply BDOF processing to each block.

[0019] Effects of the Invention

[0020] According to the present invention, it is possible to provide an image decoding device, an image decoding method, and a program that can reduce the amount of processing related to BDOF processing when implemented by hardware or software. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 FIG. 1 is a diagram showing an example of a configuration of an image processing system 10 according to an embodiment.

[0022] Figure 2 FIG. 1 is a diagram showing an example of functional blocks of an image encoding device 100 according to an embodiment.

[0023] Figure 3 FIG. 1 is a diagram showing an example of a functional block of the inter prediction unit 111 of the image encoding device 100 according to an embodiment.

[0024] Figure 4 1 is a flowchart showing an example of a processing procedure of the refinement unit 111C of the inter prediction unit 111 of the moving picture decoding device 30 according to one embodiment.

[0025] Figure 5 1 is a flowchart showing an example of a processing procedure of the prediction signal generation unit 111D of the inter prediction unit 111 of the moving picture decoding device 30 according to one embodiment.

[0026] Figure 6 FIG. 1 is a diagram showing an example of a functional block of the in-loop filter processing unit 150 of the image encoding device 100 according to an embodiment.

[0027] Figure 7 This is a diagram for explaining an example of determination performed by the boundary strength determination unit 153 of the in-loop filter processing unit 150 of the image encoding device 100 according to one embodiment.

[0028] Figure 8 This is a diagram showing an example of functional blocks of an image decoding device 200 according to an embodiment.

[0029] Fig. 9This is a diagram showing an example of a functional block of the inter prediction unit 241 of the image decoding device 200 according to one embodiment.

[0030] Fig.10 FIG. 1 is a diagram showing an example of functional blocks of the in-loop filter processing unit 250 of the image decoding device 200 according to one embodiment. DETAILED DESCRIPTION

[0031] Hereinafter, the embodiments of the present invention will be described with reference to the accompanying drawings. In addition, the constituent elements in the following embodiments may be appropriately replaced with existing constituent elements, etc., and various modifications including combinations with other existing constituent elements may be performed. Therefore, the description of the following embodiments is not intended to limit the content of the invention described in the claims.

[0032] (First Embodiment)

[0033] Below, refer to Figures 1 to 10 An image processing system 10 according to a first embodiment of the present invention will be described. Figure 1 1 is a diagram showing an image processing system 10 according to the present embodiment.

[0034] like Figure 1 As shown, the image processing system 10 includes an image encoding device 100 and an image decoding device 200 .

[0035] The image encoding device 100 is configured to generate encoded data by encoding an input image signal. The image decoding device 200 is configured to generate an output image signal by decoding the encoded data.

[0036] Such coded data may be transmitted from the image coding apparatus 100 to the image decoding apparatus 200 via a transmission path. Alternatively, the coded data may be provided from the image coding apparatus 100 to the image decoding apparatus 200 after being stored in a storage medium.

[0037] (Image Coding Device 100)

[0038] Below, refer to Figure 2 The image encoding device 100 according to this embodiment will be described. Figure 2 1 is a diagram showing an example of functional blocks of the image encoding device 100 according to the present embodiment.

[0039] like Figure 2 As shown, the image encoding device 100 has an inter-frame prediction unit 111, an intra-frame prediction unit 112, a subtractor 121, an adder 122, a transform and quantization unit 131, an inverse transform and inverse quantization unit 132, an encoding unit 140, an in-loop filter processing unit 150 and a frame buffer 160.

[0040] The inter-frame prediction unit 111 is configured to generate a prediction signal through inter-frame prediction.

[0041] Specifically, the inter-frame prediction unit 111 is configured to determine a reference block included in a reference frame by comparing a coding target frame (hereinafter referred to as a target frame) with a reference frame stored in a frame buffer 160 , and determine a motion vector for the determined reference block.

[0042] In addition, the inter prediction unit 111 is configured to generate a prediction signal included in the prediction block for each prediction block based on the reference block and the motion vector. The inter prediction unit 111 is configured to output the prediction signal to the subtractor 121 and the adder 122. The reference frame is a frame different from the target frame.

[0043] The intra prediction unit 112 is configured to generate a prediction signal through intra-frame prediction.

[0044] Specifically, the intra prediction section 112 is configured to determine a reference block included in the target frame and generate a prediction signal for each prediction block based on the determined reference block. In addition, the intra prediction section 112 is configured to output the prediction signal to the subtractor 121 and the adder 122.

[0045] The reference block is a block that is referenced for predicting a target block (hereinafter referred to as a target block). For example, the reference block is a block adjacent to the target block.

[0046] The subtractor 121 is configured to subtract the prediction signal from the input image signal and output the prediction residual signal to the transform and quantization section 131. The subtractor 121 is configured to generate a prediction residual signal which is a difference between the prediction signal generated by intra prediction or inter prediction and the input image signal.

[0047] The adder 122 is configured to add the prediction signal and the prediction residual signal output from the inverse transform and inverse quantization section 132 to generate a pre-filtering decoded signal, and output such a pre-filtering decoded signal to the intra prediction section 112 and the in-loop filter processing section 150 .

[0048] The decoded signal before filtering constitutes a reference block used in the intra prediction unit 112 .

[0049] The transform and quantization unit 131 is configured to perform a transform process on the prediction residual signal and obtain coefficient level values. Furthermore, the transform and quantization unit 131 may also be configured to perform quantization of the coefficient level values.

[0050] The transform process is a process of transforming the prediction residual signal into a frequency component signal. In such a transform process, a basic pattern (transformation matrix) corresponding to a discrete cosine transform (DCT) or a basic pattern (transformation matrix) corresponding to a discrete sine transform (DST) can be used.

[0051] The inverse transform and inverse quantization unit 132 is configured to perform an inverse transform process on the coefficient level values ​​output from the transform and quantization unit 131. The inverse transform and inverse quantization unit 132 may also be configured to perform an inverse quantization on the coefficient level values ​​before the inverse transform process.

[0052] Among them, the inverse transform processing and inverse quantization are performed according to the reverse process of the transform processing and quantization performed by the transform and quantization unit 131.

[0053] The encoding section 140 is configured to encode the coefficient level values ​​output from the transform and quantization section 131 and output the encoded data.

[0054] Among them, for example, the encoding is entropy encoding that assigns codes of different lengths based on the probability of occurrence of coefficient level values.

[0055] In addition, the encoding unit 140 is configured to encode control data used in the decoding process in addition to the coefficient level values.

[0056] The control data may include size data such as the coding block (CU: Coding Unit) size, the prediction block (PU: Prediction Unit) size, and the transform block (TU: Transform Unit) size.

[0057] The in-loop filter processing section 150 is configured to perform a filter process on the decoded signal before the filter process output from the adder 122 , and output the decoded signal after the filter process to the frame buffer 160 .

[0058] Among them, for example, the filtering process is a deblocking filtering process that reduces distortion generated at a boundary portion of a block (coding block, prediction block, or transform block).

[0059] The frame buffer 160 is configured to accumulate reference frames used in the inter prediction section 111 .

[0060] The decoded signal after filtering constitutes a reference frame used in the inter-frame prediction unit 111 .

[0061] (Inter-frame prediction unit 111)

[0062] Below, refer to Figure 3The inter-frame prediction unit 111 of the image encoding device 100 according to this embodiment will be described. Figure 3 FIG. 1 is a diagram showing an example of functional blocks of the inter prediction unit 111 of the image encoding device 100 according to the present embodiment.

[0063] like Figure 3 As shown, the inter prediction section 111 includes a motion vector search section 111A, a motion vector encoding section 111B, a refinement section 111C, and a prediction signal generation section 111D.

[0064] The inter prediction section 111 is an example of a prediction section configured to generate a prediction signal included in a prediction block based on a motion vector.

[0065] The motion vector search unit 111A is configured to determine a reference block included in the reference frame by comparing the target frame with the reference frame, and to search for a motion vector for the determined reference block.

[0066] In addition, the above search process is performed on multiple reference frame candidates, and the reference frame and motion vector used for prediction in the prediction block are determined. A maximum of two reference frames and two motion vectors can be used for one block. The case where only one set of reference frames and motion vectors is used for one block is called single prediction, and the case where two sets of reference frames and motion vectors are used is called dual prediction. Hereinafter, the first set is called L0, and the second set is called L1.

[0067] Furthermore, the motion vector search unit 111A is configured to determine a coding method for the reference frame and the motion vector. In addition to the usual method of transmitting the reference frame and the motion vector information separately, the coding method includes the merge mode described below, the Symmetric MVD mode described in Non-Patent Document 2, and the like.

[0068] In addition, regarding the method of searching for a motion vector, the method of determining a reference frame, and the method of determining the encoding method of a reference frame and a motion vector, known methods can be used, and thus detailed descriptions thereof are omitted.

[0069] The motion vector encoding section 111B is configured to encode information of the reference frame and the motion vector also determined by the motion vector search section 111A using the encoding method determined by the motion vector search section 111A.

[0070] In the case where the encoding method of the block is the merge mode, a merge list for the block is first created. The merge list is a list that lists multiple combinations of reference frames and motion vectors. An index is assigned to each combination, and instead of encoding the information of the reference frame and motion vector separately, only the above index is encoded and transmitted to the decoding side. The method for creating the merge list is shared by the encoding side and the decoding side, so that the information of the reference frame and motion vector can be decoded only from the index information on the decoding side. Regarding the method for creating the merge list, a known method can be used, so its detailed description is omitted.

[0071] The Symmetric MVD mode is an encoding method that can be used only when bi-prediction is performed in the block. In the Symmetric MVD mode, only the L0 motion vector (differential motion vector) of the two (L0, L1) reference frames and two (L0, L1) motion vectors to be transmitted to the decoding side is encoded. The remaining L1 motion vector and the information of the two reference frames are uniquely determined on the encoding side and the decoding side respectively by a predetermined method.

[0072] Regarding the encoding of motion vector information, first, a predicted value of an encoding target motion vector, namely a predicted motion vector, is generated, and a differential motion vector, namely a difference value between the predicted motion vector and a motion vector actually to be encoded, is encoded.

[0073] In the Symmetric MVD mode, a vector obtained by inverting the code of the coded L0 difference motion vector is used as the L1 difference motion vector. As a specific method, for example, the method described in Non-Patent Document 2 can be used.

[0074] The thinning section 111C is configured to perform a thinning process (eg, DMVR) of correcting the motion vector encoded by the motion vector encoding section 111B.

[0075] Specifically, the refinement section 111C is configured to perform refinement processing by setting a search range based on a reference position determined by a motion vector encoded by the motion vector encoding section 111B, determining a correction reference position having a minimum predetermined cost from the search range, and correcting the motion vector based on the correction reference position.

[0076] Figure 4 11 is a flowchart showing an example of the processing procedure of the refinement unit 111C.

[0077] like Figure 4As shown, in step S41, the refinement unit 111C determines whether the predetermined conditions for applying the refinement process are satisfied. If all the predetermined conditions are satisfied, the process proceeds to step S42. If any one of the predetermined conditions is not satisfied, the process proceeds to step S45 to end the refinement process.

[0078] The predetermined condition includes a condition that the block is a block for bi-prediction.

[0079] Further, the predetermined condition may also include a condition that the motion vector is encoded in a merge mode.

[0080] In addition, the predetermined condition may also include a condition that the motion vector is encoded in the Symmetric MVD mode.

[0081] Furthermore, the predetermined condition may also include the following condition: the motion vector is encoded in the Symmetric MVD mode, and the size of the differential motion vector (MVD of L0) transmitted in the Symmetric MVD mode is within a preset threshold.

[0082] The size of the differential motion vector may be defined, for example, by the absolute values ​​of the horizontal and vertical components of the differential motion vector.

[0083] As such a threshold value, different threshold values ​​may be used for the horizontal and vertical components of the motion vector, respectively, or a common threshold value may be used for the horizontal and vertical components. In addition, the value of the threshold value may be set to 0. In this case, it means that the predicted motion vector and the motion vector to be encoded are the same value. In addition, the threshold value may also be defined by a minimum value and a maximum value. In this case, as a predetermined condition, the value or absolute value of the differential motion vector is greater than a predetermined minimum value and less than a maximum value.

[0084] In addition, the predetermined condition may also include a condition that the motion vector is encoded in a merge mode or a Symmetric MVD mode. Similarly, the predetermined condition may also include the following condition: the motion vector is encoded in a merge mode or a Symmetric MVD mode, and the size of the differential motion vector transmitted when encoded in the Symmetric MVD mode is within a preset threshold.

[0085] In step S42 , the thinning unit 111C generates a search image based on the motion vector encoded by the motion vector encoding unit 111B and the information of the reference frame.

[0086] In the case where the motion vector points to a non-integer pixel position, the thinning unit 111C applies a filter to the pixel value of the reference frame to interpolate the pixel at the non-integer pixel position. At this time, the thinning unit 111C can reduce the amount of calculation by using an interpolation filter with a smaller number of taps than the interpolation filter used in the prediction signal generation unit 111D described later. For example, the thinning unit 111C can interpolate the pixel value at the non-integer pixel position by bilinear interpolation.

[0087] In step S43, the thinning unit 111C performs a search at integer pixel accuracy using the search image generated in step S42. Integer pixel accuracy means that only points at integer pixel intervals are searched based on the motion vector encoded by the motion vector encoding unit 111B.

[0088] The refinement unit 111C determines the corrected motion vector at the integer pixel interval position by searching in step S42. As a search method, a known method can be used. For example, the refinement unit 111C may search by searching only for a point that is a combination obtained by inverting the code of the differential motion vector on the L0 side and the L1 side. However, as a result of the search in step S43, it is possible that the value is the same as the motion vector before the search.

[0089] In step S44, the thinning unit 111C uses the corrected motion vector at the integer pixel accuracy determined in step S43 as an initial value to perform a motion vector search at the non-integer pixel accuracy. As a method for searching a motion vector, a known method can be used.

[0090] Alternatively, the refinement unit 111C may not actually perform a search, but may use the result of step S43 as an input and determine a vector with non-integer pixel accuracy using a parameter model such as parabola fitting.

[0091] The thinning unit 111C determines the corrected motion vector at the non-integer pixel precision in step S44, and then proceeds to step S45 to terminate the thinning process. Here, for convenience, the expression of the corrected motion vector at the non-integer pixel precision is used, but depending on the search result of step S44, the result may be the same value as the motion vector at the integer pixel precision obtained in step S43.

[0092] The thinning unit 111C may divide a block larger than a predetermined threshold into smaller sub-blocks and perform thinning processing on each sub-block. For example, the thinning unit 111C may pre-set the unit for thinning processing to 16×16 pixels, and when the size of a block in the horizontal direction or the vertical direction is larger than 16 pixels, divide the block into smaller blocks of 16 pixels. In this case, as the motion vector used as the reference for thinning processing, the motion vector of the block encoded by the motion vector encoding unit 111B is used for all sub-blocks in the same block.

[0093] When processing each sub-block, the refinement unit 111C may perform Figure 4 In addition, the refinement unit 111C may perform only Figure 4 Specifically, the refinement unit 111C may perform a Figure 4 The processing of steps S41 and S42 is performed on each sub-block, and only the processing of steps S43 and S44 is performed on each sub-block.

[0094] The prediction signal generation section 111D is configured to generate a prediction signal based on the corrected motion vector output from the thinning section 111C.

[0095] As will be described later, the prediction signal generation unit 111D is configured to determine whether to perform the BDOF process on each block based on information (eg, search cost) calculated during the above-mentioned thinning process.

[0096] Specifically, the prediction signal generation unit 111D is configured to generate a prediction signal based on the motion vector encoded by the motion vector encoding unit 111B when the motion vector is not corrected. On the other hand, the prediction signal generation unit 111D is configured to generate a prediction signal based on the motion vector corrected by the refinement unit 111C when the motion vector is corrected.

[0097] Figure 5 2 is a flowchart showing an example of a processing procedure of the prediction signal generating unit 111D.

[0098] However, when the refinement unit 111C performs refinement processing in units of sub-blocks, the processing of the prediction signal generation unit 111D is also performed in units of sub-blocks. In this case, the word "block" in the following description can be replaced with "sub-block" as appropriate.

[0099] like Figure 5 As shown, in step S51, the prediction signal generation unit 111D generates a prediction signal.

[0100] Specifically, the prediction signal generation unit 111D receives the motion vector encoded by the motion vector encoding unit 111B or the motion vector encoded by the thinning unit 111C as input, and when the position pointed to by such a motion vector is a non-integer pixel position, applies a filter to the pixel value of the reference frame to interpolate the pixel of the non-integer pixel position. As for the specific filter, a horizontally and vertically separable filter with a maximum of 8 taps disclosed in Non-Patent Document 3 can be applied.

[0101] When the block is a bi-predicted block, a prediction signal based on a first (hereinafter referred to as L0) reference frame and motion vector and a prediction signal based on a second (hereinafter referred to as L1) reference frame and motion vector are both generated.

[0102] In step S52 , the prediction signal generation unit 111D checks whether a BDOF (Bi-Directional Optical Flow) application condition described later is satisfied.

[0103] As such an application condition, the condition described in non-patent document 3 can be applied. The application condition includes at least the condition that the block is a block for bi-prediction. In addition, the application condition may also include the condition that the motion vector of the block is not encoded in Symmetric MVD mode as described in non-patent document 1.

[0104] In addition, the application condition may also include the following condition: the motion vector of the block is not encoded in the Symmetric MVD mode, or the size of the differential motion vector transmitted when encoded in the Symmetric MVD mode is within a preset threshold. The size of the differential motion vector can be determined by the same method as in the above step S41. The value of the threshold may also be set to 0 as in the above step S41.

[0105] If the application condition is not satisfied, the process proceeds to step S55 and ends the process. At this time, the prediction signal generation unit 111D outputs the prediction signal generated in step S51 as the final prediction signal.

[0106] On the other hand, when all the application conditions are satisfied, the present process proceeds to step S53. In step S53, the present process actually determines whether to execute the BDOF process of step S54 for the blocks satisfying the application conditions.

[0107] For example, the prediction signal generating unit 111D calculates the sum of absolute value differences between the prediction signal of L0 and the prediction signal of L1, and determines not to perform the BDOF process when the sum of absolute value differences is equal to or smaller than a predetermined threshold.

[0108] However, for the block subjected to the thinning process by the thinning unit 111C, the prediction signal generation unit 111D may use the result of the thinning process for determining whether or not to apply the BDOF.

[0109] For example, as a result of the thinning process, when the difference between the motion vectors before and after the correction is less than a predetermined threshold, the prediction signal generation unit 111D can determine that BDOF is not applied. When such thresholds are set to "0" for both the horizontal and vertical direction components, the result of the thinning process is equivalent to determining that BDOF is not applied when the motion vector is unchanged from before the correction.

[0110] The prediction signal generator 111D may use the search cost calculated during the above-mentioned thinning process (eg, the sum of absolute value differences between the pixel values ​​of the reference block on the L0 side and the pixel values ​​of the reference block on the L1 side) to determine whether to apply BDOF.

[0111] In addition, the following uses the absolute value difference sum as the search cost as an example, but other indicators can also be used as the search cost. For example, any indicator value used to determine the similarity between image signals, such as the absolute value difference sum or square error sum between signals after removing the local average value, can be used.

[0112] For example, in the integer pixel position search in step S43, when the absolute value difference sum of the search point with the smallest search cost (absolute value difference sum) is smaller than a predetermined threshold, the prediction signal generation unit 111D may determine not to apply BDOF.

[0113] The prediction signal generator 111D may determine whether to apply the BFOF process by combining the following methods: a method using the change in motion vector before and after the thinning process; and a method using the search cost of the thinning process.

[0114] For example, when the difference in motion vectors before and after the thinning process is less than a predetermined threshold and the search cost of the thinning process is less than a predetermined threshold, the predicted signal generating unit 111D may determine not to apply the BDOF process.

[0115] When the threshold of the difference between the motion vectors before and after refinement is set to 0, the search cost is determined as the sum of the absolute value differences between the reference blocks pointed to by the motion vector before refinement (= motion vector after refinement).

[0116] Alternatively, the prediction signal generator 111D may make a determination using a method based on the result of the thinning process in the block on which the thinning process has been performed, and may make a determination using a method based on the absolute value difference sum in other blocks.

[0117] In addition, the prediction signal generation unit 111D may adopt a configuration in which, as described above, the process of recalculating the absolute value difference sum of the prediction signal on the L0 side and the prediction signal on the L1 side is not performed, and only the information obtained from the result of the thinning process is used to determine whether to apply BDOF. In this case, in step S53, the prediction signal generation unit 111D determines to always apply BDOF to the block that has not been subjected to the thinning process.

[0118] According to such a configuration, in this case, it is not necessary to perform the calculation process of the absolute value difference sum in the prediction signal generating unit 111D, so the processing amount and processing delay can be reduced from the viewpoint of hardware implementation.

[0119] In addition, according to such a configuration, from the perspective of software implementation, by using the result of the thinning process so as not to perform the BDOF process on a block where the effect of the BDOF process is estimated to be low, it is possible to maintain encoding efficiency and shorten the processing time of the entire image.

[0120] Furthermore, the determination process itself using the result of the above-mentioned thinning process is executed inside the thinning unit 111C, and information indicating the result is transmitted to the prediction signal generating unit 111D, whereby the prediction signal generating unit 111D can also determine whether to apply the BDOF process.

[0121] For example, as described above, the values ​​of the motion vector and the search cost before and after the refinement processing are determined, and a flag is prepared in advance. The flag is "1" when the condition for not applying BDOF is met, and is "0" when the condition for not applying BDOF is not met and the refinement processing is not applied. The prediction signal generation unit 111D can refer to the value of such a flag to determine whether BDOF is applied or not.

[0122] In addition, for the sake of convenience, step S52 and step S53 are described as different steps here, but the determination in step S52 and step S53 may be performed at the same time.

[0123] In the above determination, for blocks determined not to be BDOF-applied, the present process proceeds to step S55. For other blocks, the present process proceeds to step S54.

[0124] In step S54, the prediction signal generation unit 111D performs BDOF processing. The BDOF processing itself can use a known method, so detailed description is omitted. After the BDOF processing is performed, the process proceeds to step S55 and ends the process.

[0125] (In-loop filter processing unit 150)

[0126] Hereinafter, the in-loop filter processing unit 150 according to the present embodiment will be described. Figure 6 It is a diagram showing the in-loop filter processing unit 150 according to the present embodiment.

[0127] like Figure 6 As shown, the in-loop filter processing unit 150 includes a target block boundary detection unit 151 , an adjacent block boundary detection unit 152 , a boundary strength determination unit 153 , a filter determination unit 154 , and a filter processing unit 155 .

[0128] Among them, the structure with "A" added at the end is a structure related to the deblocking filtering processing of the block boundary in the vertical direction, and the structure with "B" added at the end is a structure related to the deblocking filtering processing of the block boundary in the horizontal direction.

[0129] Hereinafter, an example will be given of a case where the deblocking filtering process is performed on the block boundary in the vertical direction and then the deblocking filtering process is performed on the block boundary in the horizontal direction.

[0130] As described above, the deblocking filter process can be applied to a coding block, a prediction block, or a transform block. In addition, it can also be applied to sub-blocks obtained by dividing the above blocks. That is, the target block and the adjacent block can be a coding block, a prediction block, a transform block, or sub-blocks obtained by dividing them.

[0131] The definition of a sub-block includes the sub-block described as a processing unit of the refinement unit 111C and the prediction signal generation unit 111D. When a deblocking filter is applied to a sub-block, a block in the following description can be replaced with a sub-block as appropriate.

[0132] Since the deblocking filtering process for the block boundary in the vertical direction and the deblocking filtering process for the block boundary in the horizontal direction are the same process, the deblocking filtering process for the block boundary in the vertical direction will be described below.

[0133] The target block boundary detection section 151A is configured to detect the boundary of the target block based on control data indicating the block size of the target block.

[0134] The adjacent block boundary detection section 152A is configured to detect the boundary of adjacent blocks based on control data indicating the block sizes of the adjacent blocks.

[0135] The boundary strength determination unit 153A is configured to determine the boundary strength of the block boundary between the target block and the adjacent block.

[0136] In addition, the boundary strength determination unit 153A may be configured to determine the boundary strength of the block boundary based on control data indicating whether the target block and the adjacent block are intra-prediction blocks.

[0137] For example, Figure 7As shown, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of the block boundary is "2" when at least one of the target block and the adjacent block is an intra-frame prediction block (i.e., when at least one of the blocks on both sides of the block boundary is an intra-frame prediction block).

[0138] In addition, the boundary strength determination unit 153A may also be configured to determine the boundary strength of the block boundary based on control data indicating whether the target block and the adjacent block contain non-zero (zero) orthogonal transform coefficients and whether the block boundary is a transform block boundary.

[0139] For example, Figure 7 As shown, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of the block boundary is "1" when at least one of the target block and the adjacent block contains a non-zero orthogonal transform coefficient and the block boundary is the boundary of the transform block (i.e., when at least one of the blocks on both sides of the block boundary contains a non-zero transform coefficient and is a TU boundary).

[0140] In addition, the boundary strength determination unit 153A may be configured to determine the boundary strength of the block boundary based on control data indicating whether the absolute value of the difference between the motion vectors of the target block and the adjacent block is greater than a threshold value (eg, 1 pixel).

[0141] For example, Figure 7 As shown, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of the block boundary is "1" when the absolute value of the difference between the motion vectors of the target block and the adjacent block is greater than a threshold value (e.g., 1 pixel) (i.e., when the absolute value of the difference between the motion vectors of the blocks on both sides of the block boundary is greater than a threshold value (e.g., 1 pixel)).

[0142] In addition, the boundary strength determination unit 153A may be configured to determine the boundary strength of the block boundary based on control data indicating whether the reference blocks referred to in the prediction of the motion vectors of the target block and the adjacent block are different.

[0143] For example, Figure 7 As shown, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of the block boundary is "1" when the reference blocks referenced in the prediction of the motion vectors of the target block and the adjacent block are different (i.e., when the reference images in the blocks on both sides of the block boundary are different).

[0144] The boundary strength determination unit 153A may also be configured to determine the boundary strength of the block boundary based on control data indicating whether the number of motion vectors of the target block and the adjacent block is different.

[0145] For example, Figure 7As shown, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of the block boundary is "1" when the number of motion vectors of the target block and the adjacent block is different (i.e., when the number of motion vectors in the blocks on both sides of the block boundary is different).

[0146] The boundary strength determination section 153A may also be configured to determine the boundary strength of the block boundary according to whether the thinning process performed by the thinning section 111C is applied to the target block and the adjacent blocks.

[0147] For example, Figure 7 As shown, the boundary strength determination section 153A may also be configured to determine that the boundary strength of the block boundary is "1" when both the target block and the adjacent block are blocks to which the thinning process is applied by the thinning section 111C.

[0148] The boundary strength determination unit 153A may also be configured as follows: Figure 4 When all the predetermined conditions in step S41 are satisfied, it is judged that "thinning processing is applied" in the block. In addition, a Figure 4 The boundary strength determination unit 153A may be configured to determine whether or not to apply the thinning process based on the value of the flag, by using a flag as the determination result of step S41.

[0149] In addition, the boundary strength determination unit 153A may be configured to determine that the boundary strength of the block boundary is “1” when at least one of the target block and the adjacent block is a block to which the thinning process performed by the thinning unit 111C is applied.

[0150] Alternatively, the boundary strength determination unit 153A may be configured to determine that the boundary strength of a block boundary is “1” when the boundary is a sub-block boundary being thinned by the thinning unit 111C.

[0151] Furthermore, the boundary strength determination unit 153A may be configured to determine that the boundary strength of the block boundary is “1” when DMVR is applied to at least one of the blocks on both sides of the block boundary.

[0152] For example, Figure 7 As shown, the boundary strength determination unit 153A may be configured to determine that the boundary strength of the block boundary is "0" when none of the above conditions are satisfied.

[0153] Furthermore, the larger the value of the boundary strength is, the higher the possibility that a large block distortion will be generated at the block boundary is.

[0154] The above-mentioned boundary strength determination method can be determined using a method common to the brightness signal and the color difference signal, or can be determined using partially different conditions. For example, the conditions related to the above-mentioned thinning process can be applied to both the brightness signal and the color difference signal, or can be applied only to the brightness signal or only to the color difference signal.

[0155] In addition, a header called SPS (Sequence Parameter Set) or PPS (Picture Parameter Set) may include a flag for controlling whether or not to consider the result of thinning processing when determining the boundary strength.

[0156] The filter determination unit 154A is configured to determine the type of filtering process (eg, deblocking filtering process) to be applied to a block boundary.

[0157] For example, the filter determination unit 154A may be configured to determine whether to apply filtering to block boundaries, and which filtering to apply, weak filtering or strong filtering, based on the boundary strength of the block boundaries, quantization parameters contained in the target block and the adjacent blocks, and the like.

[0158] The filter determination unit 154A may be configured to determine not to apply filtering processing when the boundary strength of the block boundary is “0”.

[0159] The filter processing unit 155A is configured to perform processing on the image before deblocking based on the decision of the filter decision unit 154A. The processing on the image before deblocking is non-filter processing, weak filter processing, strong filter processing, or the like.

[0160] (Image Decoding Device 200)

[0161] Below, refer to Figure 8 The image decoding device 200 according to this embodiment is described. Figure 8 This is a diagram showing an example of functional blocks of the image decoding device 200 according to the present embodiment.

[0162] like Figure 8 As shown, the image decoding device 200 includes a decoding unit 210 , an inverse transform and inverse quantization unit 220 , an adder 230 , an inter prediction unit 241 , an intra prediction unit 242 , an in-loop filter processing unit 250 , and a frame buffer 260 .

[0163] The decoding section 210 is configured to decode the encoded data generated by the image encoding device 100 and decode the coefficient level values.

[0164] Here, for example, the decoding is entropy decoding which is a process opposite to the entropy encoding performed by the encoding unit 140 .

[0165] In addition, the decoding unit 210 may also be configured to obtain the control data by decoding the encoded data.

[0166] In addition, as described above, the control data may also include size data such as the coding block size, the prediction block size, and the transform block size.

[0167] The inverse transform and inverse quantization unit 220 is configured to perform an inverse transform process on the coefficient level values ​​output from the decoding unit 210. The inverse transform and inverse quantization unit 220 may also be configured to perform an inverse quantization process on the coefficient level values ​​before the inverse transform process.

[0168] Among them, the inverse transform processing and inverse quantization are performed according to the reverse process of the transform processing and quantization performed by the transform and quantization unit 131.

[0169] The adder 230 is configured to add the prediction signal and the prediction residual signal output from the inverse transform and inverse quantization section 220 to generate a pre-filtering decoded signal, and output the pre-filtering decoded signal to the intra prediction section 242 and the in-loop filter processing section 250 .

[0170] Among them, the decoded signal before filtering constitutes a reference block used in the intra-frame prediction unit 242.

[0171] The inter-frame prediction unit 241 is configured to generate a prediction signal by inter-frame prediction, similarly to the inter-frame prediction unit 111 .

[0172] Specifically, the inter prediction section 241 is configured to generate a prediction signal for each prediction block based on a motion vector decoded from the encoded data and a reference signal included in a reference frame. The inter prediction section 241 is configured to output the prediction signal to the adder 230 .

[0173] The intra prediction unit 242 is configured to generate a prediction signal by intra-frame prediction, similarly to the intra prediction unit 112 .

[0174] Specifically, the intra prediction unit 242 is configured to determine a reference block included in the target frame, and generate a prediction signal for each prediction block based on the determined reference block. The intra prediction unit 242 is configured to output the prediction signal to the adder 230.

[0175] Similar to the in-loop filter processing unit 150 , the in-loop filter processing unit 250 is configured to perform filtering on the pre-filtering decoded signal output from the adder 230 and output the post-filtering decoded signal to the frame buffer 260 .

[0176] Among them, for example, the filtering process is a deblocking filtering process that reduces distortion generated at a boundary portion of a block (a coding block, a prediction block, a transform block, or a sub-block obtained by dividing these blocks).

[0177] The frame buffer 260 , similarly to the frame buffer 160 , is configured to accumulate reference frames used in the inter prediction unit 241 .

[0178] The decoded signal after filtering constitutes a reference frame used in the inter-frame prediction unit 241 .

[0179] (Inter-frame prediction unit 241)

[0180] Below, refer to Fig. 9 The inter-frame prediction unit 241 according to this embodiment will be described. Fig. 9 1 is a diagram showing an example of functional blocks of the inter prediction unit 241 according to the present embodiment.

[0181] like Fig. 9 As shown, the inter prediction section 241 includes a motion vector decoding section 241B, a refinement section 241C, and a prediction signal generation section 241D.

[0182] The inter prediction section 241 is an example of a prediction section configured to generate a prediction signal included in a prediction block based on a motion vector.

[0183] The motion vector decoding section 241B is configured to acquire a motion vector by decoding the control data received from the image encoding device 100 .

[0184] The thinning unit 241C is configured to perform thinning processing of the corrected motion vector similarly to the thinning unit 111C.

[0185] The prediction signal generating unit 241D is configured to generate a prediction signal based on a motion vector, similarly to the prediction signal generating unit 111D.

[0186] (In-loop filter processing unit 250)

[0187] Hereinafter, the in-loop filter processing unit 250 according to the present embodiment will be described. Fig.10 It is a diagram showing the in-loop filter processing unit 250 according to the present embodiment.

[0188] like Fig.10 As shown, the in-loop filter processing unit 250 includes a target block boundary detection unit 251 , an adjacent block boundary detection unit 252 , a boundary strength determination unit 253 , a filter determination unit 254 , and a filter processing unit 255 .

[0189] Among them, the structure with "A" added at the end is a structure related to the deblocking filtering processing of the block boundary in the vertical direction, and the structure with "B" added at the end is a structure related to the deblocking filtering processing of the block boundary in the horizontal direction.

[0190] Here, an example is given of a case where the deblocking filtering process for the block boundary in the vertical direction is performed and then the deblocking filtering process for the block boundary in the horizontal direction is performed.

[0191] As described above, the deblocking filter process can be applied to a coding block, a prediction block, or a transform block. In addition, it can also be applied to sub-blocks obtained by dividing the above blocks. That is, the target block and the adjacent block can be a coding block, a prediction block, a transform block, or sub-blocks obtained by dividing them.

[0192] Since the deblocking filtering process for the block boundary in the vertical direction and the deblocking filtering process for the block boundary in the horizontal direction are the same process, the deblocking filtering process for the block boundary in the vertical direction will be described below.

[0193] The target block boundary detection unit 251A is configured to detect the boundary of the target block based on the control data indicating the block size of the target block, similarly to the target block boundary detection unit 151A.

[0194] Similar to the adjacent block boundary detection unit 152A, the adjacent block boundary detection unit 252A is configured to detect the boundary between adjacent blocks based on the control data indicating the block size of the adjacent blocks.

[0195] The boundary strength determination unit 253A is configured to determine the boundary strength of the block boundary between the target block and the adjacent block, similarly to the boundary strength determination unit 153A. The method of determining the boundary strength of the block boundary is as described above.

[0196] The filter determination unit 254A is configured to determine the type of deblocking filter processing to be applied to the block boundary, similarly to the filter determination unit 154A. The method of determining the type of deblocking filter processing is as described above.

[0197] The filter processing unit 255A is configured to process the image before deblocking based on the decision of the filter decision unit 254A, similarly to the filter processing unit 155A. The processing of the image before deblocking is non-filter processing, weak filter processing, strong filter processing, or the like.

[0198] According to the image encoding device 100 and the image decoding device 200 of the present embodiment, when determining the boundary strength, the boundary strength determination units 153 and 253 consider whether the thinning process performed by the thinning units 111C and 241C is applied to the blocks adjacent to the boundary.

[0199] For example, as described above, when the thinning process is applied to at least one of two blocks adjacent to the boundary, the boundary strength of the boundary is set to 1.

[0200] For a boundary whose boundary strength is “1” or higher, the filter decision units 154 and 254 decide whether to apply a deblocking filter to the block boundary and the type of deblocking filter in consideration of parameters such as a quantization parameter.

[0201] By adopting this structure, even if the value of the refined motion vector cannot be used in the above-mentioned boundary strength judgment due to hardware implementation limitations, the deblocking filter can be appropriately applied to the boundary of the block subjected to the refinement processing, and the block noise can be suppressed to improve the subjective image quality.

[0202] There are many methods for determining whether to apply a deblocking filter and determining the boundary strength. For example, Patent Document 1 discloses a technique for omitting the application of a deblocking filter based on syntax information such as whether the block is in skip mode. In addition, for example, Patent Document 2 discloses a technique for omitting the application of a deblocking filter using a quantization parameter.

[0203] However, none of them considers whether or not to apply the refinement process. One of the problems to be solved by the present invention is a problem unique to the refinement process, that is, in the refinement process of correcting the value of the motion vector decoded from the syntax on the decoding side, when the value of the corrected motion vector cannot be used for determining the application of the deblocking filter, the application determination cannot be properly performed.

[0204] Therefore, the problem cannot be solved by the methods of Patent Documents 1 and 2 that do not consider whether to apply the thinning process. On the other hand, the methods of Patent Documents 1 and 2 can be combined with the application determination of the deblocking filter of the present invention.

[0205] According to the image encoding device 100 and the image decoding device 200 of the present embodiment, as a condition for executing the thinning process in the thinning units 111C and 241C, whether or not the motion vector serving as a reference of the process is encoded in the Symmetric MVD mode is considered.

[0206] In such refinement processing, by searching only for points where the absolute values ​​of differential motion vectors on the L0 side and the L1 side are the same and the codes are reversed, it is possible to obtain the same motion vector as when the motion vector is transmitted in the Symmetric MVD mode.

[0207] Therefore, when the motion vector originally transmitted is obtained through the above-mentioned refinement processing, the differential motion vector transmitted in the Symmetric MVD mode is reduced as much as possible, thereby reducing the amount of code related to the differential motion vector.

[0208] In particular, in the case of an encoding method that only transmits a flag when the differential motion vector is a specific value (such as 0), and directly encodes the differential value for other values, the effect of reducing the amount of code may be large even if the difference in the motion vector corrected in such refinement processing is small.

[0209] In addition, as the execution condition of the above-mentioned thinning processing, the condition that the motion vector serving as the basis for such thinning processing is encoded in Symmetric MVD mode and the value of the difference motion vector is less than a predetermined threshold is set, thereby implicitly switching whether to execute the thinning processing according to the value of the difference motion vector.

[0210] Similarly, in the application conditions of the BDOF processing in the prediction signal generating units 111D and 241D, the condition that the motion vector of the block is encoded in the Symmetric MVD mode and the value of the differential motion vector is below a predetermined threshold is used, thereby implicitly switching whether to apply the BDOF processing according to the value of the differential motion vector.

[0211] According to the image encoding device 100 and the image decoding device 200 of the present embodiment, the result of the above-mentioned thinning process is considered in the judgment of whether to perform the BDOF process in the prediction signal generation unit 111D, 241D. In Non-Patent Document 3, the absolute value difference value of the prediction signal of L0 and L1 is used for judgment, but by making a judgment based on the result of the thinning process, the process of calculating the above-mentioned absolute value difference can be reduced.

[0212] Furthermore, the above-described image encoding apparatus 100 and image decoding apparatus 200 may be realized by a program that causes a computer to execute each function (each step).

[0213] In addition, in the above-mentioned embodiments, the present invention is described by taking the application to the image encoding device 100 and the image decoding device 200 as an example, but the present invention is not limited thereto and can also be applied to an encoding / decoding system having the functions of an encoding device and a decoding device.

[0214] According to the present invention, even when the thinned motion vector cannot be used for application determination of the deblocking filter, the deblocking filter can be appropriately applied to the boundary of the thinned block, and block noise can be suppressed to improve subjective image quality.

[0215] Explanation of symbols

[0216] 10…image processing system; 100…image encoding device; 111, 241…inter-frame prediction unit; 111A…motion vector search unit; 111B…motion vector encoding unit; 111C, 241C…thinning unit; 111D, 241D…prediction signal generation unit; 112, 242…intra-frame prediction unit; 121…subtractor; 122, 230…adder; 131…transformation and quantization unit; 132, 220…inverse transformation and inverse Quantization unit; 140 ... encoding unit; 150, 250 ... in-loop filter processing unit; 151, 251 ... target block boundary detection unit; 152, 252 ... adjacent block boundary detection unit; 153, 253 ... boundary strength determination unit; 154, 254 ... filter determination unit; 155, 255 ... filter processing unit; 160, 260 ... frame buffer; 200 ... image decoding device; 210 ... decoding unit; 241B ... motion vector decoding unit

Claims

1. An image decoding device, It is characterized in that have: a motion vector decoding unit configured to decode a motion vector from the encoded data; A refinement unit configured to perform refinement processing to correct the decoded motion vector; as well as a prediction signal generating unit configured to generate a prediction signal based on the corrected motion vector output from the thinning unit, The thinning unit performs the thinning process on each sub-block obtained by dividing the block. The prediction signal generating unit is configured as follows: Determining whether each of the sub-blocks satisfies a BDOF application condition; For the sub-blocks satisfying the BDOF application condition, determining whether to perform BDOF processing on each of the sub-blocks based on information calculated during the thinning process; When the search cost in the refinement process is below a predetermined threshold, determining not to perform the BDOF process in each of the sub-blocks; as well as The BDOF process is always applied to the sub-blocks that meet the BDOF application condition and have not been subjected to the thinning process.

2. The image decoding device according to claim 1, Features: The prediction signal generation unit is configured to determine not to perform the BDOF process on each of the sub-blocks when the difference between the motion vectors before and after the refinement process is 0 and the search cost in the reference block indicated by the motion vectors before and after the refinement process is below a predetermined threshold.

3. An image decoding method, It is characterized in that have: Step A, decoding motion vectors from encoded data; Step B, performing refinement processing on the motion vector obtained by correction decoding; as well as Step C, generating a prediction signal based on the corrected motion vector output in step B, In the step B, the thinning process is performed on each sub-block into which the block is divided, In step C: Determining whether each of the sub-blocks satisfies a BDOF application condition; For the sub-blocks satisfying the BDOF application condition, determining whether to perform BDOF processing on each of the sub-blocks based on information calculated during the thinning process; When the search cost in the refinement process is below a predetermined threshold, determining not to perform the BDOF process in each of the sub-blocks; as well as The BDOF process is always applied to the sub-blocks that meet the BDOF application condition and have not been subjected to the thinning process.

4. A computer-readable storage medium, It is characterized in that The computer-readable storage medium has a program for use in an image decoding device, and when the program is executed by a computer, the computer executes the following steps: Step A, decoding motion vectors from encoded data; Step B, performing refinement processing on the motion vector obtained by correction decoding; as well as Step C, generating a prediction signal based on the corrected motion vector output in step B, In the step B, the thinning process is performed on each sub-block into which the block is divided, In step C: Determining whether each of the sub-blocks satisfies a BDOF application condition; For the sub-blocks satisfying the BDOF application condition, determining whether to perform BDOF processing on each of the sub-blocks based on information calculated during the thinning process; When the search cost in the refinement process is below a predetermined threshold, determining not to perform the BDOF process in each of the sub-blocks; as well as The BDOF process is always applied to the sub-blocks that meet the BDOF application condition and have not been subjected to the thinning process.

Citation Information

Patent Citations

  • Image encoding device, image decoding device, image encoding method, and image decoding method

    CN102656889A

  • Inter-prediction method and apparatus in image coding system

    WO2017188566A1