Image decoding device, image decoding method, and program
By performing detailed processing before the BDOF processing and controlling the application of the BDOF processing using the calculated information, the problem of the inability to shorten the BDOF processing time in both software and hardware implementation in the prior art is solved, and efficient processing in both implementation methods is achieved.
Patent Information
- Application Number
- CN202310232278.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-11
- Filing Date
- 2020-03-04
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-03-04
AI Technical Summary
The prior art cannot shorten the execution time of BDOF processing when implemented by software, and the execution time is also increased during hardware implementation, so the processing time cannot be effectively shortened in both implementation methods.
By performing refinement processing before the BDOF processing, the calculated information is used to control whether the BDOF processing is applied, thereby reducing the processing amount in the hardware implementation, and reducing the number of blocks to apply the BDOF processing in the software implementation, shortening the processing time.
It is realized that the processing volume related to BDOF processing can be reduced in both hardware and software implementation, thereby shortening processing time and improving efficiency.
Smart Images

Figure CN116193144B_ABST
Abstract
Description
[0001] This application is a divisional application of the application with application number 202080019761.5, application date March 4, 2020, and invention name "Image Decoding Device, Image Decoding Method, and Program". Technical Field
[0002] The present invention relates to an image decoding device, an image decoding method, and a program. Background Art
[0003] Conventionally, regarding the BODF (Bi-Directional Optical Flow) technology, the following technology has been disclosed: in order to shorten the execution time of software, the absolute value difference sum of pixel values between two reference images used in the BDOF process is calculated, and when such an absolute value difference sum is less than a predetermined threshold, the BDOF process in the block is skipped (for example, refer to Non-Patent Document 1).
[0004] On the other hand, from the viewpoint of reducing the processing delay in hardware implementation, a technology for deleting the skip process of BDOF based on the calculation of the above absolute value difference sum has also been disclosed (for example, refer to Non-Patent Document 2).
[0005] Prior Art Documents
[0006] Non-Patent Documents
[0007] Non-Patent Document 1: Versatile Video Coding (Draft 4), JVET-M1001
[0008] Non-Patent Document 2: CE9-related: BDOF buffer reduction and enabling VPDU based application, JVET-M0890 Summary of the Invention
[0009] Problems to be Solved by the Invention
[0010] However, for example, in the technology disclosed in Non-Patent Document 1, there are the following problems: when implementing such a technology by software, the execution time can be shortened, but when implemented by hardware, the execution time increases.
[0011] On the other hand, in the technology disclosed in Non-Patent Document 2, there are the following problems: the execution time in hardware can be shortened, but the execution time in software increases.
[0012] Therefore, in the prior art as described above, there is a problem that the processing time cannot be shortened both in software implementation and in hardware implementation.
[0013] Accordingly, the present invention has been made in view of the above problems, and an object thereof is to provide an image decoding apparatus, an image decoding method, and a program that use information calculated during a refinement process performed before BDOF processing to control whether or not to apply BDOF processing. Thus, from the viewpoint of hardware implementation, the processing amount can be reduced by reusing the already calculated values, and from the viewpoint of software implementation, the processing time can be shortened by reducing the number of blocks to which BDOF processing is applied.
[0014] Means for Solving the Problems
[0015] The gist of the first feature of the present invention lies in an image decoding apparatus having: a motion vector decoding unit configured to decode a motion vector from encoded data; a refinement unit configured to perform a refinement process for correcting the decoded motion vector; and a prediction signal generation unit configured to generate a prediction signal based on the corrected motion vector output from the refinement unit, the prediction signal generation unit being configured to determine whether to apply BDOF processing to each block based on information calculated during the refinement process.
[0016] The gist of the second feature of the present invention lies in an image decoding apparatus having: a motion vector decoding unit configured to decode a motion vector from encoded data; a refinement unit configured to perform a refinement process for correcting the decoded motion vector; and a prediction signal generation unit configured to generate a prediction signal based on the corrected motion vector output from the refinement unit, the prediction signal generation unit being configured to apply BDOF processing when an application condition is satisfied, the application condition being that the motion vector is encoded in the Symmetric MVD mode and the magnitude of the differential motion vector transmitted in the Symmetric MVD mode is within a preset threshold.
[0017] The gist of the third feature of the present invention lies in having: step A of decoding a motion vector from encoded data; step B of performing a refinement process for correcting the decoded motion vector; and step C of generating a prediction signal based on the corrected motion vector output from the refinement unit, in step C, determining whether to apply BDOF processing to each block based on information calculated during the refinement process.
[0018] The gist of the fourth feature of the present invention lies in a program used in an image decoding device, which causes a computer to execute the following steps: Step A, decoding a motion vector from encoded data; Step B, performing a refinement process for correcting the decoded motion vector; and Step C, generating a prediction signal based on the corrected motion vector output from the refinement unit. In Step C, it is determined whether to apply BDOF processing to each block based on the information calculated during the refinement process.
[0019] Effect of the Invention
[0020] According to the present invention, it is possible to provide an image decoding device, an image decoding method, and a program that can reduce the processing amount related to BDOF processing both in hardware implementation and software implementation. Description of the Drawings
[0021] Figure 1 FIG. is an example showing the structure of an image processing system 10 according to an embodiment.
[0022] Figure 2 FIG. is an example showing the functional blocks of an image encoding device 100 according to an embodiment.
[0023] Figure 3 FIG. is an example showing the functional blocks of an inter-frame prediction unit 111 of an image encoding device 100 according to an embodiment.
[0024] Figure 4 FIG. is a flowchart showing an example of the processing procedure of a refinement unit 111C of an inter-frame prediction unit 111 of a moving image decoding device 30 according to an embodiment.
[0025] Figure 5 FIG. is a flowchart showing an example of the processing procedure of a prediction signal generation unit 111D of an inter-frame prediction unit 111 of a moving image decoding device 30 according to an embodiment.
[0026] Figure 6 FIG. is an example showing the functional blocks of an in-loop filtering processing unit 150 of an image encoding device 100 according to an embodiment.
[0027] Figure 7 FIG. is an example for explaining the determination performed by a boundary strength determination unit 153 of an in-loop filtering processing unit 150 of an image encoding device 100 according to an embodiment.
[0028] Figure 8 FIG. is an example showing the functional blocks of an image decoding device 200 according to an embodiment.
[0029] Figure 9FIG. is an example of a functional block diagram of an inter prediction unit 241 of an image decoding apparatus 200 showing one embodiment.
[0030] Figure 10 FIG. is an example of a functional block diagram of an in-loop filtering processing unit 250 of an image decoding apparatus 200 showing one embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In addition, the components in the following embodiments can be appropriately replaced with existing components, etc., and various modifications including combinations with other existing components can be made. Therefore, the description of the following embodiments does not limit the content of the invention described in the claims.
[0032] (First Embodiment)
[0033] Hereinafter, Figures 1 to 10 an image processing system 10 according to a first embodiment of the present invention will be described. Figure 1 FIG. is a diagram showing the image processing system 10 of the present embodiment.
[0034] As Figure 1 shown, the image processing system 10 includes an image encoding apparatus 100 and an image decoding apparatus 200.
[0035] The image encoding apparatus 100 is configured to generate encoded data by encoding an input image signal. The image decoding apparatus 200 is configured to generate an output image signal by decoding the encoded data.
[0036] Among them, such encoded data can be transmitted from the image encoding apparatus 100 to the image decoding apparatus 200 via a transmission path. In addition, the encoded data can also be provided from the image encoding apparatus 100 to the image decoding apparatus 200 after being stored in a storage medium.
[0037] (Image Encoding Apparatus 100)
[0038] Hereinafter, Figure 2 the image encoding apparatus 100 of the present embodiment will be described. Figure 2 FIG. is an example of a functional block diagram of the image encoding apparatus 100 of the present embodiment.
[0039] As Figure 2 shown, the image encoding apparatus 100 includes an inter prediction unit 111, an intra prediction unit 112, a subtractor 121, an adder 122, a transform and quantization unit 131, an inverse transform and inverse quantization unit 132, an encoding unit 140, an in-loop filtering processing unit 150, and a frame buffer 160.
[0040] The inter-frame prediction unit 111 is configured to generate a prediction signal through inter-frame prediction.
[0041] Specifically, the inter-frame prediction unit 111 is configured to determine a reference block included in a reference frame by comparing an encoding target frame (hereinafter referred to as a target frame) with the reference frame stored in the frame buffer 160, and determine a motion vector for the determined reference block.
[0042] In addition, the inter-frame prediction unit 111 is configured to generate a prediction signal included in each prediction block based on the reference block and the motion vector. The inter-frame prediction unit 111 is configured to output the prediction signal to the subtracter 121 and the adder 122. Herein, the reference frame is a frame different from the target frame.
[0043] The intra-frame prediction unit 112 is configured to generate a prediction signal through intra-frame prediction.
[0044] Specifically, the intra-frame prediction unit 112 is configured to determine a reference block included in the target frame, and generate a prediction signal for each prediction block based on the determined reference block. In addition, the intra-frame prediction unit 112 is configured to output the prediction signal to the subtracter 121 and the adder 122.
[0045] Herein, the reference block is a block referred to for a prediction target block (hereinafter referred to as a target block). For example, the reference block is a block adjacent to the target block.
[0046] The subtracter 121 is configured to subtract the prediction signal from the input image signal, and output a prediction residual signal to the transform and quantization unit 131. Herein, the subtracter 121 is configured to generate a prediction residual signal, which is a difference between the prediction signal generated through intra-frame prediction or inter-frame prediction and the input image signal.
[0047] The adder 122 is configured to add the prediction signal and the prediction residual signal output from the inverse transform and inverse quantization unit 132 to generate a pre-filtering decoded signal, and output such a pre-filtering decoded signal to the intra-frame prediction unit 112 and the in-loop filtering unit 150.
[0048] Herein, the pre-filtering decoded signal constitutes a reference block used in the intra-frame prediction unit 112.
[0049] The transform and quantization unit 131 is configured to perform a transform process on the prediction residual signal, and obtain a coefficient level value. Further, the transform and quantization unit 131 may also be configured to perform quantization of the coefficient level value.
[0050] Among them, the transformation process is a process of transforming the prediction residual signal into a frequency component signal. In such a transformation process, a basic pattern (transformation matrix) corresponding to the discrete cosine transform (DCT: Discrete Cosine Transform) can be used, or a basic pattern (transformation matrix) corresponding to the discrete sine transform (DST: Discrete Sine Transform) can be used.
[0051] The inverse transform and inverse quantization unit 132 is configured to perform an inverse transform process on the coefficient level value output from the transform and quantization unit 131. Among them, the inverse transform and inverse quantization unit 132 can also be configured to perform inverse quantization of the coefficient level value before the inverse transform process.
[0052] Among them, the inverse transform process and inverse quantization are performed in a process opposite to the transform process and quantization performed by the transform and quantization unit 131.
[0053] The encoding unit 140 is configured to encode the coefficient level value output from the transform and quantization unit 131 and output encoded data.
[0054] Among them, for example, encoding is entropy encoding that assigns codes of different lengths based on the occurrence probability of the coefficient level value.
[0055] In addition, the encoding unit 140 is configured to encode control data used in the decoding process in addition to the coefficient level value.
[0056] Among them, the control data may include size data such as the coding unit (CU) size, prediction unit (PU) size, transform unit (TU) size, etc.
[0057] The in-loop filtering processing unit 150 is configured to perform filtering processing on the decoded signal before filtering output from the adder 122 and output the decoded signal after filtering to the frame buffer 160.
[0058] Among them, for example, the filtering processing is deblocking filtering processing that reduces distortion generated at the boundary portion of a block (coding block, prediction block, or transform block).
[0059] The frame buffer 160 is configured to accumulate reference frames used in the inter-frame prediction unit 111.
[0060] Among them, the decoded signal after filtering constitutes the reference frame used in the inter-frame prediction unit 111.
[0061] (Inter-frame prediction unit 111)
[0062] Hereinafter, with reference to Figure 3The inter-frame prediction unit 111 of the image encoding device 100 according to this embodiment will be described. Figure 3 FIG. is an example of a functional block diagram showing the inter-frame prediction unit 111 of the image encoding device 100 according to this embodiment.
[0063] As Figure 3 shown, the inter-frame prediction unit 111 includes a motion vector search unit 111A, a motion vector encoding unit 111B, a refinement unit 111C, and a prediction signal generation unit 111D.
[0064] The inter-frame prediction unit 111 is an example of a prediction unit configured to generate a prediction signal included in a prediction block based on a motion vector.
[0065] The motion vector search unit 111A is configured to determine a reference block included in a reference frame by comparing a target frame with the reference frame, and search for a motion vector for the determined reference block.
[0066] In addition, the above search process is performed on multiple reference frame candidates, and the reference frame and motion vector used for prediction in the prediction block are determined. Up to two reference frames and motion vectors can be used for one block. The case where only one set of reference frame and motion vector is used for one block is called single prediction, and the case where two sets of reference frames and motion vectors are used is called dual prediction. Hereinafter, the first set is called L0, and the second set is called L1.
[0067] Furthermore, the motion vector search unit 111A is configured to determine an encoding method for the reference frame and the motion vector. In addition to the normal method of separately transmitting the information of the reference frame and the motion vector, the encoding method also includes the merge mode described later, the Symmetric MVD mode described in Non-Patent Document 2, etc.
[0068] In addition, regarding the motion vector search method, the reference frame determination method, and the method for determining the encoding method of the reference frame and the motion vector, known methods can be adopted, so their detailed descriptions are omitted.
[0069] The motion vector encoding unit 111B is configured to encode the information of the reference frame and the motion vector, which are also determined by the motion vector search unit 111A, using the encoding method determined by the motion vector search unit 111A.
[0070] When the encoding method of the block is the merge mode, first, a merge list for the block is created. The merge list is a list that enumerates multiple combinations of reference frames and motion vectors. An index is assigned to each combination. Instead of encoding the information of the reference frame and the motion vector separately, only the above index is encoded and transmitted to the decoding side. The method for creating the merge list is shared on the encoding side and the decoding side. Thus, on the decoding side, the information of the reference frame and the motion vector can be decoded only from the index information. Regarding the method for creating the merge list, a known method can be adopted, and thus its detailed description is omitted.
[0071] The Symmetric MVD mode is an encoding method that can be used only when dual prediction can be performed in the block. In the Symmetric MVD mode, only the motion vector (differential motion vector) of L0 among the two (L0, L1) reference frames and the two (L0, L1) motion vectors, which is the information to be transmitted to the decoding side, is encoded. The remaining motion vector of L1 and the information of the two reference frames are uniquely determined on the encoding side and the decoding side respectively by a predetermined method.
[0072] Regarding the encoding of the motion vector information, first, a predicted value of the motion vector to be encoded, that is, a predicted motion vector, is generated, and the difference value between the predicted motion vector and the actual motion vector to be encoded, that is, the differential motion vector, is encoded.
[0073] In the Symmetric MVD mode, as the differential motion vector of L1, a vector obtained by inverting the code of the encoded differential motion vector of L0 is used. Regarding the specific method, for example, the method described in Non-Patent Document 2 can be used.
[0074] The refinement unit 111C is configured to perform a refinement process (e.g., DMVR) for correcting the motion vector encoded by the motion vector encoding unit 111B.
[0075] Specifically, the refinement unit 111C is configured to perform the following refinement process: set a search range based on the reference position determined by the motion vector encoded by the motion vector encoding unit 111B, determine a corrected reference position with the smallest predetermined cost from the search range, and correct the motion vector based on the corrected reference position.
[0076] Figure 4 is a flowchart showing an example of the processing procedure of the refinement unit 111C.
[0077] As Figure 4As shown, in step S41, the refinement unit 111C determines whether a predetermined condition for applying the refinement process is satisfied. If all the predetermined conditions are satisfied, this processing proceeds to step S42. If any of the predetermined conditions is not satisfied, this processing proceeds to step S45 to end the refinement process.
[0078] Among them, the predetermined conditions include the condition that the block is a block for dual prediction.
[0079] Furthermore, the predetermined conditions may also include the condition that the motion vector is encoded in the merge mode.
[0080] In addition, the predetermined conditions may also include the condition that the motion vector is encoded in the Symmetric MVD mode.
[0081] Furthermore, the predetermined conditions may also include the following condition: the motion vector is encoded in the Symmetric MVD mode, and the magnitude of the differential motion vector (MVD of L0) transmitted in the Symmetric MVD mode is within a preset threshold.
[0082] Among them, the magnitude of the differential motion vector can be defined, for example, by the absolute values of the horizontal and vertical direction components of the differential motion vector.
[0083] As such a threshold, different thresholds can be used for the horizontal and vertical direction components of the motion vector respectively, or a common threshold can be used for the horizontal and vertical direction components. In addition, the value of the threshold can be set to 0. In this case, it means that the predicted motion vector and the motion vector to be encoded are the same value. In addition, the threshold can also be defined by a minimum value and a maximum value. In this case, as a predetermined condition, it includes the condition that the value or absolute value of the differential motion vector is equal to or greater than a predetermined minimum value and equal to or less than a maximum value.
[0084] In addition, the predetermined conditions may also include the condition that the motion vector is encoded in the merge mode or the Symmetric MVD mode. Similarly, the predetermined conditions may also include the following condition: the motion vector is encoded in the merge mode or the Symmetric MVD mode, and the magnitude of the differential motion vector transmitted in the case of being encoded in the Symmetric MVD mode is within a preset threshold.
[0085] In step S42, the refinement unit 111C generates a search image based on the motion vector encoded by the motion vector encoding unit 111B and the information of the reference frame.
[0086] Among them, when the motion vector points to a non-integer pixel position, the refinement unit 111C applies a filter to the pixel values of the reference frame to interpolate the pixels at the non-integer pixel positions. At this time, the refinement unit 111C can reduce the amount of computation by using an interpolation filter with a smaller number of taps than the interpolation filter used in the prediction signal generation unit 111D described later. For example, the refinement unit 111C can interpolate the pixel values at the non-integer pixel positions by bilinear interpolation.
[0087] In step S43, the refinement unit 111C uses the search image generated in step S42 to perform a search with integer pixel accuracy. Here, integer pixel accuracy means searching only for points that are at integer pixel intervals based on the motion vector encoded by the motion vector encoding unit 111B.
[0088] The refinement unit 111C determines the corrected motion vector at the integer pixel interval position through the search in step S42. As a search method, a known method can be used. For example, the refinement unit 111C can also perform the search by the following method: only search for points that are combinations obtained by inverting only the codes of the differential motion vectors on the L0 side and the L1 side. Among them, as a result of the search in step S43, it is also possible that the value is the same as the motion vector before the search.
[0089] In step S44, the refinement unit 111C uses the corrected motion vector with integer pixel accuracy determined in step S43 as an initial value to perform a motion vector search with non-integer pixel accuracy. As a search method for the motion vector, a known method can be used.
[0090] In addition, the refinement unit 111C can actually not perform a search, but use a parametric model such as parabola fitting with the result of step S43 as an input to determine the vector with non-integer pixel accuracy.
[0091] The refinement unit 111C determines the corrected motion vector with non-integer pixel accuracy in step S44, and then enters step S45 to end the refinement process. Here, for convenience, the expression of the corrected motion vector with non-integer pixel accuracy is used, but according to the search result of step S44, the result may also be the same as the motion vector with integer pixel accuracy obtained in step S43.
[0092] The refinement unit 111C can also divide a block larger than a predetermined threshold into smaller sub-blocks and perform a refinement process on each sub-block. For example, the refinement unit 111C can preset the execution unit of the refinement process to 16×16 pixels, and when the size in the horizontal or vertical direction of the block is greater than 16 pixels, divide it into 16 pixels or less respectively. At this time, as the motion vector serving as the basis for the refinement process, for all sub-blocks within the same block, the motion vector of this block encoded by the motion vector encoding unit 111B is used.
[0093] When processing each sub-block, the refinement unit 111C can perform Figure 4 all processes. In addition, the refinement unit 111C can only perform Figure 4 a part of the Figure 4 processing on each sub-block. Specifically, the refinement unit 111C can perform the processing of steps S41 and S42 of
[0094] for each block, and only perform the processing of steps S43 and S44 for each sub-block.
[0095] The prediction signal generation unit 111D is configured to generate a prediction signal based on the corrected motion vector output from the refinement unit 111C.
[0096] Among them, as will be described later, the prediction signal generation unit 111D is configured to determine whether to perform BDOF processing on each block based on the information (such as search cost) calculated during the above refinement process.
[0097] Figure 5 is a flowchart showing an example of the processing procedure of the prediction signal generation unit 111D.
[0098] Among them, when the refinement unit 111C performs refinement processing in units of sub-blocks, the processing of the prediction signal generation unit 111D is also performed in units of sub-blocks. In this case, words such as block in the following description can be appropriately replaced with sub-block.
[0099] As Figure 5 shown, in step S51, the prediction signal generation unit 111D generates a prediction signal.
[0100] Specifically, the prediction signal generation unit 111D takes as input the motion vector encoded by the motion vector encoding unit 111B or the motion vector encoded by the refinement unit 111C, and when the position pointed to by such a motion vector is a non-integer pixel position, applies a filter to the pixel values of the reference frame to interpolate the pixels at the non-integer pixel position. Among them, regarding the specific filter, a horizontal-vertical separable filter with up to 8 taps disclosed in Non-Patent Document 3 can be applied.
[0101] When this block is a block for bi-prediction, both a prediction signal based on the first (hereinafter referred to as L0) reference frame and the motion vector, and a prediction signal based on the second (hereinafter referred to as L1) reference frame and the motion vector are generated.
[0102] In step S52, the prediction signal generation unit 111D confirms whether the following BDOF (Bi-Directional Optical Flow) application conditions are satisfied.
[0103] As such application conditions, the conditions described in Non-Patent Document 3 can be applied. The application conditions at least include the condition that this block is a block for bi-prediction. In addition, the application conditions can also include, as described in Non-Patent Document 1, the condition that the motion vector of this block is not encoded in the Symmetric MVD mode.
[0104] In addition, the application conditions can also include the following conditions: the motion vector of this block is not encoded in the Symmetric MVD mode, or when it is encoded in the Symmetric MVD mode, the magnitude of the transmitted differential motion vector is within a preset threshold. Among them, the magnitude of the differential motion vector can be determined by the same method as in step S41 above. The value of the threshold can also be set to 0 in the same way as in step S41 above.
[0105] When the application conditions are not satisfied, this processing procedure goes to step S55 and ends the processing. At this time, the prediction signal generation unit 111D outputs the prediction signal generated in step S51 as the final prediction signal.
[0106] On the other hand, when all the application conditions are satisfied, this processing procedure enters step S53. In step S53, this processing procedure actually determines whether to perform the BDOF processing of step S54 for the block that satisfies the application conditions.
[0107] For example, the prediction signal generation unit 111D calculates the absolute value difference sum of the prediction signal of L0 and the prediction signal of L1, and determines not to perform the BDOF processing when its value is below a predetermined threshold.
[0108] Among them, for the block that has been refined by the refinement unit 111C, the prediction signal generation unit 111D can also use the result of the refinement process for the application of the presence or absence of BDOF.
[0109] For example, as a result of the refinement process, when the difference between the motion vectors before and after correction is equal to or less than a predetermined threshold, the prediction signal generation unit 111D can determine not to apply BDOF. When such a threshold is set to "0" for both the horizontal and vertical direction components, the result of the refinement process is equivalent to determining not to apply the BDOF process when the motion vector has not changed compared to before correction.
[0110] The prediction signal generation unit 111D can also use the search cost calculated during the above refinement process (e.g., the absolute value difference sum between the pixel values of the reference block on the L0 side and the reference block on the L1 side) to determine whether to apply BDOF.
[0111] In addition, hereinafter, the case where the absolute value difference sum is used as the search cost will be described as an example, but other metrics can also be used as the search cost. For example, as long as it is a metric value for determining the similarity between image signals, such as the absolute value difference sum or the sum of squared errors between the signals after removing the local average.
[0112] For example, in the integer pixel position search in step S43, when the absolute value difference sum at the search point with the minimum above-mentioned search cost (absolute value difference sum) is less than a predetermined threshold, the prediction signal generation unit 111D can determine not to apply BDOF.
[0113] In addition, the prediction signal generation unit 111D can also combine the following methods to determine whether to apply the BFOF process: the method using the change in the motion vector before and after the above refinement process; and the method using the search cost of the above refinement process.
[0114] For example, when the difference between the motion vectors before and after the refinement process is equal to or less than a predetermined threshold and the search cost of the refinement process is equal to or less than a predetermined threshold, the prediction signal generation unit 111D can determine not to apply the BDOF process.
[0115] Among them, when the threshold for the difference between the motion vectors before and after the refinement process is set to 0, the above-mentioned search cost is determined to be the absolute value difference sum between the reference blocks pointed to by the motion vector before the refinement process (= the motion vector after the refinement process).
[0116] In addition, the prediction signal generation unit 111D can also make a judgment in the block that has been refined by a method based on the result of the refinement process, and make a judgment in other blocks by a method based on the absolute value difference sum.
[0117] In addition, the prediction signal generation unit 111D may also adopt the following structure: As described above, instead of performing the process of recalculating the absolute value difference sum between the prediction signal on the L0 side and the prediction signal on the L1 side, only the information obtained from the result of the refinement process is used to determine whether to apply BDOF. In this case, in step S53, the prediction signal generation unit 111D determines to always apply BDOF to the blocks that have not undergone the refinement process.
[0118] According to such a structure, in this case, there is no need to perform the calculation process of the absolute value difference sum in the prediction signal generation unit 111D. Therefore, from the perspective of hardware implementation, the processing amount and processing delay can be reduced.
[0119] In addition, according to such a structure, from the perspective of software implementation, the result of the refinement process is used so that BDOF processing is not performed in the blocks where the effect of BDOF processing is speculated to be low. Thus, the coding efficiency can be maintained and the processing time of the entire image can be shortened.
[0120] In addition, the determination process itself using the result of the above refinement process is executed inside the refinement unit 111C, and the information indicating the result is transmitted to the prediction signal generation unit 111D. Thus, the prediction signal generation unit 111D can also determine whether to apply BDOF processing.
[0121] For example, as described above, the motion vectors and search cost values before and after the refinement process are determined, and a flag is prepared in advance. This flag is "1" when the condition for not applying BDOF is met, and "0" when the condition for not applying BDOF is not met and the refinement process has not been applied. Thus, the prediction signal generation unit 111D can refer to the value of such a flag to determine whether to apply BDOF.
[0122] In addition, here, for the sake of convenience, steps S52 and S53 have been described as different steps, but the determinations in steps S52 and S53 can also be performed simultaneously.
[0123] In the above determination, for the blocks determined not to apply BDOF, this processing procedure proceeds to step S55. For other blocks, this processing procedure enters step S54.
[0124] In step S54, the prediction signal generation unit 111D performs BDOF processing. Since the BDOF processing itself can use known methods, the detailed description is omitted. After the BDOF processing is implemented, this processing procedure enters step S55 and ends the processing.
[0125] (In-loop filter processing unit 150)
[0126] Hereinafter, the in-loop filtering processing unit 150 of the present embodiment will be described. Figure 6 FIG. is a diagram showing the in-loop filtering processing unit 150 of the present embodiment.
[0127] As Figure 6 shown, the in-loop filtering processing unit 150 includes a target block boundary detection unit 151, an adjacent block boundary detection unit 152, a boundary strength determination unit 153, a filter determination unit 154, and a filtering processing unit 155.
[0128] Among them, the structure with "A" appended at the end is related to the deblocking filtering processing of the vertical block boundary, and the structure with "B" appended at the end is related to the deblocking filtering processing of the horizontal block boundary.
[0129] Hereinafter, a case where the deblocking filtering processing of the horizontal block boundary is performed after the deblocking filtering processing of the vertical block boundary will be exemplified.
[0130] As described above, the deblocking filtering processing can be applied to coded blocks, prediction blocks, and transform blocks. In addition, it can also be applied to sub-blocks obtained by dividing the above-mentioned blocks. That is, the target block and the adjacent block can be coded blocks, prediction blocks, transform blocks, or sub-blocks obtained by dividing them.
[0131] The definition of the sub-block includes the sub-blocks described as the processing units of the refinement unit 111C and the prediction signal generation unit 111D. When applying the deblocking filter to the sub-block, the blocks in the following description can be appropriately replaced with sub-blocks.
[0132] Since the deblocking filtering processing of the vertical block boundary and the deblocking filtering processing of the horizontal block boundary are the same processing, the deblocking filtering processing of the vertical block boundary will be described below.
[0133] The target block boundary detection unit 151A is configured to detect the boundary of the target block based on control data indicating the block size of the target block.
[0134] The adjacent block boundary detection unit 152A is configured to detect the boundary of the adjacent block based on control data indicating the block size of the adjacent block.
[0135] The boundary strength determination unit 153A is configured to determine the boundary strength of the block boundary between the target block and the adjacent block.
[0136] In addition, the boundary strength determination unit 153A may also be configured to determine the boundary strength of the block boundary based on control data indicating whether the target block and the adjacent block are intra-prediction blocks.
[0137] For example, as Figure 7As shown, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of a block boundary is "2" when at least one of the target block and the adjacent block is an intra prediction block (i.e., when at least one of the blocks on both sides of the block boundary is an intra prediction block).
[0138] In addition, the boundary strength determination unit 153A may also be configured to determine the boundary strength of a block boundary based on control data indicating whether the target block and the adjacent block contain non-zero (zero) orthogonal transform coefficients and whether the block boundary is the boundary of a transform block.
[0139] For example, as Figure 7 shown, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of a block boundary is "1" when at least one of the target block and the adjacent block contains non-zero orthogonal transform coefficients and the block boundary is the boundary of a transform block (i.e., when there are non-zero transform coefficients in at least one of the blocks on both sides of the block boundary and it is the boundary of a TU).
[0140] In addition, the boundary strength determination unit 153A may also be configured to determine the boundary strength of a block boundary based on control data indicating whether the absolute value of the difference between the motion vectors of the target block and the adjacent block is greater than or equal to a threshold value (e.g., 1 pixel).
[0141] For example, as Figure 7 shown, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of a block boundary is "1" when the absolute value of the difference between the motion vectors of the target block and the adjacent block is greater than or equal to a threshold value (e.g., 1 pixel) (i.e., when the absolute value of the difference between the motion vectors of the blocks on both sides of the block boundary is greater than or equal to a threshold value (e.g., 1 pixel)).
[0142] In addition, the boundary strength determination unit 153A may also be configured to determine the boundary strength of a block boundary based on control data indicating whether the reference blocks referred to in the prediction of the motion vectors of the target block and the adjacent block are different.
[0143] For example, as Figure 7 shown, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of a block boundary is "1" when the reference blocks referred to in the prediction of the motion vectors of the target block and the adjacent block are different (i.e., when the reference images are different in the blocks on both sides of the block boundary).
[0144] The boundary strength determination unit 153A may also be configured to determine the boundary strength of a block boundary based on control data indicating whether the number of motion vectors of the target block and the adjacent block is different.
[0145] For example, as Figure 7As shown, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of the block boundary is "1" when the number of motion vectors of the target block and the adjacent block is different (i.e., when the number of motion vectors in the blocks on both sides of the block boundary is different).
[0146] The boundary strength determination unit 153A may also be configured to determine the boundary strength of the block boundary according to whether the refinement process performed by the refinement unit 111C is applied to the target block and the adjacent block.
[0147] For example, as Figure 7 shown, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of the block boundary is "1" when both the target block and the adjacent block are blocks to which the refinement process is applied by the refinement unit 111C.
[0148] Among them, the boundary strength determination unit 153A may also determine that "the refinement process is applied" in the block when all the predetermined conditions in Figure 4 step S41 are satisfied. In addition, a flag indicating the determination result of Figure 4 step S41 is prepared in advance, and the boundary strength determination unit 153A may also be configured to determine whether the refinement process is applied according to the value of the flag.
[0149] In addition, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of the block boundary is "1" when at least one of the target block and the adjacent block is a block to which the refinement process performed by the refinement unit 111C is applied.
[0150] Alternatively, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of the block boundary is "1" when the boundary is a sub-block boundary in the refinement process performed by the refinement unit 111C.
[0151] Furthermore, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of the block boundary is "1" when DMVR is applied to at least one of the blocks on both sides of the block boundary.
[0152] For example, as Figure 7 shown, the boundary strength determination unit 153A may also be configured to determine that the boundary strength of the block boundary is "0" when the above conditions are not satisfied.
[0153] In addition, the larger the value of the boundary strength, the higher the possibility of large block distortion occurring at the block boundary.
[0154] The above boundary strength determination method can be determined using a method common to the luminance signal and the color difference signal, or can be determined using partially different conditions. For example, the conditions related to the above refinement process can be applied to both the luminance signal and the color difference signal, or can be applied only to the luminance signal or only to the color difference signal.
[0155] In addition, headers called SPS (Sequence Parameter Set) and PPS (Picture Parameter Set) can have a flag for controlling whether to consider the result of the refinement process in the determination of the boundary strength.
[0156] The filter determination unit 154A is configured to determine the type of filtering process (e.g., deblocking filtering process) applied to the block boundary.
[0157] For example, the filter determination unit 154A can also be configured to determine whether to apply a filtering process to the block boundary and which filtering process of weak filtering process and strong filtering process to apply based on the boundary strength of the block boundary, quantization parameters included in the target block and the adjacent block, etc.
[0158] The filter determination unit 154A can also be configured to determine not to apply a filtering process when the boundary strength of the block boundary is "0".
[0159] The filtering process unit 155A is configured to process the pre-deblocking image based on the determination of the filter determination unit 154A. The processing of the pre-deblocking image is no filtering process, weak filtering process, strong filtering process, etc.
[0160] (Image decoding device 200)
[0161] Hereinafter, with reference to Figure 8 the image decoding device 200 of the present embodiment will be described. Figure 8 is a diagram showing an example of the functional blocks of the image decoding device 200 of the present embodiment.
[0162] As Figure 8 shown, the image decoding device 200 includes a decoding unit 210, an inverse transform and inverse quantization unit 220, an adder 230, an inter-frame prediction unit 241, an intra-frame prediction unit 242, an in-loop filtering process unit 250, and a frame buffer 260.
[0163] The decoding unit 210 is configured to decode the encoded data generated by the image encoding device 100 and decode the coefficient level values.
[0164] Among them, for example, decoding is entropy decoding, which is a process opposite to the entropy encoding performed by the encoding unit 140.
[0165] In addition, the decoding unit 210 may also be configured to obtain control data by decoding the encoded data.
[0166] In addition, as described above, the control data may also include size data such as the encoded block size, the predicted block size, and the transformed block size.
[0167] The inverse transform and inverse quantization unit 220 is configured to perform an inverse transform process on the coefficient level values output from the decoding unit 210. Among them, the inverse transform and inverse quantization unit 220 may also be configured to perform inverse quantization of the coefficient level values before the inverse transform process.
[0168] Among them, the inverse transform process and the inverse quantization are performed in a process opposite to the transform process and the quantization performed by the transform and quantization unit 131.
[0169] The adder 230 is configured to add the prediction signal and the prediction residual signal output from the inverse transform and inverse quantization unit 220 to generate a decoded signal before filtering processing, and output the decoded signal before filtering processing to the intra-frame prediction unit 242 and the in-loop filtering processing unit 250.
[0170] Among them, the decoded signal before filtering processing constitutes a reference block used in the intra-frame prediction unit 242.
[0171] The inter-frame prediction unit 241, similar to the inter-frame prediction unit 111, is configured to generate a prediction signal through inter-frame prediction.
[0172] Specifically, the inter-frame prediction unit 241 is configured to generate a prediction signal for each prediction block based on the motion vector decoded from the encoded data and the reference signal included in the reference frame. The inter-frame prediction unit 241 is configured to output the prediction signal to the adder 230.
[0173] The intra-frame prediction unit 242, similar to the intra-frame prediction unit 112, is configured to generate a prediction signal through intra-frame prediction.
[0174] Specifically, the intra-frame prediction unit 242 is configured to determine the reference block included in the target frame, and generate a prediction signal for each prediction block based on the determined reference block. The intra-frame prediction unit 242 is configured to output the prediction signal to the adder 230.
[0175] The in-loop filtering processing unit 250, similar to the in-loop filtering processing unit 150, is configured to perform filtering processing on the decoded signal before filtering processing output from the adder 230, and output the decoded signal after filtering processing to the frame buffer 260.
[0176] Among them, for example, the filtering process is a deblocking filtering process that reduces distortion generated at the boundary portion of a block (encoding block, prediction block, transform block, or sub-block formed by dividing these blocks).
[0177] Similar to the frame buffer 160, the frame buffer 260 is configured to accumulate reference frames used in the inter-frame prediction unit 241.
[0178] Among them, the decoded signal after the filtering process constitutes the reference frame used in the inter-frame prediction unit 241.
[0179] (Inter-frame prediction unit 241)
[0180] Hereinafter, Figure 9 the inter-frame prediction unit 241 of the present embodiment will be described. Figure 9 FIG. is an example of a functional block diagram showing the inter-frame prediction unit 241 of the present embodiment.
[0181] As Figure 9 shown, the inter-frame prediction unit 241 includes a motion vector decoding unit 241B, a refinement unit 241C, and a prediction signal generation unit 241D.
[0182] The inter-frame prediction unit 241 is an example of a prediction unit configured to generate a prediction signal included in a prediction block based on a motion vector.
[0183] The motion vector decoding unit 241B is configured to obtain a motion vector by decoding control data received from the image encoding device 100.
[0184] Similar to the refinement unit 111C, the refinement unit 241C is configured to perform a refinement process for correcting a motion vector.
[0185] Similar to the prediction signal generation unit 111D, the prediction signal generation unit 241D is configured to generate a prediction signal based on a motion vector.
[0186] (In-loop filtering processing unit 250)
[0187] Hereinafter, the in-loop filtering processing unit 250 of the present embodiment will be described. Figure 10 FIG. is a diagram showing the in-loop filtering processing unit 250 of the present embodiment.
[0188] As Figure 10 shown, the in-loop filtering processing unit 250 includes a target block boundary detection unit 251, an adjacent block boundary detection unit 252, a boundary strength determination unit 253, a filter determination unit 254, and a filtering processing unit 255.
[0189] Among them, the structure with "A" appended at the end is related to the deblocking filtering process of the vertical block boundary, and the structure with "B" appended at the end is related to the deblocking filtering process of the horizontal block boundary.
[0190] Here, a case where the deblocking filtering process of the horizontal block boundary is performed after the deblocking filtering process of the vertical block boundary is exemplified.
[0191] As described above, the deblocking filtering process can be applied to coded blocks, prediction blocks, and transform blocks. In addition, it can also be applied to sub-blocks obtained by dividing the above-mentioned blocks. That is, the target block and the adjacent block can be coded blocks, prediction blocks, transform blocks, or sub-blocks obtained by dividing them.
[0192] Since the deblocking filtering process of the vertical block boundary and the deblocking filtering process of the horizontal block boundary are the same process, the deblocking filtering process of the vertical block boundary will be described below.
[0193] Similar to the target block boundary detection unit 151A, the target block boundary detection unit 251A is configured to detect the boundary of the target block based on control data indicating the block size of the target block.
[0194] Similar to the adjacent block boundary detection unit 152A, the adjacent block boundary detection unit 252A is configured to detect the boundary of the adjacent block based on control data indicating the block size of the adjacent block.
[0195] Similar to the boundary strength determination unit 153A, the boundary strength determination unit 253A is configured to determine the boundary strength of the block boundary between the target block and the adjacent block. The method for determining the boundary strength of the block boundary is as described above.
[0196] Similar to the filter determination unit 154A, the filter determination unit 254A is configured to determine the type of deblocking filtering process applied to the block boundary. The method for determining the type of deblocking filtering process is as described above.
[0197] Similar to the filtering process unit 155A, the filtering process unit 255A is configured to process the pre-deblocking image based on the determination of the filter determination unit 254A. The processing of the pre-deblocking image includes no filtering process, weak filtering process, strong filtering process, etc.
[0198] In the image encoding device 100 and the image decoding device 200 according to the present embodiment, in the boundary strength determination units 153 and 253, when determining the boundary strength, it is considered whether the refinement process performed by the refinement units 111C and 241C is applied to the blocks adjacent to the boundary.
[0199] For example, as described above, when the refinement process is applied to at least one of the two blocks adjacent to the boundary, the boundary strength of the boundary is set to 1.
[0200] For boundaries with a boundary strength of "1" or more, in the filter determination units 154 and 254, parameters such as quantization parameters are considered to determine whether to apply a deblocking filter to the block boundary and the type of deblocking filter.
[0201] By adopting such a structure, even when it is impossible to use the refined motion vector value in the determination of the above boundary strength due to hardware implementation limitations, the deblocking filter can be appropriately applied to the boundary of the block where the refinement process is performed, and block noise can be suppressed to improve the subjective image quality.
[0202] Multiple methods for determining whether to apply a deblocking filter and determining the boundary strength have been proposed. For example, Patent Document 1 discloses a technique for omitting the application of a deblocking filter based on information such as whether the block is in a skip mode or the like. Additionally, for example, Patent Document 2 discloses a technique for omitting the application of a deblocking filter using quantization parameters.
[0203] However, neither considers whether the refinement process is applied. One of the problems to be solved by the present invention is a problem specific to the refinement process, that is, in the refinement process of correcting the value of the motion vector decoded from the syntax on the decoding side, when the corrected motion vector value cannot be used for the application determination of the deblocking filter, the application determination cannot be appropriately performed.
[0204] Therefore, the problems cannot be solved by the methods of Patent Document 1 and Patent Document 2 that do not consider whether the refinement process is applied. On the other hand, the methods of Patent Document 1 and Patent Document 2 can be combined with the application determination of the deblocking filter of the present invention.
[0205] According to the image encoding device 100 and the image decoding device 200 of the present embodiment, as the execution condition of the refinement process in the above refinement units 111C and 241C, it is considered whether the motion vector serving as the basis for the process is encoded in the Symmetric MVD mode.
[0206] In such a refinement process, when only searching for points where the absolute values of the differential motion vectors on the L0 side and the L1 side are the same and the codes are inverted, motion vectors similar to those in the case of transmitting the motion vector in the Symmetric MVD mode can be obtained.
[0207] Therefore, when the originally transmitted motion vector is obtained through the above refinement process, the differential motion vector transmitted in the Symmetric MVD mode is reduced as much as possible, thereby reducing the code amount related to the differential motion vector.
[0208] In particular, in the case of an encoding method in which only a flag is transmitted when the differential motion vector has a specific value (such as 0), and the differential value is directly encoded for other values, even if the difference in the motion vector corrected in such a refinement process is small, the effect of reducing the code amount may be large.
[0209] In addition, as an execution condition of the above refinement process, a condition is set such that the motion vector serving as a basis for such a refinement process is encoded in the Symmetric MVD mode and the value of the differential motion vector is below a predetermined threshold. Thus, it is also possible to implicitly switch whether to execute the refinement process according to the value of the differential motion vector.
[0210] Similarly, in the application conditions of the BDOF process in the prediction signal generation units 111D and 241D, a condition is used in which the motion vector of the block is encoded in the Symmetric MVD mode and the value of the differential motion vector is below a predetermined threshold. Thus, it is possible to implicitly switch whether to apply the BDOF process according to the value of the differential motion vector.
[0211] In the image encoding apparatus 100 and the image decoding apparatus 200 according to the present embodiment, in the determination of whether to execute the BDOF process in the prediction signal generation units 111D and 241D, the result of the above refinement process is considered. In Non-Patent Document 3, the absolute value difference value of the prediction signals of L0 and L1 is used for the determination. On the contrary, by performing the determination based on the result of the refinement process, the process of calculating the above absolute value difference can be reduced.
[0212] In addition, the above image encoding apparatus 100 and image decoding apparatus 200 can also be implemented by a program that causes a computer to execute each function (each step).
[0213] Furthermore, in each of the above embodiments, the present invention has been described by taking the application to the image encoding apparatus 100 and the image decoding apparatus 200 as an example, but the present invention is not limited thereto, and can also be applied to an encoding / decoding system having each function of an encoding apparatus and a decoding apparatus.
[0214] According to the present invention, even when the refined motion vector cannot be used for the application determination of the deblocking filter, the deblocking filter can be appropriately applied to the boundary of the block subjected to the refinement process, and block noise can be suppressed to improve the subjective image quality.
[0215] Symbol Description
[0216] 10…Image processing system; 100…Image encoding device; 111, 241…Inter-frame prediction unit; 111A…Motion vector search unit; 111B…Motion vector encoding unit; 111C, 241C…Refinement unit; 111D, 241D…Predicted signal generation unit; 112, 242…Intra-frame prediction unit; 121…Subtractor; 122, 230…Adder; 131…Transformation and quantization unit; 132, 220…Inverse transformation and inverse quantization unit; 140…Encoding unit; 150, 250…In-loop filtering processing unit; 151, 251…Target block boundary detection unit; 152, 252…Adjacent block boundary detection unit; 153, 253…Boundary strength determination unit; 154, 254…Filter determination unit; 155, 255…Filtering processing unit; 160, 260…Frame buffer; 200…Image decoding device; 210…Decoding unit; 241B…Motion vector decoding unit.
Claims
1. An image decoding device, characterized in that, it has: a motion vector decoding unit configured to decode a motion vector from encoded data; a refinement unit configured to perform a refinement process for correcting the decoded motion vector; and a prediction signal generation unit configured to generate a prediction signal based on the corrected motion vector output from the refinement unit, wherein the refinement unit performs the refinement process on each sub-block obtained by dividing a block, and the prediction signal generation unit is configured to: determine for each of the sub-blocks whether a BDOF application condition is satisfied; for the sub-blocks that satisfy the BDOF application condition, determine whether to perform BDOF processing on each of the sub-blocks based on information calculated during the refinement process; when the search cost during the refinement process is below a predetermined threshold, determine that no BDOF processing is performed on each of the sub-blocks; and always apply the BDOF processing to the sub-blocks that satisfy the BDOF application condition and have not been subjected to the refinement process, wherein the size of the sub-blocks in the horizontal and vertical directions is 16 pixels or less respectively.
2. The image decoding device according to claim 1, characterized in that: the prediction signal generation unit is configured to determine that no BDOF processing is performed on each of the sub-blocks when the difference between the motion vectors before and after the refinement process is 0 and the search cost in the reference blocks indicated by the motion vectors before and after the refinement process is below a predetermined threshold.
3. An image decoding method, characterized in that, it has: Step A: Decode a motion vector from encoded data; Step B: Perform a refinement process for correcting the decoded motion vector; and Step C: Generate a prediction signal based on the corrected motion vector output in Step B, wherein in Step B, the refinement process is performed on each sub-block obtained by dividing a block, and in Step C: determine for each of the sub-blocks whether a BDOF application condition is satisfied; for the sub-blocks that satisfy the BDOF application condition, determine whether to perform BDOF processing on each of the sub-blocks based on information calculated during the refinement process; when the search cost during the refinement process is below a predetermined threshold, determine that no BDOF processing is performed on each of the sub-blocks; and always apply the BDOF processing to the sub-blocks that satisfy the BDOF application condition and have not been subjected to the refinement process, wherein the size of the sub-blocks in the horizontal and vertical directions is 16 pixels or less respectively.
4. A computer-readable storage medium, characterized in that, the computer-readable storage medium has a program used in an image decoding device, and when the program is executed by a computer, it causes the computer to perform the following steps: Step A: Decode a motion vector from encoded data; Step B: Perform a refinement process for correcting the decoded motion vector; and Step C: Generate a prediction signal based on the corrected motion vector output in Step B, In the step B, the refinement process is performed on each sub-block obtained by dividing the block. In the step C: Determine whether each of the sub-blocks satisfies the BDOF application condition; For the sub-blocks that satisfy the BDOF application condition, determine whether to perform BDOF processing on each of the sub-blocks based on the information calculated during the refinement process; When the search cost in the refinement process is below a predetermined threshold, determine that the BDOF processing is not performed on each of the sub-blocks; And For the sub-blocks that satisfy the BDOF application condition and have not undergone the refinement process, always apply the BDOF processing. The sizes of the sub-blocks in the horizontal and vertical directions are each 16 pixels or less.
Citation Information
Patent Citations
Bidirectional optical flow and perceptual hash based fingertip tracking method
CN105261038A
Dynamic image encoding device, dynamic image decoding device, dynamic image encoding method, and dynamic image decoding method
CN106713930A