Image Symbolization Device, Method, and Program
The video encoding apparatus addresses the challenge of maintaining image quality and controlling code amount by selectively excluding high-frequency components based on a calculated statistic, thereby optimizing predictive encoding methods.
Patent Information
- Application Number
- JP2023552683
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-06
- Filing Date
- 2021-12-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-10
AI Technical Summary
Existing video encoding methods, such as VVC, face challenges in maintaining subjective image quality while minimizing the increase in code amount, particularly when high-frequency transform coefficients are excluded due to their small values.
A video encoding apparatus and method that select an optimal predictive encoding method by excluding certain conversion methods from candidates based on a calculated statistic, which assesses the significance of high-frequency components relative to the overall image content.
This approach effectively improves subjective image quality by reducing information loss in areas with noticeable features, while also controlling the amount of generated codes, thus balancing image quality and coding efficiency.
Smart Images

Figure 0007683719000002 
Figure 0007683719000003 
Figure 0007683719000004
Abstract
Description
Technical Field
[0001] The present invention relates to a video encoding apparatus and a video encoding method for encoding moving images.
Background Art
[0002] Non-Patent Document 1 discloses a video encoding method called VVC (Versatile Video Coding).
[0003] In VVC, each frame of a video is divided into blocks called coding tree units (CTUs), and the encoding process for each CTU is performed in raster scan order.
[0004] Each CTU is composed of a set of coding units (CUs). The encoding process is executed for each CU. A CU corresponds to a block obtained by dividing a CTU using a quad-tree (QT) structure or a multi-type tree (MMT) structure, or the CTU itself. In the quad-tree structure, the block is equally divided in the horizontal and vertical directions. In the multi-type tree structure, in the horizontal or vertical direction, the short side of the divided block is divided into two parts so that the ratio becomes 1:1. Or, in the horizontal or vertical direction, the short side of the divided block is divided into three parts so that the ratio becomes 1:2:1.
[0005] In each CU, a predicted image is generated for each prediction unit (PU) obtained by dividing the CU. Usually, the size of the PU is the same as the size of the CU. As methods for generating a predicted image (hereinafter simply referred to as prediction methods), there are intra prediction and inter prediction (hereinafter simply referred to as inter prediction) involving a motion compensation method.
[0006] The difference is calculated between the images before and after prediction for each PU, and a prediction error image for each PU is generated. A prediction error image for the corresponding CU is defined from the prediction error images of each PU.
[0007] For the prediction error image of each CU, transform coefficients are obtained by applying transform processing in units of transform units (TUs) obtained by dividing the CU. As a transform method, a frequency transform method mainly using discrete cosine transform (DCT) is used. When both the width and height of the TU are 32 or less, it is also possible to use a frequency transform method selected from a plurality of frequency transform methods such as discrete sine transform (DST). Also, in the transform processing, it is possible to select and use a transform method other than the frequency transform method called transform skip.
[0008] The obtained transform coefficients are quantized using a value determined by a quantization parameter (QP) or the like, and quantization coefficients are generated. Generally, the larger the value of the QP, the greater the amount of information loss. After the quantization coefficients are integerized, the integerized quantization coefficients are arithmetic coded.
[0009] Generally, the energy of the transform coefficients generated by frequency transform is concentrated in the low-frequency region. Therefore, the values of the transform coefficients in the low-frequency region become large, and the values of the transform coefficients in the high-frequency region become small.
[0010] When a frequency transform method is selected, when at least one of the width and height of the TU exceeds 32, the portion exceeding 32, that is, the transform coefficients of the high-frequency components, are excluded regardless of the magnitude of the values. Therefore, the number of transform coefficients to be quantized and arithmetic coded is 32×32 or less.
Prior Art Documents
Non-Patent Documents
[0011]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0012] The video encoding device selects an optimal combination from a number of combinations of a method for dividing a CTU to be encoded into CUs, a prediction method for each CU generated by the division, and a transformation method. The video encoding device performs predictive encoding using the selected combination. For example, when selecting the optimal combination, the video encoding device targets a prediction error image generated based on a prediction method corresponding to a candidate for a predictive encoding method that can be used, and performs transformation processing, quantization processing, inverse quantization processing, inverse frequency transformation processing corresponding to the transformation processing, arithmetic encoding processing, etc. by a transformation method corresponding to the predictive encoding method that can be used. Note that the predictive encoding method includes at least a prediction method and a transformation method.
[0013] The transformation method used in the video encoding device also includes a method of excluding coefficients from the targets of encoding regardless of the values of the transform coefficients. For example, as described above, when a frequency transformation method is selected as the transformation method and either the width or height of the TU is greater than 32, the process of excluding transform coefficients so that the transform coefficients become 32×32 is applied regardless of the values of the transform coefficients when using the frequency transformation method.
[0014] The transform coefficients to be excluded are transform coefficients in the high-frequency region as described above, and the values of the transform coefficients in the high-frequency region are generally small. Therefore, even if the transform coefficients in the high-frequency region are excluded, the quality of the image decoded by the video decoding device is not significantly affected (degraded) in many cases. Also, by using a TU that meets the above conditions, it is possible to reduce the amount of code generated when encoding the region corresponding to that TU compared to the case of dividing and encoding it into a plurality of TUs.
[0015] However, when the conditions for the width and height of the above TU are satisfied and the value of the conversion coefficient in the high-frequency region is relatively large, the amount of information loss due to the exclusion of the coefficient increases. As a result, the quality of the image decoded by the video decoder deteriorates.
[0016] Therefore, an object of the present invention is to provide a video encoding apparatus and a video encoding method capable of improving subjective image quality while suppressing a large increase in the amount of generated codes when selecting an optimal predictive encoding method.
Means for Solving the Problems
[0017] The video encoding apparatus according to the present invention includes a predictive encoding method selection unit that selects a predictive encoding method to be applied to a processing target block from a plurality of candidate predictive encoding methods. The candidate predictive encoding methods include a conversion method that excludes a predetermined conversion coefficient from the processing target in the conversion of the prediction error signal, and an exclusion means that excludes the conversion method from the selection targets from the candidate predictive encoding methods.
[0018] The video encoding apparatus according to the present invention includes a predictive encoding method selection unit that selects a predictive encoding method to be applied to a processing target block from a plurality of candidate predictive encoding methods. The candidate predictive encoding methods include a conversion method that excludes a predetermined conversion coefficient from the processing target in the conversion of the prediction error signal. The original signal or predicted signal of the processing target block, or a signal generated using the original signal or predicted signal is set as the calculation target signal, and block analysis means for calculating a predetermined statistic based on the calculation target signal; Exclusion means for excluding the conversion method from the selection targets from the candidate predictive encoding methods and including it is noted that the block analysis means calculates the statistic using the calculation target signal of the processing target block and the calculation target signal of a block within the same video frame or a block within another video frame, and the exclusion means excludes the conversion method from being selected from candidates of the prediction coding method when the statistic is a value within a predetermined range. .
[0019] The video encoding method according to the present invention selects a predictive encoding method to be applied to a processing target block from a plurality of candidate predictive encoding methods. The candidate predictive encoding methods include a conversion method that excludes a predetermined conversion coefficient from the processing target in the conversion of the prediction error signal. The original signal or predicted signal of the processing target block, or a signal generated using the original signal or predicted signal is set as the calculation target signal, and a predetermined statistic is calculated based on the calculation target signal. When calculating the statistic, the statistic is calculated using the calculation target signal of the processing target block and the calculation target signal of a block within the same video frame or a block within another video frame. When selecting a predictive encoding method, When the statistic is a value within a predetermined range; the conversion method is excluded from the selection targets from the candidate predictive encoding methods.
Effects of the Invention
[0020] The video encoding program according to the present invention causes a computer to select a prediction encoding method to be applied to a processing target block from a plurality of candidate prediction encoding methods. The candidate prediction encoding methods include a conversion method that excludes a predetermined conversion coefficient from the processing target in the conversion of the prediction error signal. The computer is caused to The original signal or predicted signal of the processing target block, or a signal generated using the original signal or predicted signal is set as the calculation target signal, and a predetermined statistic is calculated based on the calculation target signal. When calculating the statistic, the statistic is calculated using the calculation target signal of the processing target block and the calculation target signal of a block within the same video frame or a block within another video frame. exclude the conversion method from the selection targets from the candidate prediction encoding methods when selecting the prediction encoding method.
Brief Description of Drawings
[0021]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Modes for Carrying Out the Invention
[0022] Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0023] Embodiment 1 (Description of Configuration) FIG. 1 is a block diagram showing a configuration example of a video encoding device. The video encoding device shown in FIG. 1 includes a block division unit 101, a subtractor 102, a conversion unit 103, a quantization unit 104, an inverse quantization unit 105, an inverse conversion unit 106, an adder 107, a loop filter 108, a prediction unit 110, an arithmetic encoder 113, a coding method determination unit 114, and a code sequence generation unit 115. The prediction unit 110 includes an intra predictor 111 and an inter predictor 112.
[0024] The video signal encoding device further includes a control unit 120 including a block analysis unit 121 and an encoding method control unit 122.
[0025] The block splitting unit 101 splits the input video frame into a plurality of CTUs. Further, the block splitting unit 101 defines a set of CUs for each CTU. The set of CUs can be obtained by defining the CTU as a CU without splitting it. Alternatively, the set of CUs can be obtained by defining each block obtained by splitting the CTU using a quadtree structure or a multi-type tree structure as a CU. Also, the block splitting unit 101 defines the CU as a PU without splitting it, or defines the blocks obtained by splitting the CU as PUs. Similarly, the block splitting unit 101 defines the CU as a TU without splitting it, or defines the blocks obtained by splitting the CU as TUs.
[0026] The subtractor 102 subtracts the prediction signal from the input signal (input pixel value) for each block selected by the block splitting unit 101 to generate a prediction error signal. The prediction error signal is also called a prediction residual or a prediction residual signal.
[0027] The conversion unit 103 performs frequency conversion on the prediction error signal of the processing target block to obtain conversion coefficients. The conversion unit 103 includes a plurality of types of frequency conversion functions including a type II DCT (DCT-II) and a conversion skip function that does not perform frequency conversion on the prediction error signal. The conversion unit 103 executes any of the above conversions using the conversion method selected by the encoding method control unit 122.
[0028] The quantization unit 104 quantizes the conversion coefficients to obtain quantization coefficients (conversion quantization values). The conversion quantization values are used by the arithmetic encoder 113 and the inverse quantization unit 105.
[0029] The inverse quantization unit 105 inverse quantizes the conversion quantization values to restore the conversion coefficients. The inverse conversion unit 106 inverse frequency-converts the conversion coefficients based on the conversion method executed by the conversion unit 103 to restore the prediction error signal.
[0030] The adder 107 adds the restored prediction error signal and the prediction signal to generate a reconstructed signal (reconstructed image).
[0031] The intra predictor 111, the loop filter 108, and the encoding method determination unit 114 take the reconstructed signal as an input.
[0032] In general, a block memory for storing a reference block in the picture to be encoded is provided in front of the prediction unit 110 or in the intra predictor 111, but it is omitted in FIG. 1.
[0033] The intra predictor 111 performs intra prediction on the block to be encoded with reference to the reference block, and generates a prediction signal (in this case, an intra prediction signal).
[0034] The loop filter 108 includes, for example, a deblocking filter, a sample adaptive offset filter, and an adaptive loop filter, and performs appropriate filtering. The reconstructed signal filtered by the loop filter 108 is input to the inter predictor 112. In general, a frame memory for storing a reference picture is provided in front of the prediction unit 110 or in the inter predictor 112, but it is omitted in FIG. 1.
[0035] The inter predictor 112 performs inter prediction on the block to be encoded with reference to a reference picture different from the picture to be encoded, and generates a prediction signal (in this case, an inter prediction signal).
[0036] The arithmetic encoder 113 generates an encoded signal (code sequence: bit stream) by arithmetically encoding the transformed quantization value. The arithmetic encoder 113 binarizes the transformed quantization value, arithmetically encodes the binary signal, and generates a binary arithmetic code.
[0037] The symbolization method determination unit 114 calculates the cost when predictive coding is performed using each of a plurality of prediction methods and conversion methods. The symbolization method determination unit 114 selects an optimal predictive coding method for the block to be processed. Generally, the rate-distortion cost (RD cost) J is calculated by the following formula (1) from the estimated bitstream length R and the distortion D between the original signal and the reconstructed signal. Note that the symbolization method determination unit 114 may calculate the cost by other means than the RD cost. J = D + λR (1)
[0038] The symbol sequence generation unit 115 selects a binary arithmetic code using the optimal predictive coding method and outputs it as a bitstream. For example, the bitstream is transmitted to an image decoding device. The bitstream may be output to a storage medium (not shown) and stored in the storage medium.
[0039] The block analysis unit 121 in the control unit 120 calculates a statistic representing the degree to which the block to be processed has a predetermined feature from the signal of the block to be processed. Examples of the predetermined feature include the values of pixels representing a specific color included in the block to be processed. When calculating the statistic from such a feature, the block analysis unit 121, for example, detects pixels representing a specific color. Then, the block analysis unit 121 uses, as the statistic, the ratio of the number of pixels representing a specific color to the number of pixels included in the block to be processed. Note that there may be a plurality of features to be detected. Also, the block analysis unit 121 may calculate a statistic for each feature of interest. The block analysis unit 121 may calculate a statistic for pixels corresponding to at least one feature.
[0040] That is, the block analysis unit 121 may calculate a plurality of statistics from the block to be processed and calculate one statistic from the plurality of statistics for the purpose of calculating the statistic of the block to be processed to be transmitted to the coding method control unit 122. Note that the technique of calculating one statistic from a plurality of statistics is also applicable to the method of calculating a statistic for each sub-block, which will be described later.
[0041] The symbolization method control unit 122 determines whether the statistical quantity calculated by the block analysis unit 121 satisfies the conditions given in advance as the constraints from the size of the processing target block, the prediction method, and the conversion method. If the conditions are not satisfied, the symbolization method control unit 122 sets the conversion method and the prediction method executed by the conversion unit 103 and the prediction unit 110, respectively. If the conditions are satisfied, the symbolization method control unit 122 controls so that the processing in the conversion unit 103 and the prediction unit 110 is not performed, and sets so that the cost calculated by the symbolization method determination unit 114 becomes the maximum value.
[0042] Hereinafter, the inverse quantization unit 105, the inverse transform unit 106, and the adder 107 may be referred to as a local decoding unit.
[0043] (Explanation of statistical quantity) As described above, in VVC, when a frequency conversion method is selected as the conversion method and at least one of the width and height of the TU exceeds 32, the conversion coefficients in the portion exceeding 32 (that is, the high-frequency region) are excluded. The portion exceeding 32 corresponds to the high-frequency region. After frequency conversion, if the degree of concentration of energy in the low-frequency components is large, even if the conversion coefficients in the high-frequency region are excluded, the quality of the image decoded by the video decoding apparatus (the decoded image quality) does not deteriorate so much.
[0044] However, if the degree of concentration of energy in the low-frequency components is not so large, when the conversion coefficients in the high-frequency region are excluded, the decoded image quality deteriorates. In other words, when the conversion coefficients in the high-frequency region are excluded during frequency conversion, the amount of information of the original image is reduced. That is, information is lost. As a result, the decoded image quality deteriorates. In particular, when the TU includes two or more regions with different patterns and is a region with easily noticeable features, significant deterioration occurs. Further, when the regions include regions with features that humans pay attention to, significant deterioration occurs.
[0045] Even when such image quality degradation occurs, a predictive coding method that satisfies the above conditions (conditions regarding the width and height of the above TUs) may be selected as the optimal predictive coding method. For example, as described above, when a large-sized TU is used, it becomes possible to reduce the amount of generated codes. Therefore, when a selection method that determines the predictive coding method based only on the amount of codes, that is, a selection method that emphasizes the amount of codes, is used, the predictive coding method that satisfies the above conditions may be determined to be the optimal predictive coding method.
[0046] From the perspective of subjective image quality, it is not desirable to perform predictive coding using a TU in which at least one of the width and height exceeds 32. However, by using a relatively large TU that satisfies the above conditions (the condition that at least one of the width and height exceeds 32), it is possible to reduce the amount of generated codes. Therefore, when the predictive coding method is restricted by the above conditions, when the amount of generated codes is controlled to be a predetermined value, the process for suppressing the amount of codes to be equivalent to the case without restrictions may cause information loss. That is, it may degrade the image quality.
[0047] In other words, in order to suppress the degradation of the decoded image quality, it is desirable that the predictive coding method that satisfies the above conditions be restricted. However, as described above, since the predictive coding method that satisfies the above conditions contributes to reducing the amount of generated codes, if the use is restricted in all regions within the image, the amount of generated codes may increase. In particular, the restriction of the above predictive coding method may cause degradation of the image quality when the amount of generated codes is controlled to be a predetermined value. For example, in order to suppress the amount of codes to be equivalent to the case without the above restrictions, processing such as quantization using a larger QP value is performed. The information loss caused by such processing may cause degradation of the image quality.
[0048] In this embodiment, the block analysis unit 121 calculates the number of pixels having a predetermined characteristic or the ratio of the number of pixels included in a pixel block to the number of pixels included in a processing target block, that is, the attention area occupancy rate, as a statistic. If the attention area occupancy rate is within a predetermined value range and at least one of the width and height of a TU exceeds 32, the coding method control unit 122 does not select a predictive coding method using such a TU. As a result, it is possible to suppress the degree of degradation in an area where degradation of subjective image quality is easily noticeable while suppressing the generated code amount.
[0049] Specifically, the block analysis unit 121 calculates the region of interest occupancy rate A of an input signal I having P pixels by dividing each pixel I of the input signal by p For example, when a specific pixel value C is set as a predetermined feature of interest, the block analysis unit 121 calculates the statistics using the following formula (2).
[0050]
number
[0051] The region of interest occupancy rate may be calculated using other calculation formulas. Also, for example, pixel values may be weighted according to pixel positions within the target block.
[0052] The coding method control unit 122 may, for example, determine whether the attention area occupancy rate A is greater than or equal to a threshold value related to a predetermined upper limit (hereinafter, th min ) and the lower limit threshold (hereafter, th max ) is within the range set by th min to a smaller value, th max is set to a larger value.
[0053] Note that, as an example of the statistic, the attention area occupancy rate, which is the ratio of the number of pixels having a specific pixel value to the number of pixels in the processing target block, has been described. However, in the present invention, the statistic is not limited to such a statistic.
[0054] For example, the block analysis unit 121 can use the pixel correlation between the block in the same video frame as the processing target block as the statistic. Also, the block analysis unit 121 can use the sum of absolute differences of pixels at the same position as the pixels in the block in adjacent video frames as the statistic. In other words, the block analysis unit 121 may calculate the statistic using the calculation target signal of the processing target block (the original signal or prediction signal of the processing target block, or a signal generated using the original signal or prediction signal) and the calculation target signal of the block in the same video frame or the block in another video frame (as an example, an adjacent video frame).
[0055] Furthermore, after the block division unit 101 divides the processing target block into sub-blocks, the statistic can be calculated for each sub-block, and a value selected from the statistics of each sub-block or a value calculated from the statistics of each sub-block can be used as the statistic. Also, the block analysis unit 121 can determine whether each sub-block has a predetermined feature from the calculation target signal of each sub-block, and calculate the ratio of the number of sub-blocks determined to have the predetermined feature to the total number of sub-blocks as the statistic.
[0056] Also, the block analysis unit 121 calculates the statistic from the original signal input to the video encoding device. However, the block analysis unit 121 may calculate the statistic from the prediction signal or prediction error signal. The block analysis unit 121 may calculate the statistic from a signal obtained by performing gamma conversion or the like on the original signal.
[0057] (Description of the operation) As an example, the video encoding device includes a storage unit (not shown) that stores a candidate table in which data capable of specifying each of a plurality of types of candidate prediction encoding methods is set. When evaluating candidate prediction encoding methods, the control unit 120 sets the conversion method to be evaluated in the conversion unit 103 and sets the prediction method in the prediction unit 110.
[0058] As prediction methods set in the candidate table, regarding intra prediction, the following prediction methods can be considered. · DC prediction · Planar prediction · Each of the angular predictions (Angular prediction)
[0059] Regarding intra prediction, as candidate prediction methods, the following prediction methods (see Non-Patent Document 1) may be added. · IBC (Intra Block Copy) · MIP (Matrix-based Intra Prediction)
[0060] Regarding inter prediction, the following prediction methods can be considered. · Adaptive motion vector encoding · Merge encoding
[0061] Regarding inter prediction, as candidate prediction methods, the following prediction methods (see Non-Patent Document 1) may be added. · Affine prediction · GPM (Geometric Partitioning Mode) · CIIP (Combined inter merge / intra prediction) · SBT (Sub-block transform)
[0062] As conversion methods set in the candidate table, the following conversion methods can be considered. · DCT-II · Transform skip
[0063] As candidates for the conversion method, the following conversion methods (see Non-Patent Document 1) may be added. ·DCT-VIII ·DST-VII ·Any combination of two of DCT-II, DCT-VIII, and DST-VII ·Combination of the above conversion method and LFNST (Low frequency non-separatable transform)
[0064] Note that, in the video encoding device, it is an example that a candidate table in which data capable of specifying each of the prediction mode candidates is set is used. For example, when the video encoding device is realized by a processor, data capable of specifying each of the prediction mode candidates may be described in a program.
[0065] The operation regarding the evaluation of candidates for the optimal prediction encoding method performed for each CTU of the video encoding device will be described with reference to the flowchart of FIG. 2.
[0066] The block division unit 101 selects one division pattern from the dividable patterns of the CTU to be evaluated and generates a set of CUs (step S100). Further, the block division unit 101 selects one CU from the set of CUs (step S101). Also, the encoding method control unit 122 selects one prediction method and one conversion method from a candidate table in which the prediction method and the conversion method (specifically, data capable of specifying the prediction method and data capable of specifying the conversion method) are set (step S102).
[0067] The encoding method control unit 122 determines whether at least one of the width and height of the TU exceeds 32 for the block input from the block division unit 101 (the processing target block that is the target of evaluation of the prediction encoding method candidate) (step S103). If it is determined that neither the width nor the height of the TU exceeds 32, the process proceeds to step S106. If at least one of the width and height of the TU exceeds 32, the process proceeds to step S104.
[0068] In step S104, the block analysis unit 121 calculates the occupancy rate A of the target area of the processing target block. The block analysis unit 121 notifies the encoding method control unit 122 of the occupancy rate A of the target area.
[0069] The encoding method control unit 122 compares the notified occupancy rate A of the target area with a preset threshold th min , th max That is, the encoding method control unit 122 determines whether the relationship of th min ≦A≦th max is satisfied. If the encoding method control unit 122 determines that the relationship is not satisfied, the process proceeds to step S106. If the encoding method control unit 122 determines that the relationship is satisfied, the process proceeds to step S110. In this case, in step S110, the encoding method determination unit 114 sets the RD cost to the maximum value. Note that the maximum value is a value larger than a value assumed as the RD cost corresponding to other predictive encoding methods.
[0070] In step S106, in the prediction unit 110, the intra predictor 111 or the inter predictor 112 generates a prediction signal for the block input from the block division unit 101. Also, the subtractor 102 generates a prediction error signal.
[0071] The conversion unit 103 frequency-converts the prediction error signal to generate conversion coefficients (step S107). Note that when at least one of the width and height of the TU exceeds 32, the conversion unit 103 excludes the conversion coefficients of the portion exceeding 32 (that is, the high-frequency region). That is, assuming a two-dimensional matrix with conversion coefficients as elements, in the conversion result of the conversion unit, both the row and the column are 32 or less.
[0072] In addition, when at least one of the horizontal size and the vertical size of TU exceeds 32, the conversion unit 103 may exclude the conversion coefficients in the high-frequency region as the conversion result. Further, the conversion unit 103 may use the conversion coefficients in the entire region as the conversion result, and the quantization unit 104 may quantize the conversion coefficients in the region where both the row and the column are 32 or less, and discard the other conversion coefficients.
[0073] In step S107, the quantization unit 104 quantizes the conversion coefficients from the conversion unit 103 to generate conversion quantization values. The inverse quantization unit 105 and the arithmetic encoder 113 input the conversion quantization values.
[0074] The inverse quantization unit 105 inverse-quantizes the conversion quantization values (step S108). Further, the inverse conversion unit 106 inverse-frequency-converts the inverse-quantized conversion quantization values to restore the conversion coefficients. The arithmetic encoder 113 arithmetically encodes the conversion quantization values to generate an encoded signal (step S109).
[0075] The encoding method determination unit 114 calculates the above-described RD cost J. Note that an index other than the formula (1) may be used. As an example, the encoding method determination unit 114 may use only one of R and D. When only R is used, the arithmetic encoding process (the process in step S109) is unnecessary. Further, for example, the encoding method determination unit 114 may use the cumulative sum (total sum) of the prediction error signals instead of the sum of the squares of the differences between the original image (input signal) and the reconstructed image (reconstructed signal). Furthermore, the encoding method determination unit 114 may use the input symbol amount to the arithmetic encoder or the symbol amount estimated by some method instead of the generated symbol amount of the arithmetic encoder.
[0076] If the evaluation of all the candidates of the prediction methods and the conversion methods set in the candidate table is completed, the process proceeds to step S112. If there are unevaluated candidates, the process returns to step S102.
[0077] If the evaluation of all CUs in the set of CUs has not been completed, the process returns to step S101. If the evaluation of all CUs has been completed, the encoding method determination unit 114 calculates the cost of the CTU in the current partition pattern being evaluated.
[0078] If the evaluation of all partition patterns of the CTU to be evaluated has been completed, the process ends. If there are unevaluated partition patterns, the process returns to step S100.
[0079] For example, in the process of step S110, the encoding method determination unit 114 temporarily stores the encoding efficiency (in this example, the RD cost) of each candidate prediction encoding method. The encoding method determination unit 114 determines the prediction encoding method that exhibits the minimum encoding efficiency among the stored encoding efficiencies as the prediction encoding method to be used in the actual encoding process, that is, the prediction encoding method applied to the processing target block.
[0080] Note that the encoding method determination unit 114 may save not the encoding efficiencies of all candidate prediction encoding methods, but the minimum encoding efficiency and the prediction encoding method that exhibits it. In that case, in the process of step S110, when the encoding efficiency calculated at that time is smaller than the stored encoding efficiency, the calculated encoding efficiency and the prediction mode that exhibits it are used to update the stored encoding efficiency and the prediction encoding method.
[0081] Other Embodiment 1. In VVC, SBT (Sub-block Transform) can be used. SBT is a method of dividing a block into two sub-blocks in the horizontal or vertical direction and performing frequency conversion only on one of the sub-blocks. All prediction error signals in the other sub-block are replaced with 0. Since information loss also occurs in SBT, it is conceivable to apply each of the above embodiments.
[0082] Other Embodiment 2. In VVC, LFNST (Low-Frequency Non-Separable Transform) can be used. When encoding with intra prediction, LFNST is a method of re-transforming the transform coefficients using an orthogonal transform matrix defined for LFNST. Up to a maximum of 48 coefficients are subject to re-transformation. All coefficients other than those subject to re-transformation (976 coefficients in the case of 32×32) are set to 0. Therefore, since coefficient exclusion is performed on the coefficients of the high-frequency components, information loss will occur even with LFNST, and it is conceivable to apply each of the above embodiments.
[0083] Each of the above-described video encoding apparatuses can be configured by individual hardware circuits or integrated circuits, but can also be realized by a computer having a processor such as a CPU (Central Processing Unit) and a memory. For example, a program for implementing the method (process) in the above embodiment may be stored in a storage device (storage medium), and each function may be realized by executing the program with a CPU.
[0084] FIG. 3 is a block diagram showing an example of a computer having a CPU. The computer is mounted in a video encoding apparatus. The CPU 1000 realizes each function in the above embodiment by executing processing according to a video encoding program stored in the storage device 1001. That is, the CPU 1000 realizes the functions of the control unit 120 including the block division unit 101, the subtractor 102, the transform unit 103, the quantization unit 104, the inverse quantization unit 105, the inverse transform unit 106, the adder 107, the loop filter 108, the prediction unit 110 (intra predictor 111 and inter predictor 112), the arithmetic encoder 113, the encoding method determination unit 114, the code sequence generation unit 115, and the block analysis unit 121 and the encoding method control unit 122 in the video encoding apparatus shown in FIG. 1.
[0085] The memory device 1001 is, for example, a non-transitory computer readable medium. The non-transitory computer readable medium includes various types of tangible storage media. Specific examples of the non-transitory computer readable medium include magnetic recording media (e.g., hard disks), CD-ROMs (Compact Disc-Read Only Memory), CD-Rs (Compact Disc-Recordable), CD-R / Ws (Compact Disc-ReWritable), and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs).
[0086] Also, the program may be stored in various types of transitory computer readable media. The transitory computer readable media include, for example, those to which the program is supplied via a wired communication path or a wireless communication path, that is, via an electrical signal, an optical signal, or an electromagnetic wave.
[0087] The memory 1002 is realized by, for example, a RAM (Random Access Memory), and is a storage means that temporarily stores data when the CPU 1000 executes processing. A form is also conceivable in which the program held by the memory device 1001 or the transitory computer readable medium is transferred to the memory 1002, and the CPU 1000 executes processing based on the program in the memory 1002.
[0088] FIG. 4 is a block diagram showing the main part of a video encoding apparatus. The video encoding apparatus 10 shown in FIG. 4 includes a prediction encoding method selection unit (prediction encoding method selection means) 15 (in the embodiment, realized by an encoding method determination unit 114) that selects a prediction encoding method to be applied to a processing target block from a plurality of candidate prediction encoding methods. The candidate prediction encoding methods include a conversion method that excludes a predetermined conversion coefficient from the processing target in the conversion of the prediction error signal (for example, a conversion method applied when at least one of the width and height of a TU exceeds 32). The video encoding apparatus 10 further includes an exclusion unit (exclusion means) 16 (in the embodiment, realized by an encoding method control unit 122) that excludes the conversion method from the selection candidates of the prediction encoding method.
[0089] FIG. 5 is a block diagram showing the main part of a video encoding apparatus according to another aspect. The video encoding apparatus 10 shown in FIG. 5 further includes a block analysis unit (block analysis means) 17 (in the embodiment, realized by a block analysis unit 121) that calculates a predetermined statistic based on a source signal or a prediction signal of a processing target block, or a signal generated using the source signal or the prediction signal as a calculation target signal. The exclusion unit 16 excludes the conversion method from the selection candidates of the prediction encoding method when the statistic is within a predetermined range of values.
[0090] FIG. 6 is a block diagram showing the main part of yet another video encoding apparatus. As shown in FIG. 6, the video encoding apparatus 10 further includes a division unit (division means) 18 (in the embodiment, realized by a block division unit 101) that divides a processing target block into sub-blocks of a predetermined size. The block analysis unit 17 includes means for calculating a statistic for each sub-block from the calculation target signal of each sub-block, and calculates the statistic of the processing target block from the values of the statistics of each sub-block. Further, the block analysis unit 17 may include means for determining whether each sub-block has a predetermined feature from the calculation target signal of each sub-block, and calculate, as a statistic, the ratio of the number of sub-blocks determined to have the predetermined feature to the total number of sub-blocks.
[0091] The present invention has been described with reference to the embodiments, but the present invention is not limited to the above embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.
[0092] This application claims priority based on Japanese Patent Application No. 2021-164585 filed on October 6, 2021, and incorporates the entire disclosure thereof herein.
Explanation of Reference Numerals
[0093] 10 Video Encoding Device 15 Prediction Encoding Method Selection Unit (Prediction Encoding Method Selection Means) 16 Exclusion Unit (Exclusion Means) 17 Block Analysis Unit (Block Analysis Means) 18 Division Unit (Division Means) 101 Block Division Unit 102 Subtractor 103 Conversion Unit 104 Quantization Unit 105 Inverse Quantization Unit 106 Inverse Conversion Unit 107 Adder 108 Loop Filter 110 Prediction Unit 111 Intra Predictor 112 Inter Predictor 113 Arithmetic Encoder 114 Encoding Method Determination Unit 115 Code Sequence Generation Unit 120 Control Unit 121 Block Analysis Unit 122 Encoding Method Control Unit 1000 CPU 1001 Storage Device 1002 Memory
Claims
1. A prediction coding method selection unit that selects a prediction coding method to be applied to a processing target block from a plurality of candidate prediction coding methods, wherein the candidates for the prediction coding method include a conversion method that excludes a predetermined conversion coefficient from the processing target in the conversion of the prediction error signal, block analysis means for calculating a calculation target signal using the original signal or prediction signal of the processing target block, or a signal generated using the original signal or the prediction signal, and calculating a predetermined statistic based on the calculation target signal, excluding means for excluding the conversion method from the selection candidates from the candidates for the prediction coding method, wherein the block analysis means calculates the statistic using the calculation target signal of the processing target block and the calculation target signal of a block in the same video frame or a block in another video frame, wherein the excluding means excludes the conversion method from the selection candidates from the candidates for the prediction coding method when the statistic is a value within a predetermined range A video coding device.
2. The block analysis means includes means for detecting pixels having a predetermined feature from the calculation target signal of the processing target block, and calculates, as a statistic, the ratio of the number of detected pixels to the number of pixels included in the processing target block. The video coding device according to claim 1.
3. It includes dividing means for dividing the processing target block into sub-blocks of a predetermined size, The block analysis means includes means for calculating a statistic for each sub-block from the calculation target signal of each sub-block, and calculates the statistic of the processing target block from the values of the statistics of each sub-block. The video coding device according to claim 1.
4. A prediction coding method selection unit that selects a prediction coding method to be applied to a processing target block from a plurality of candidate prediction coding methods, wherein the candidates for the prediction coding method include a conversion method that excludes a predetermined conversion coefficient from the processing target in the conversion of the prediction error signal, block analysis means for calculating a calculation target signal using the original signal or prediction signal of the processing target block, or a signal generated using the original signal or the prediction signal, and calculating a predetermined statistic based on the calculation target signal, dividing means for dividing the processing target block into sub-blocks of a predetermined size, excluding means for excluding the conversion method from the selection candidates from the candidates for the prediction coding method, The block analysis means means for determining whether each sub-block has a predetermined feature from the signal to be calculated for each sub-block; calculating, as the statistic, the ratio of the number of sub-blocks determined to have the predetermined feature to the total number of sub-blocks; the excluding means excludes the conversion method from the selection target from the candidates of the prediction encoding method when the statistic is a value within a predetermined range; A video encoding device.
5. The block analysis means calculates a plurality of the statistics from the processing target block, and calculates the statistic of one processing target block from the plurality of statistics. The video encoding device according to any one of Claims 1 to 4.
6. selecting a prediction encoding method to be applied to a processing target block from a plurality of prediction encoding method candidates; the candidates of the prediction encoding method include a conversion method for excluding a predetermined conversion coefficient from the processing target in the conversion of the prediction error signal; using the original signal or the prediction signal of the processing target block, or a signal generated using the original signal or the prediction signal as the signal to be calculated, calculating a predetermined statistic based on the signal to be calculated; when calculating the statistic, calculating the statistic using the signal to be calculated of the processing target block and the signal to be calculated of a block in the same video frame or a block in another video frame; when selecting a prediction encoding method, excluding the conversion method from the selection target from the candidates of the prediction encoding method when the statistic is a value within a predetermined range; A video encoding method.
7. Selecting a prediction encoding method to be applied to a processing target block from a plurality of prediction encoding method candidates; the candidates of the prediction encoding method include a conversion method for excluding a predetermined conversion coefficient from the processing target in the conversion of the prediction error signal; using the original signal or the prediction signal of the processing target block, or a signal generated using the original signal or the prediction signal as the signal to be calculated, calculating a predetermined statistic based on the signal to be calculated; dividing the processing target block into sub-blocks of a predetermined size; when calculating the statistic, determining whether each sub-block has a predetermined feature from the signal to be calculated for each sub-block, and calculating, as the statistic, the ratio of the number of sub-blocks determined to have the predetermined feature to the total number of sub-blocks; When selecting a prediction encoding method, when the statistic is a value within a predetermined range, exclude the conversion method from the selection candidates of the prediction encoding method. Video encoding method. **Claim 8** Cause a computer to select a prediction encoding method to be applied to a processing target block from among a plurality of prediction encoding method candidates. The candidates for the prediction encoding method include a conversion method that excludes a predetermined conversion coefficient from the processing target in the conversion of the prediction error signal. Cause the computer to Use the original signal or prediction signal of the processing target block, or a signal generated using the original signal or the prediction signal as a calculation target signal, and calculate a predetermined statistic based on the calculation target signal. When calculating the statistic, calculate the statistic using the calculation target signal of the processing target block and the calculation target signal of a block within the same video frame or a block within another video frame. When selecting a prediction encoding method, when the statistic is a value within a predetermined range, exclude the conversion method from the selection candidates of the prediction encoding method. Video encoding program therefor. **Claim 9** Cause a computer to select a prediction encoding method to be applied to a processing target block from among a plurality of prediction encoding method candidates. The candidates for the prediction encoding method include a conversion method that excludes a predetermined conversion coefficient from the processing target in the conversion of the prediction error signal. Cause the computer to Use the original signal or prediction signal of the processing target block, or a signal generated using the original signal or the prediction signal as a calculation target signal, and calculate a predetermined statistic based on the calculation target signal. Divide the processing target block into sub-blocks of a predetermined size. When calculating the statistic, determine whether each sub-block has a predetermined feature from the calculation target signal of each sub-block, and calculate the ratio of the number of sub-blocks determined to have the predetermined feature to the total number of sub-blocks as the statistic. When selecting a prediction encoding method, when the statistic is a value within a predetermined range, exclude the conversion method from the selection candidates of the prediction encoding method. Video encoding program therefor.
Citation Information
Patent Citations
System and method for low-complexity forward transform with zeroed-out coefficients
JP2017513342A