Image processing device, image processing method, and program

The image processing apparatus addresses the challenge of maintaining a target bit rate by dynamically adjusting image resolution and quantization steps, ensuring consistent bit rate control even at extreme quantization levels.

JP2025080590APending Publication Date: 2025-05-26CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023193847
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-26

AI Technical Summary

Technical Problem

Existing technologies struggle to maintain a target bit rate during image compression, particularly when the quantization step reaches its limits, leading to inconsistent bit rate control.

Method used

An image processing apparatus that dynamically adjusts the image resolution and quantization step using a control mechanism to maintain a constant bit rate, by changing the resolution based on the quantization step and adjusting the quantization step accordingly.

Benefits of technology

Effectively maintains the bit rate at the target value by dynamically adjusting the image resolution and quantization step, preventing bit rate deviations caused by extreme quantization steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025080590000001_ABST
    Figure 2025080590000001_ABST
Patent Text Reader

Abstract

To provide an image processing device capable of appropriately storing a target bit rate in accordance with a quantization step.SOLUTION: In an image distribution device 100, an image processing device 105 comprises: an expansion and reduction part that changes a resolution of an image; a coding part that performs quantization of an image of which the resolution is converted on the basis of a quantization step to be compressed; and a coding amount control part that controls the change of the resolution of the image by the expansion and reduction part on the basis of the quantization step. The coding amount control part determines whether or not a magnification ratio at the time of changing the resolution of the image on the basis of the quantization step is changed, and changes the quantization step in the case where it is determined that the magnification ratio is changed. The coding amount control part further determines whether or not the magnification ratio at the time of changing the resolution of the image is changed on the basis of the quantization step and the resolution of the image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing method, and a program.

Background Art

[0002] In the compression of moving images, the amount of coded data after compression varies depending on the nature of the input image. When storing compressed images, it is desirable that the amount of compressed coded data be controlled to be constant in order to determine the storage time available from the free capacity. A method of controlling the amount of compressed coded data to be constant is generally called CBR (Constant Bit Rate) (hereinafter, CBR). Patent Document 1 describes a method of performing CBR control by controlling the quantization step. Patent Document 2 performs control to control the amount of coded data by the quantization step and change the resolution according to the processing capacity of the decoder.

[0003] Also, as an encoding method for compression recording of moving images, a VVC (Versatile Video Coding) encoding method (hereinafter, VVC) is known. In VVC, a technique called RPR (Reference Picture Resampling) (hereinafter, RPR) has been introduced to improve the encoding efficiency. RPR is a technique that can use an image with a different resolution as a reference image for the image to be decoded and enables the resolution to be changed even in the case of inter-frame compression.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] In Patent Document 1 above, the case where the quantization step reaches the upper limit or the lower limit and the target bit rate cannot be maintained is not considered. In Patent Document 2 above, although controlling the resolution and the bit rate is described, there is no description regarding quantization.

[0006] As described above, in the above-described technology, there is a problem that the target bit rate cannot be appropriately maintained according to the quantization step.

[0007] Therefore, an object of the present invention is to maintain the bit rate at the target bit rate according to the quantization step.

Means for Solving the Problem

[0008] To solve this problem, for example, an image processing apparatus of the present invention has the following configuration. That is, a changing means for changing the resolution of an image, an encoding means for quantizing and compressing the image with the converted resolution based on a quantization step, a control means for controlling the change of the resolution of the image by the changing means based on the quantization step, and is characterized by comprising:

Effects of the Invention

[0009] According to the present invention, the bit rate can be maintained at the target bit rate according to the quantization step.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Best Mode for Carrying Out the Invention

[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are given the same reference numerals, and redundant explanations are omitted.

[0012] (First Embodiment) FIG. 1 is a block diagram showing the hardware configuration of the image distribution apparatus 100 according to the first embodiment. The image distribution apparatus 100 includes an image processing apparatus 105 and an imaging unit 120. The image processing apparatus 105 includes a CPU 101, a ROM 102, a RAM 103, a storage 104, a bus 110, a camera signal processing unit 130, a motor control unit 140, a zoom unit 150, an encoding unit 160, a code amount control unit 170, and an IP communication unit 180. The CPU 101, the ROM 102, the RAM 103, the storage 104, the camera signal processing unit 130, the motor control unit 140, the zoom unit 150, the encoding unit 160, the code amount control unit 170, and the IP communication unit 180 are connected to each other via the bus 110 so as to be able to transmit and receive data such as images. In the following description, the term "image" may be used as a concept including still images, moving images, videos, and image data such as those data.

[0013] The CPU 101 is an abbreviation for Central Processing Unit and is also called a central arithmetic processing unit. The CPU 101 reads out programs stored in the ROM 102 and the storage 104 and expands them in the RAM 103 to realize various functions. The CPU 101 overall controls the entire image distribution apparatus 100 by various functions.

[0014] The ROM 102 is an abbreviation for Read Only Memory and is a non-volatile memory such as an EEPROM or a flash memory. The ROM 102 stores programs executed by the CPU 101.

[0015] The RAM 103 is an abbreviation for Random Access Memory and is a volatile memory such as an SRAM (Static Random Access Memory) or a DRAM (Dynamic Random Access Memory). The RAM 103 functions as a work area when the CPU 101 executes a program.

[0016] Storage 104 is a non-volatile memory such as a HDD (Hard Disk Drive) and an SSD (Solid State Drive). Storage 104 stores programs executed by CPU 101 and data necessary for the programs.

[0017] The imaging unit 120 captures a subject, generates an analog image signal as an image, and outputs it to the image processing apparatus 105. The imaging unit 120 includes an imaging device 124 composed of a focus lens 121, a fixed lens 122, a diaphragm 123, and an image sensor, etc., and a lens driving unit 125. The focus lens 121 moves along the optical axis by the lens driving unit 125. The diaphragm 123 is driven by the lens driving unit 125 to operate. The imaging device 124 photoelectrically converts the light that has passed through the focus lens 121 and the diaphragm 123 to generate an analog image signal. The generated analog image signal is output to the camera signal processing unit 130 after amplification processing by sampling processing such as correlated double sampling.

[0018] The camera signal processing unit 130 acquires an analog image signal from the imaging device 124. For example, the camera signal processing unit 130 sequentially acquires analog image signals of images constituting a video one frame at a time. The camera signal processing unit 130 converts the acquired analog image signal into a digital image signal by A / D conversion, and then executes various digital image processes to generate image data. The various digital image processes are, for example, offset processing, gamma correction processing, gain processing, RGB interpolation processing, noise reduction processing, contour correction processing, color tone correction processing, light source type determination processing, etc. The camera signal processing unit 130 stores the image data after digital image processing in the RAM 103 via the bus 110. The motor control unit 140 controls the aforementioned lens driving unit 125.

[0019] The enlargement / reduction unit (modification unit) 150 changes the resolution of the image stored in the RAM 103 from the camera signal processing unit 130 based on the enlargement / reduction ratio. The enlargement / reduction ratio is a value indicating the ratio before and after the conversion of the resolution of the image. The enlargement / reduction ratio is changed by the sign amount control unit 170. Thus, the enlargement / reduction unit 150 functions as a modification means for changing the resolution of the image. Also, as will be described later, the sign amount control unit 170 controls the change in the resolution of the image by the modification means by changing the enlargement / reduction ratio.

[0020] Here, the enlargement / reduction ratio is assumed to be the magnification when enlarging the resolution of the image or the magnification when reducing the resolution of the image. When reducing, the magnification can be represented by a fraction or a decimal such as 0.5, 1 / 2, 1 / 3, etc. Each embodiment is applicable to a technique for only enlarging the resolution of the image. Also, each embodiment is applicable to a technique for only reducing the resolution of the image. That is, each embodiment is applicable to any technique for changing the resolution of the image.

[0021] Also, if the current resolution is known, the magnification may be represented by the resolution (image size) after the change. For example, when changing an image equivalent to 4K (3840 pixels × 2160 pixels) to an image equivalent to full high vision (1920 pixels × 1080 pixels), it may be expressed as a magnification of 1 / 2 (0.5), or the magnification may be expressed as the resolution after the change, "1920 pixels × 1080 pixels".

[0022] The encoding unit 160 compresses the image with the converted resolution using a quantization step, generates it as a bitstream, and stores it in the RAM 103 via the bus 110. In this embodiment, the image compression process by the encoding unit 160 is assumed to be performed based on the VVC standard, but the image compression process is not limited to this.

[0023] The symbol amount control unit 170 acquires quantization steps, and based on the quantization steps, controls by changing the scaling ratio in the enlargement / reduction unit 150 and the quantization steps in the encoding unit 160. Thereby, the symbol amount control unit 170 controls so that the bit rate of the bit stream output from the encoding unit 160 is maintained in the vicinity of the target value.

[0024] The IP communication unit 180 is connected to the network 190 via a LAN. The bit stream stored in the RAM 103 is distributed from this IP communication unit 180 through the network 190.

[0025] Part or all of the enlargement / reduction unit 150, the encoding unit 160, the symbol amount control unit 170, and the IP communication unit 180 may be realized as functions of a processor that has read a program, or may be realized by a circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0026] FIG. 2 is a block diagram showing the configuration of the encoding unit 160 of the present embodiment.

[0027] The image analysis unit 200 analyzes the angle-of-view value of the input frame, and outputs the analyzed result as image analysis information to the RPR control information generation unit 210 and the bit stream generation unit 270. The image analysis unit 200 outputs the image as a tile image combined with tile information for dividing the image into spatial regions according to image characteristics and external inputs to the prediction unit 220.

[0028] The RPR control information generation unit 210 generates information such as the scaling ratio and the offset position of the motion vector necessary for decoding using RPR (Reference Picture Resampling).

[0029] The prediction unit 220 performs prediction processing on the image data in tile units to generate prediction image data, and outputs and stores it in the frame memory 250. The prediction processing may be processing based on at least any one of intra prediction which is in-frame prediction and inter prediction which is inter-frame prediction. Further, the prediction unit 220 calculates a prediction error from the input image data and the prediction image data, and outputs it to the conversion / quantization unit 230. Also, the prediction unit 220 may output information necessary for prediction, such as information on a prediction mode and a motion vector, to the image reproduction unit 240 and the entropy encoding unit 260 together with the prediction error. Hereinafter, the information necessary for this prediction is referred to as prediction information.

[0030] The conversion / quantization unit 230 performs orthogonal conversion on the prediction error in block units to obtain conversion coefficients. The conversion / quantization unit 230 obtains a predetermined quantization step or a quantization step changed by the coding amount control unit 170. The conversion / quantization unit 230 performs quantization on the conversion coefficients using the quantization step to obtain quantization coefficients. The conversion / quantization unit 230 outputs the quantization coefficients to the entropy encoding unit 260 and the inverse quantization / inverse conversion unit 231.

[0031] The inverse quantization / inverse conversion unit 231 inverse-quantizes the quantization coefficients output from the conversion / quantization unit 230 to reproduce the conversion coefficients, and further performs inverse orthogonal conversion to reproduce the prediction error. The inverse quantization / inverse conversion unit 231 outputs the reproduced prediction error to the image reproduction unit 240.

[0032] The image reproduction unit 240 appropriately refers to the frame memory 250 based on the prediction information output from the prediction unit 220 to obtain prediction image data. The image reproduction unit 240 generates reproduced image data from the prediction image data and the input prediction error. The image reproduction unit 240 outputs and stores the reproduced image data in the frame memory 250.

[0033] The frame memory 250 stores the predicted image data generated by the prediction unit 220, the image data reproduced by the image reproduction unit 240, and the filtered image data that has undergone in-loop filter processing by the in-loop filter unit 251.

[0034] The in-loop filter unit 251 performs in-loop filter processing such as a deblocking filter and a sample adaptive offset on the reproduced image data stored in the frame memory 250. The in-loop filter unit 251 outputs and stores the filtered image data that has undergone in-loop filter processing in the frame memory 250.

[0035] The entropy encoding unit 260 entropy-encodes the quantized coefficients output from the transform / quantization unit 230 and the prediction information output from the prediction unit 220 to generate coded data, and outputs it to the bitstream generation unit 270.

[0036] The bitstream generation unit 270 encodes the image analysis information output by the image analysis unit 200 and the RPR control information output by the RPR control information generation unit 210 to generate header coded data. The bitstream generation unit 270 further combines the coded data output from the entropy encoding unit 260 with the header coded data to generate and output a bitstream.

[0037] The image encoding operation in the encoding unit 160 will be described below. In this embodiment, the configuration is such that moving image data is input in units of frames.

[0038] The image analysis unit 200 acquires the image data for one frame and obtains the viewing angle change value. The viewing angle change value is the ratio of the viewing angles of the reference frame and the frame to be encoded when a certain arbitrary frame is used as the reference frame.

[0039] Next, the RPR control information generation unit 210 signals by setting the sps_ref_pic_resampling_enabled_flag in the Sequence parameter set (hereinafter referred to as SPS) to 1 to indicate the use of RPR. Also, the RPR control information generation unit 210 calculates the pixels of the input frame and stores the number of pixels in the vertical direction and the number of pixels in the horizontal direction of the luminance as pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples in the Picture Parameter Set (hereinafter referred to as PPS), respectively.

[0040] The prediction unit 220 cuts out the image data input from the image analysis unit 200 into a plurality of blocks and performs prediction processing in units of blocks. As a result of the prediction processing, the prediction unit 220 generates a prediction error and outputs it to the transform / quantization unit 230. Also, the prediction unit 220 generates prediction information and outputs it to the entropy encoding unit 260 and the image reproduction unit 240.

[0041] Here, the prediction processing executed by the prediction unit 220 and the prediction information output by the prediction unit 220 will be described in more detail. In image encoding technologies such as VVC, in order to reduce the data amount of the encoded bitstream while maintaining the image quality of the reproduced image, prediction processing is used to predict the pixels of the block to be encoded using the pixels of the encoded block. The prediction processing includes intra prediction using the pixels of the encoded block in the same frame, and inter prediction using the pixels of the blocks in different encoded frames. Also, in VVC, a technology called RPR is standardized so that decoding can be performed even when the resolution of the reference encoded frame and the resolution of the frame to be encoded are different in inter prediction.

[0042] Here, as an explanation of RPR, inter prediction when the resolution of the reference encoded frame and the resolution of the frame to be encoded are different will be further described.

[0043] Figure 3 shows an example of frames when the resolution is changed in the middle of a stream by RPR. Frames 301 to 309 are arranged in chronological order in this sequence. Frame 301 is an I-frame composed only of intra references. Frames 302 to 309 are P-frames that include a reference to the previous frame. Between frames 303 and 304, the vertical and horizontal dimensions of the frame are each reduced to 2 / 3. Similarly, between frames 306 and 307, the vertical and horizontal dimensions of the frame are each reduced to 2 / 3. The vertical and horizontal dimensions after frame 307 are 4 / 9 each compared to frame 301. Specifically, for example, the number of pixels from frame 301 to frame 303 is 2880x1620. The number of pixels from frame 304 to frame 306 is 1920x1080. The number of pixels from frame 304 to frame 306 is 1280x720.

[0044] At the location where the resolution change occurs, the encoded frame of the reference target is scaled and inter prediction is performed with the same resolution as the resolution of the frame to be encoded. An example of the method for scaling the encoded frame of the reference target is shown below. For simplicity, the explanation is given only for the luminance value. Regarding the color difference, since the discussion can proceed in the same way as for the luminance value considering the sampling number, the explanation is omitted.

[0045] (Step 1) The vertical scaling ratio and the horizontal scaling ratio are obtained as scalingRatio[0] and scalingRatio[1], respectively. Hereinafter, scalingRatio[0] and scalingRatio[1] are collectively denoted as scalingRatio[x]. scalingRatio[x] is determined by the ratio between the size of the scaling window of the encoded frame of the reference target and the scaling window of the frame to be encoded. In this embodiment, scalingRatio[x] is obtained as RPR control information from the RPR control information generation unit 210.

[0046] (Step 2) Figures 4 to 6 show the table of filter coefficients for the first embodiment. The interpolation filter used for scaling is determined. For example, the coefficients of the interpolation filter are selected according to the value of scalingRatio[x]. When scalingRatio[x] exceeds a ratio of 1.75 times, for example, the filter coefficients in Table 1 of Figure 4 are used. When scalingRatio[x] is less than 1.75 times and exceeds 1.25 times, for example, the filter coefficients in Table 2 of Figure 5 are used. Otherwise, for example, the filter coefficients in Table 3 of Figure 6 are used. The coefficients of the interpolation filter from each table are determined by the sample position p to be calculated. The filter coefficients of Tables 1 to 3 in Figures 4 to 6 are the coefficients of the interpolation filter defined for luma scaling in the VVC standard. Table 1 is created based on the cut-off frequency when the ratio is 2 times, and Table 2 is created based on the cut-off frequency when the ratio is 1.5 times.

[0047] The sample position p is an integer in the range from 0 to 15, which is the numerator value when the minimum unit of the sample is divided into 1 / 16 units. For example, when a certain sample point A and a certain sample point A + 1 are divided into 16 parts, for example, the coefficients of the interpolation filter at the third sample point (A + 3 / 16, p = 3) are, referring to Table 1, fL[3][i]=[-4, -1, 16, 29, 23, 7, -4, -2].

[0048] Hereinafter, when the coefficients of the interpolation filter obtained in this way are used horizontally, they are denoted as fLH[p][i] (= fL[p][i]), and when used vertically, they are denoted as fLV[p][i] (= fL[p][i]).

[0049] (Step 3) Resample the reference image generated from the coded frame to be referenced to the same resolution as the frame to be coded according to scalingRatio[x]. For example, when scaling the frame to be coded by scalingRatio[x], determine the position of each pixel with 1 / 16 pixel accuracy. Also, interpolate the reference image 16 times using the filter obtained in Step 2. For example, when scalingRatio[x] ≥ 1.25, obtain new sampling points (x3 + px, y3 + py) using Equation 1 and Equation 2. Here, using the notation where the coordinates of the pixel of interest in the reference image are (xi, yi), the coordinates of the pixel adjacent to the left are (x(i - 1), yi), and the coordinates of the pixel adjacent to the bottom are (xi, y(i - 1)). Also, px and py are integers modulo 16, and are the indices of the coordinates obtained by dividing the coordinates of the adjacent pixels in the horizontal and vertical directions by 16, respectively, and L(x, y) represents the luminance value at the coordinates (x, y). Let a be a normalization constant.

[0050]

Number

[0051]

Number

[0052] Here, fLH[p][i] and fLV[p][i] are generated from Table 2. Also, yn are the coordinate values from y0 to y7. The samples generated as described above will be referred to as the upsampled image. Thinning out this upsampled image generates a reference image after resampling with the same resolution as the frame to be coded from the reference image.

[0053] Figures 7 and 8 are schematic diagrams for simply explaining the thinning method. The thinning method does not vary in the vertical and horizontal directions. Therefore, for simplicity, only the horizontal direction of the thinning method will be described. The reference image 701 shown in Fig. 7(a) and the frame to be coded 702 shown in Fig. 7(b) are composed of a plurality of pixel blocks 703 arranged in two directions. It is assumed that each pixel block has a defined pixel value. The reference image 701 is a 6×6 image, and the frame to be coded 702 is a 4×4 image. The origin is the upper left vertex of the entire image, and the coordinates of each pixel value are the coordinates of the upper left vertex of the pixel block.

[0054] Here, the coordinate value of each pixel of the frame to be coded is multiplied by scalingRatio[x] (3 / 2 in the example of Fig. 7). For the x coordinates of the frame to be coded 702 being (0, 1, 2, 3), the coordinate values (0, 3 / 2, 6 / 2, 9 / 2) can be calculated respectively, and the calculated coordinate values are stored in, for example, the coordinate array H[x]. Here, the image 704 shown in Fig. 7(c) is resampled so that the number of pixels is the same as that of the frame to be coded. The pixel value of each pixel block of the resampled image 704 may be configured using the pixel value corresponding to the coordinates of the enlarged image from the upsampled image.

[0055] Here, an example of the method of constructing the image 704 will be described with reference to Fig. 8. The pixel block sequence 801 shown in Fig. 8(a) depicts only the pixel blocks of the image 704 at y = 0. The pixel block sequence 804 shown in Fig. 8(b) depicts only the pixel blocks of the reference image 701 at y = 0. Here, let the luminance value of the reference image 701 be Y[x] and the luminance value of the resampled image be Y’[x]. Here, x is the x coordinate value of the pixel block. Then, the luminance value at the coordinate position x of the resampled image can be obtained by referring to H[x] in the coordinate array 803 shown in Fig. 8(c), obtaining the coordinate position of the reference image, and then obtaining the luminance value at that coordinate position. The luminance value, limited to the description of only the x coordinate, can be written as follows.

[0056]

Equation

[0057] For example, since the coordinate of the desired luminance value of pixel block 802 is 1, set x = 1, obtain the coordinate matrix H[1] = 3 / 2, and it is sufficient to obtain the value of Y[3 / 2]. Pixel block column 805 is a diagram that clearly shows sampling points where pixel blocks are interpolated by an interpolation filter between the coordinate value 1 and the coordinate value 2 of pixel block column 804. Since the reference image has been interpolated in 1 / 16 increments like pixel block column 805 in advance, it is sufficient to obtain the luminance value of the coordinate corresponding to 3 / 2 (=1 + 8 / 16) of pixel block 806. In this way, the reference image after resampling at the same resolution can be configured by thinning out the upsampled image of the reference image except for the pixel values corresponding to the enlarged pixel positions of the frame to be encoded.

[0058] Inter prediction is a process of predicting the pixels of a block to be encoded by referring to the pixels of an encoded frame or, when the number of pixels of the encoded frame and the number of pixels of the frame to be encoded are different, the pixels of the resampled image configured by the above method. Here, for simplicity, the encoded frame and the resampled image are collectively referred to as an inter prediction target image. For example, when there is no motion between the encoded frame to be referred to and the inter prediction target image, the pixels of the block to be encoded are predicted using the pixels at the same position in the inter prediction target image. In such a case, a (0, 0) motion vector indicating no motion is included as prediction information. On the other hand, when motion occurs between frames for the block to be encoded, the motion vector (MVx, MVy) is included in the prediction information.

[0059] FIG. 9 is a flowchart of the amount-of-code control for one frame of the amount-of-code control unit 170 according to the first embodiment. In the embodiment described from here, the resolution of the stream is set such that the standard resolution is 2880x1620, the maximum resolution changed by the amount-of-code control is 6480x3645, and the minimum resolution is 1280x720.

[0060] In step S900, the quantization control unit 170 obtains the size output from the compressed bitstream generation unit 270 of the previous frame.

[0061] In step S910, a CBR process is performed in which the quantization control unit 170 determines the quantization step (Q value) of the next frame based on the size of the previous frame obtained in step S900. Since various known examples of the CBR processing method are widely known, they are omitted here.

[0062] In step S920, the quantization control unit 170 executes a branch process according to the Q value obtained in step S910. When the quantization control unit 170 determines that the Q value is a value greater than or equal to a first threshold (here, 48 or more) indicating that the Q value is approaching the maximum value of 51, it executes step S930. When the quantization control unit 170 determines that the Q value is a value less than or equal to a second threshold (here, 4 or less) smaller than the first threshold indicating that the Q value is approaching the minimum value of 0, it executes step S940. The quantization control unit 170 determines whether the Q value is away from both the maximum value and the minimum value. For example, when the quantization control unit 170 determines that the Q value is a value greater than or equal to a fourth threshold (here, 16 or more) greater than the second threshold and less than or equal to a third threshold (here, 36 or less) smaller than the first threshold, it executes step S950. When the quantization control unit 170 determines that the Q value is in other transitional states (4 < Q < 16 or 36 < Q < 48), it uses the Q value obtained in step S910 as it is and ends the quantization control process of the frame without changing the scaling factor.

[0063] In step S930, when the quantization control unit 170 determines that the current resolution has already reached the minimum resolution of 1280x720, since the resolution cannot be further decreased, the Q value obtained in step S910 is used as it is, and the quantization control process for the frame is terminated. When the quantization control unit 170 determines that the current resolution is not the minimum, in step S931, the scaling factor is changed to reduce the resolution to 2 / 3, and it is sent to the scaling unit 150. The reason for choosing 2 / 3 as the scaling factor here is that, due to the relationship between the cut-off frequencies of the filters in FIG. 5 prepared in the VVC standard, reducing the resolution to 2 / 3 may be efficient. However, even if other scaling factors are selected, it does not deviate from the gist of this embodiment. Subsequently, in step S932, the quantization control unit 170 subtracts 6 from the Q value obtained in step S910 so as to match the fact that the resolution has been reduced to 2 / 3 and the area ratio has become 4 / 9 (about 1 / 2). Regarding how to determine the Q value to be adjusted, since the optimal value may vary depending on the use case, it is not mentioned here. Steps S931 and S932 are an example of the first change process.

[0064] In step S940, when the quantization control unit 170 checks the current resolution and determines that it has already reached the maximum resolution of 6480x3645, since the resolution cannot be further increased, the Q value obtained in step S910 is used as it is, and the quantization control process for the frame is terminated. When the quantization control unit 170 determines that the current resolution is not the maximum, in step S941, the scaling factor is changed to increase the resolution to 3 / 2, and it is sent to the scaling unit 150. Subsequently, in step S942, the quantization control unit 170 adds 6 to the Q value obtained in step S910 so as to match the fact that the resolution has been increased to 3 / 2 and the area ratio has become 9 / 4 (about 2 times). Steps S941 and S942 are an example of the second change process.

[0065] Step S950 is for returning to the standard resolution direction when the Q value has settled down from a state where the resolution is different from the standard resolution of 2880x1620 through the processing of steps S931 and S941 in the past frame. Therefore, when the coding amount control unit 170 determines that the resolution is larger than the standard resolution, it proceeds to step S931 to control in the direction of decreasing the resolution. When the coding amount control unit 170 determines that the resolution is smaller than the standard resolution, it proceeds to step S941 to control in the direction of increasing the resolution. When the coding amount control unit 170 determines that the resolution has already become the standard resolution, it uses the Q value obtained in step S910 as it is and ends the coding amount control process for the frame.

[0066] Figure 10 is a diagram of a control example when the Q value is near the maximum value in the first embodiment. Also, Figure 10 is a diagram simulating an actual use case of coding amount control. The three graphs on the left are based on conventional control, and the three graphs on the right are those to which this embodiment is applied. The topmost graph of complexity (motion amount) simulates the distribution of frequency components and the motion amount within the screen that affect the compression efficiency, and there is no standard index that can be specifically quantified. The upward-sloping graph shows that the input image is changing such that the coding efficiency is decreasing. The middle graph shows the transition of the Q value as a result of performing CBR control. In the conventional control, after the Q value reaches the maximum value, it just sticks to the maximum value. As a result, in the graph showing the transition of the bit rate at the bottom, in the conventional control, the target value of the bit rate is exceeded after the Q value sticks to the maximum value. In this case, for the user, when the target bit rate set for the input image is too low, a notice will be given that the target value may be exceeded.

[0067] In the graphs of this embodiment on the right, when the Q value reaches a threshold near the maximum value, the resolution is reduced to 2 / 3 times, and at that time, the Q value is also decreased by 6, so that the bit rate continues to stay near the target value.

[0068] FIG. 11 is a diagram of a control example when the Q value is near the minimum value in the first embodiment. FIG. 11, contrary to FIG. 10, shows a state where the complexity gradually decreases and the Q value by CBR control approaches the minimum value. In the conventional example, when the Q value sticks to the minimum value, the bit rate falls below the target value. In this case, conventionally, it has been done to allow the bit rate to fall below the target value or to add dummy data called Filler to make it appear as if it stays near the target value. Filler is just filled and has no meaningful information and is discarded during decoding.

[0069] In the graph of this embodiment on the right, when the Q value reaches a threshold near the minimum value, the resolution is increased by 3 / 2 times, and at that time, the Q value is also increased by 6, so that the bit rate continues to stay near the target value.

[0070] As described so far, there is an effect that the range in which the bit rate can be kept constant with respect to changes in the input image is wider than in the conventional control.

[0071] In this embodiment, based on the Q value (quantization step), it is determined whether or not to change the scaling ratio, which is the ratio of enlargement or reduction of the resolution. Thus, it appropriately determines whether or not to change the scaling ratio based on the Q value and changes the resolution. Thereby, this embodiment can appropriately keep the bit rate at the target bit rate based on the Q value.

[0072] When this embodiment determines to change the scaling ratio based on the Q value, it changes the Q value. Thereby, this embodiment can suppress the Q value from sticking to the maximum value or the minimum value.

[0073] In this embodiment, together with the Q value, it is determined whether or not to change the scaling ratio based on the resolution of the image, so that the accuracy of the necessity determination can be further improved.

[0074] This embodiment is a case where the Q value is 48 or more, that is, in the vicinity of the maximum value, and when the resolution is not the minimum value, the scaling factor is decreased and the Q value is decreased. Further, this embodiment is a case where the Q value is 4 or less, that is, in the vicinity of the minimum value, and when the resolution is not the maximum value, the scaling factor is increased and the Q value is increased. Thereby, this embodiment can accurately maintain the bit rate at the target bit rate according to the Q value and the resolution.

[0075] This embodiment decreases the resolution and the Q value when the Q value is in the transient period and the resolution is larger than the standard resolution. Further, this embodiment increases the resolution and the Q value when the Q value is in the transient period and the resolution is smaller than the standard resolution. Thereby, this embodiment can adjust the bit rate in advance, so that it is possible to suppress the bit rate from deviating from the target bit rate.

[0076] (Second Embodiment) The image distribution apparatus according to the second embodiment will be described with reference to FIGS. 12 and 13. Since FIGS. 1 to 8 are common to the first embodiment, the description thereof will be omitted. FIG. 12 is a flowchart of the coding amount control according to the second embodiment.

[0077] FIG. 12 is a flowchart of the coding amount control according to the second embodiment. In FIG. 12, with respect to the flowchart of FIG. 9 of the first embodiment, the processes from step S1200 to S1202 and from step S1210 to S1212 are added. The processes added in this embodiment are considered in view of the response when the complexity suddenly increases.

[0078] In step S1200, the coding amount control unit 170 determines whether the Q value has increased rapidly as a result of the CBR control in step S910. As a determination method, it is determined whether the increase amount or the average of the increase amounts of the Q values in the past several frames is equal to or greater than the threshold increase amount. An example of the threshold increase amount is 6. This criterion is an example, and using other criteria does not deviate from the gist of this embodiment.

[0079] When the increase amount of the Q value or the average of the increase amounts is less than the threshold increase amount, the symbol amount control unit 170 determines that the Q value is not increasing rapidly and executes the same processing as after step S931 of the first embodiment. When the increase amount of the Q value or the average of the increase amounts is equal to or greater than the threshold increase amount, the symbol amount control unit 170 determines that the Q value is increasing rapidly, changes the scaling factor so as to reduce the resolution by a factor of 1 / 2 in step S1201, and sends it to the enlargement / reduction unit 150. Subsequently, in step S1202, the symbol amount control unit 170 subtracts 12 from the Q value obtained in step S910 so as to match the fact that the resolution has been halved to 1 / 2 and the area ratio has become 1 / 4 in step S1201. Note that when the resolution is slightly larger than the minimum resolution, if the process of reducing the resolution by a factor of 1 / 2 is executed, the resolution will become smaller than the minimum resolution, but this is allowed in this embodiment. Steps S1201 and S1202 are an example of the third change process.

[0080] In step S1210, the symbol amount control unit 170 determines whether the resolution has been changed to 1 / 2 in step S1201. When the symbol amount control unit 170 determines that the resolution has been changed to 1 / 2, it executes steps S1211 and subsequent steps, and otherwise, executes the same processing as after step S941 of the first embodiment.

[0081] In step S1211, the symbol amount control unit 170 changes the scaling factor so as to increase the resolution by a factor of 4 / 3 and sends it to the enlargement / reduction unit 150. The meaning of 4 / 3 is that by multiplying by 4 / 3 after reducing the resolution by a factor of 1 / 2 in step S1201, it becomes 1 / 2 × 4 / 3 = 2 / 3, which is equivalent to executing step S931 instead of step S1201. In step S1212, the symbol amount control unit 170 adds 5 to the Q value obtained in step S910 so as to match the fact that the resolution has been increased by a factor of 4 / 3 in step S941 and the area ratio has become 16 / 9 (about 1.7 times). Steps S1211 and S1212 are an example of the fourth change process.

[0082] FIG. 13 is a diagram of a control example near the maximum value where the Q value of the second embodiment has increased rapidly. FIG. 13 is a diagram simulating the use case of the control in FIG. 12. As shown in the graph of complexity, the complexity increases rapidly, and as shown in the graph of the Q value, the Q value increases rapidly near the maximum value. In this embodiment, in such a case, as shown in the flowchart of FIG. 12, control is performed to reduce the resolution to 1 / 2 and reduce the Q value by 12. The graph of the Q value on the right side of FIG. 13 shows this state. As shown in FIG. 13, in the case of only CBR, when the Q value increases rapidly and reaches the maximum value, the bit rate deviates greatly from the target bit rate. On the other hand, this embodiment can suppress the bit rate from deviating from the target bit rate even when the Q value increases rapidly and reaches the maximum value.

[0083] In this embodiment, there is an effect that it is easy to respond even when the Q value increases rapidly.

[0084] This embodiment is a case where the Q value (quantization step) is 48 or more, that is, near the maximum value, and the resolution is not the minimum value, and when the Q value is increasing rapidly, the resolution and the Q value are made smaller than when it is not increasing rapidly. Thereby, this embodiment can keep the bit rate at the target bit rate even when the Q value is increasing rapidly.

[0085] This embodiment is a case where the Q value is in the transient period, the resolution is smaller than the standard resolution, and the scaling factor is 1 / 2 set in step S1201, and the resolution and the Q value are increased with values smaller than steps S941 and S942. Thereby, this embodiment can adjust the bit rate in advance even when the Q value is in the transient period and the resolution is smaller than the standard resolution. As a result, this embodiment can suppress a rapid change in the bit rate and suppress the bit rate from deviating from the target bit rate.

[0086] (Other Examples) The present invention can also be realized by supplying a program that implements one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. It can also be realized by a circuit (for example, ASIC) that implements one or more functions.

[0087] Further, in the above-described embodiment, it is determined whether or not to change the magnification factor and the quantization step based on the quantization step and the resolution. However, it may be determined whether or not to change the magnification factor and the quantization step based only on the quantization step.

[0088] The disclosure of this specification includes the following image processing apparatus, image processing method, and program. (Item 1) Change means for changing the resolution of an image, Encoding means for quantizing and compressing the image with the converted resolution based on a quantization step, Control means for controlling the change in the resolution of the image by the change means based on the quantization step, An image processing apparatus comprising the same. (Item 2) The control means determines whether to change the magnification factor when changing the resolution of the image based on the quantization step, When the control means determines to change the magnification factor, the control means changes the quantization step The image processing apparatus according to Item 1, characterized in that. (Item 3) The control means determines whether to change the magnification factor when changing the resolution of the image based on the quantization step and the resolution of the image, according to the image processing apparatus according to Item 1 or Item 2, characterized in that. (Item 4) The control means When it is determined that the quantization step is equal to or greater than a first threshold value and that the resolution of the image is not the minimum value of a predetermined resolution, a first change process of reducing the magnification and the quantization step when changing the resolution of the image is executed. When it is determined that the quantization step is equal to or less than a second threshold value that is less than the first threshold value and that the resolution of the image is not the maximum value of a predetermined resolution, a second change process of increasing the magnification and the quantization step is executed. The image processing apparatus according to any one of Items 1 to 3, characterized in that. (Item 5) The control means When it is determined that the quantization step is equal to or less than a third threshold value that is less than the first threshold value and equal to or greater than a fourth threshold value that is greater than the second threshold value, When it is determined that the resolution of the image is greater than a predetermined standard resolution, the first change process is executed. When it is determined that the resolution of the image is less than a predetermined standard resolution, the second change process is executed. The image processing apparatus according to Item 4, characterized in that. (Item 6) The control means When it is determined that the quantization step is equal to or greater than a first threshold value and that the resolution of the image is not the minimum value of a predetermined resolution, When it is determined that the increase amount of the quantization step or the average of the increase amounts is less than a threshold increase amount, a first change process of reducing the magnification and the quantization step when changing the resolution of the image is executed. When it is determined that the increase amount of the quantization step or the average of the increase amounts is equal to or greater than the threshold increase amount, a third change process of reducing the magnification and the quantization step more than the first change process is executed. When it is determined that the quantization step is equal to or less than a second threshold value that is less than the first threshold value and that the resolution of the image is not the maximum value of a predetermined resolution, a second change process of increasing the magnification and the quantization step is executed. The image processing apparatus according to any one of Items 1 to 5, characterized in that... (Item 7) The control means when it is determined that the quantization step is equal to or less than a third threshold value smaller than the first threshold value and equal to or greater than a fourth threshold value larger than the second threshold value, when it is determined that the resolution of the image is larger than a predetermined standard resolution, the first change process is executed, when it is determined that the resolution of the image is smaller than a predetermined standard resolution, when it is determined that the magnification is not a value changed by the third change process, the second change process is executed, when it is determined that the magnification is a value changed by the third change process, a fourth change process of increasing the magnification and the quantization step is executed, the increase amounts of the magnification and the quantization step in the second change process are larger than the increase amounts of the magnification and the quantization step in the fourth change process The image processing apparatus according to Item 6, characterized in that... (Item 8) When the control means increases the resolution, it changes the magnification when changing the resolution of the image based on a value prepared when performing RPR in the VVC standard. The image processing apparatus according to any one of Items 1 to 7, characterized in that... (Item 9) The image processing apparatus according to any one of Items 1 to 8, imaging means for imaging a subject to generate an image, communication means for distributing the compressed image, An image distribution apparatus, characterized by comprising... (Item 10) a change step of changing the resolution of the image, an encoding step of quantizing and compressing the image with the converted resolution based on a quantization step, a control step of controlling the change in the resolution of the image in the change step based on the quantization step An image processing method characterized by comprising (Item 11) A program for causing a computer to function as each means of the image processing apparatus according to any one of Items 1 to 8.

[0089] The invention is not limited to the above embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Therefore, claims are attached to disclose the scope of the invention.

Explanation of Reference Numerals

[0090] 100... Image distribution apparatus, 105... Image processing apparatus, 120... Imaging unit, 150... Enlargement / reduction unit, 160... Encoding unit, 170... Code amount control unit.

Claims

1. changing means for changing the resolution of an image; encoding means for quantizing and compressing the image with the converted resolution based on a quantization step; control means for controlling the change in the resolution of the image by the changing means based on the quantization step; An image processing apparatus comprising the same.

2. The control means determines whether to change the magnification when changing the resolution of the image based on the quantization step, and when the control means determines to change the magnification, the control means changes the quantization step The image processing apparatus according to claim 1, characterized in that.

3. The image processing apparatus according to claim 1, characterized in that the control means determines whether to change the magnification when changing the resolution of the image based on the quantization step and the resolution of the image.

4. The control means, when it is determined that the quantization step is equal to or greater than a first threshold value and the resolution of the image is not the minimum value of a predetermined resolution, executes a first change process of reducing the magnification and the quantization step when changing the resolution of the image, when it is determined that the quantization step is equal to or less than a second threshold value smaller than the first threshold value and the resolution of the image is not the maximum value of a predetermined resolution, executes a second change process of increasing the magnification and the quantization step The image processing apparatus according to claim 1, characterized in that.

5. The control means, when it is determined that the quantization step is equal to or less than a third threshold value smaller than the first threshold value and equal to or greater than a fourth threshold value greater than the second threshold value, when it is determined that the resolution of the image is greater than a predetermined standard resolution, executes the first change process, when it is determined that the resolution of the image is smaller than a predetermined standard resolution, executes the second change process The image processing apparatus according to claim 4, characterized in that.

6. The control means, when it is determined that the quantization step is equal to or greater than a first threshold value and the resolution of the image is not the minimum value of a predetermined resolution, when it is determined that the increase amount of the quantization step or the average of the increase amounts is less than a threshold increase amount, executes a first change process of reducing the magnification and the quantization step when changing the resolution of the image, When it is determined that the increased number of quantization steps or the average of the increased numbers is equal to or greater than a threshold increased number, execute a third change process that makes the magnification and the quantization step smaller than the first change process. When it is determined that the quantization step is equal to or less than a second threshold smaller than the first threshold and that the resolution of the image is not the maximum value of a predetermined resolution, execute a second change process that increases the magnification and the quantization step. The image processing apparatus according to claim 1, characterized in that.

7. The control means This is the case where it is determined that the quantization step is equal to or less than a third threshold smaller than the first threshold and equal to or greater than a fourth threshold greater than the second threshold. When it is determined that the resolution of the image is greater than a predetermined standard resolution, execute the first change process. This is the case where it is determined that the resolution of the image is smaller than a predetermined standard resolution. When it is determined that the magnification is not a value changed by the third change process, execute the second change process. When it is determined that the magnification is a value changed by the third change process, execute a fourth change process that increases the magnification and the quantization step. The increased numbers of the magnification and the quantization step in the second change process are greater than the increased numbers of the magnification and the quantization step in the fourth change process. The image processing apparatus according to claim 6, characterized in that.

8. When increasing the resolution, the control means changes the magnification when changing the resolution of the image based on a value prepared when performing RPR according to the VVC standard. The image processing apparatus according to claim 1, characterized in that.

9. An image processing apparatus according to any one of claims 1 to 8, Imaging means for imaging a subject to generate an image, Communication means for distributing the compressed image, An image distribution apparatus, characterized by comprising.

10. A change step of changing the resolution of the image, An encoding step of quantizing and compressing the image whose resolution has been converted based on a quantization step, A control step of controlling the change in the resolution of the image in the change step based on the quantization step, An image processing method, characterized by comprising.

11. A program for causing a computer to function as each means of the image processing apparatus according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Jikihetsudoshiirudokeesuno naimenkakoho

    JP1976097574A

  • Video stream supply system and apparatus, and video stream receiving apparatus

    JP2007088539A