Encoding device, encoding method, and computer program

The encoding device stabilizes video streams by converting frame resolutions before encoding when encoding and display orders match, and referencing without conversion when they differ, addressing image instability in VVC with RPR.

JP2026009656APending Publication Date: 2026-01-21CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024109688
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-08
Publication Date
2026-01-21

AI Technical Summary

Technical Problem

Existing video encoding methods using Versatile Video Coding (VVC) with Reference Picture Resampling (RPR) cause image instability due to frequent resolution conversions, especially in streams with backward-referenced B frames, leading to mixed display and encoding orders and unstable video resolution.

Method used

An encoding device that converts the resolution of image frames before encoding if the encoding and display orders are the same, and references frames without conversion if they are different, using RPR to stabilize the stream delivery.

Benefits of technology

Reduces screen disturbances and delivers stable video streams by managing resolution changes effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026009656000001_ABST
    Figure 2026009656000001_ABST
Patent Text Reader

Abstract

To provide an encoding device or the like capable of performing stable stream distribution by reducing disturbance or the like of a screen accompanying switching of resolution.SOLUTION: The encoding apparatus includes an encoding unit configured to, when sequentially encoding a plurality of image frames, if an encoding order and a display order of the image frames are the same, convert a resolution of the image frame having a different resolution and refer to the image frame to encode a difference frame, and if the encoding order is different from the display order, refer to the image frame having the different resolution without converting the resolution to encode the difference frame.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an encoding device, an encoding method, a computer program, and the like. [Background technology]

[0002] When compressing moving images, the amount of code after compression varies depending on the characteristics of the input video. On the other hand, when the network that transmits the video stream is the Internet or a wireless network, the communication bandwidth fluctuates depending on the situation. If the network communication bandwidth is narrower than the variable amount of code after compression of the input video, the compressed coded data cannot be transmitted in its entirety.

[0003] In Patent Document 1, the resolution of the video is converted to match the processing capacity of the decoder, and the quantization step is changed based on packet loss in the network.In Patent Document 2, the timing for changing the resolution is determined based on the bit amount of the stream stored in the transmission means, and the encoding mode and quantization step of the encoding means are controlled according to this timing.

[0004] Furthermore, the Versatile Video Coding (VVC) coding method (hereinafter referred to as VVC) is known as a coding method for compressing and recording moving images. In VVC, a technique called Reference Picture Resampling (RPR) (hereinafter referred to as RPR) is introduced to improve coding efficiency.

[0005] RPR is a technology that allows the decoding target image and an image with a different resolution to be used as reference images, making it possible to change the resolution even in the case of inter-frame compression. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2007-088539 [Patent Document 2] Japanese Patent Application Laid-Open No. 2013-214894 Summary of the Invention [Problem to be solved by the invention]

[0007] In the above Patent Document 1, the resolution or quantization step is changed so as to prevent the destination from being unable to decode or receive the signal completely, but the new VVC encoding method is not taken into consideration.

[0008] In the above Patent Document 2, the resolution and quantization step are changed according to the communication bandwidth, but the timing of switching the resolution is for each Group of Picture (hereinafter, GOP).Since the RPR adopted in the VVC encoding method is not taken into consideration, only I frames are supported.

[0009] On the other hand, using RPR makes it possible to reduce the resolution of any frame within a GOP to reduce data volume depending on the communication bandwidth, but frequent resolution conversion can cause image distortion. In particular, streams that include backward-referenced B frames have different display and encoding orders, causing the original resolution and reduced resolution to become mixed up and making the stream unstable.

[0010] As described above, when resolution conversion using RPR occurs on a stream whose coding order differs from its display order due to a coding order change, there is a problem in that the video resolution during display is unstable.

[0011] In view of the above-mentioned problems, an object of the present invention is to provide an encoding device and the like that can reduce screen disturbances that occur when switching resolutions and that can deliver stable streams. [Means for solving the problem]

[0012] In order to achieve the above object, the encoding device according to claim 1 comprises: The image encoding device is characterized by having an encoding means that, when sequentially encoding a plurality of image frames, if the encoding order and display order of the image frames are the same, converts the resolution of the image frames having different resolutions before referring to them and encodes the difference frames, and, if the encoding order and display order are different, references the image frames having different resolutions without converting their resolution and encodes the difference frames. [Effects of the Invention]

[0013] According to the present invention, it is possible to provide an encoding device that can reduce screen disturbances that occur when switching resolutions and that can deliver stable streams. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a functional block diagram illustrating an example of the configuration of a video distribution device 100 according to a first embodiment. [Figure 2] 2 is a functional block diagram illustrating an example of the configuration of a compression encoding unit 130 according to the first embodiment. FIG. [Figure 3] 10A and 10B are diagrams illustrating an example of coding order information generated by the coding order determination unit 160 according to the first embodiment. [Figure 4] FIG. 10 is a diagram showing an example in which the video streams of the first embodiment are rearranged in coding order. [Figure 5] 10A and 10B are diagrams for explaining an example of converting the resolution of a video stream rearranged in the encoding order according to the first embodiment. [Figure 6] FIG. 10 is a diagram showing an example of a display order of video streams using RPR according to the first embodiment. [Figure 7] FIG. 2 is a diagram illustrating an example of a filter coefficient table 1 according to the first embodiment. [Figure 8] FIG. 4 is a diagram illustrating an example of a filter coefficient table 2 according to the first embodiment. [Figure 9] FIG. 3 is a diagram illustrating an example of a filter coefficient table 3 according to the first embodiment. [Figure 10] 10(A) to 10(C) are diagrams illustrating an example of a thinning method according to the first embodiment. [Figure 11]10A and 10B are diagrams illustrating an example of a method for thinning out one line according to the first embodiment. [Figure 12] 1 is a flowchart showing an example of processing of an encoding method according to the first embodiment. [Figure 13] 10 is a flowchart illustrating an example of an RPR encoding process according to the first embodiment. [Figure 14] 10 is a flowchart showing an example of the RPR encoding process according to the second embodiment. [Figure 15] 11 is a flowchart showing an example of the RPR encoding process according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the present invention is not limited to the following embodiments. In each drawing, the same members or elements are designated by the same reference numerals, and duplicate descriptions will be omitted or simplified.

[0016] (Embodiment 1) Fig. 1 is a functional block diagram showing an example of the configuration of a video distribution device 100 according to a first embodiment of the present invention. Note that some of the functional blocks shown in Fig. 1 are realized by causing a CPU or other computer included in the video distribution device 100 to execute a computer program stored in a memory serving as a storage medium.

[0017] However, some or all of these functions may be implemented by hardware, which may be a dedicated circuit (ASIC) or a processor (reconfigurable processor, DSP).

[0018] 1 may not be contained in the same housing, but may be configured as separate devices connected to each other via signal paths. The above explanation regarding FIG. 1 also applies to FIG. 2.

[0019] A video distribution device 100 functions as an encoding device, and includes a CPU 101 as a computer, a ROM (non-volatile memory) such as an EEPROM or flash memory 102, and a RAM (volatile memory) such as an SRAM or DRAM 103.

[0020] Computer programs for realizing the functions according to this embodiment and data used when executing the computer programs are stored in ROM 102. These programs and data are appropriately loaded into RAM 103 via bus 104 under the control of CPU 101, and executed by CPU 101.

[0021] The communication unit 105 is connected to a network such as Ethernet, and receives image-related setting commands such as instructions for setting image size and frame rate, exposure control for the subject image, and white balance from a video receiving device connected via the network.

[0022] The image setting command is analyzed by the CPU 101 and input to the imaging unit 110, the signal processing unit 120, etc. Image setting information is stored in the ROM 102, and at startup, the CPU 101 sets the imaging unit 110 and the signal processing unit 120 in accordance with the image setting information stored in the ROM 102.

[0023] The imaging unit 110 is composed of an imaging element and an imaging signal output converter, and light rays from a subject form an image on the imaging element via a lens 140. The subject image is photoelectrically converted into an imaging signal by the imaging element, and the imaging signal is converted into a video signal by the imaging signal output converter and output to the signal processing unit 120.

[0024] The signal processing unit 120 performs various digital image processing on the input video signal, such as offset processing, gamma correction processing, gain processing, RGB interpolation processing, noise reduction processing, contour correction processing, color correction processing, image resolution enlargement / reduction processing, etc. The digitally processed image is stored in the RAM 103 via the bus 104 and is compression-encoded by the compression encoding unit 130.

[0025] The communication unit 105 receives setting commands such as the compression code amount of the transmitted video stream and transmission start / stop commands from an external video receiving device. The setting commands for the transmitted video stream are analyzed by the CPU 101 and input to the compression encoding unit 130.

[0026] In addition, transmission video stream setting information is stored in ROM 102, and when video distribution device 100 is started up, CPU 101 performs various settings on compression encoding unit 130 in accordance with the transmission video stream setting information stored in ROM 102.

[0027] The transmission start command is analyzed by the CPU 101, and the compression encoding unit 130 is controlled to read the images stored in the RAM 103, compress them using the VVC encoding method, and store the generated video stream in the RAM 103. The CPU 101 starts transmitting the video stream stored in the RAM 103 via the communication unit 105 to the external video receiving device that has requested the transmission.

[0028] The transmission stop command from the external video receiving device is analyzed by the CPU 101, which stops transmission of the video stream compressed by the compression-encoding unit 130 using the VVC encoding method to the external video receiving device and also stops generation of the video stream by the compression-encoding unit 130. In this embodiment, image compression by the compression-encoding unit 130 is performed based on the VVC standard, but the present invention is not limited to this.

[0029] The communication unit 105 receives control commands such as zoom magnification and focus position from an external video receiving device, and the received commands are analyzed by the CPU 101. The lens 140 includes a variator lens that changes the zoom magnification, a focus lens that changes the focus position, and their actuators.

[0030] The CPU 101 controls the actuator to drive the lens 140 in accordance with control commands for zoom magnification and focus position.

[0031] The communication unit 105 receives PAN and TILT control commands from the video receiving device, and the received commands are analyzed by the CPU 101. The camera platform 150 mechanically supports the imaging unit 110 and lens 140, and has an actuator that rotates them in the PAN and TILT directions.

[0032] The CPU 101 controls the actuator in accordance with the analyzed PAN and TILT control commands to rotate the imaging unit 110 and the lens 140 in the PAN and TILT directions.

[0033] The coding order determination unit 160 calculates motion vectors between multiple images for a GOP (Group of Pictures) length stored in the RAM 103. Then, it outputs coding order information to the compression coding unit 130 so that the sum of the motion vectors of the multiple images of the GOP length is minimized. The coding order information includes picture type information and display order information for each image.

[0034] 2 is a functional block diagram showing an example of the configuration of the compression encoding unit 130 of embodiment 1. Reference numeral 200 denotes an image analysis unit that analyzes the angle of view value of an input frame and outputs the analysis results as image analysis information, and also generates tile information for dividing the image into spatial regions based on image characteristics and external input, and outputs a tile image that combines the image and tile information.

[0035] Reference numeral 210 denotes an RPR control information generator, which generates information on the scaling ratio and offset position of the motion vector required for decoding using RPR.

[0036] A prediction unit 220 performs intra-frame prediction or inter-frame prediction on tile-based image data to generate predicted image data. Furthermore, the prediction unit 220 calculates and outputs a prediction error from the input image data and the predicted image data.

[0037] Additionally, information necessary for prediction, such as prediction mode, motion vector, etc., is also output together with the prediction error from the prediction unit 220. Hereinafter, this information necessary for prediction will be referred to as prediction information.

[0038] A transform / quantization unit 230 performs orthogonal transform on the prediction error in units of blocks to calculate transform coefficients, and then performs quantization to obtain quantized coefficients. A dequantization / inverse transform unit 231 dequantizes the quantized coefficients output from the transform / quantization unit 230 to regenerate transform coefficients, and then performs inverse orthogonal transform on them to regenerate the prediction error.

[0039] Reference numeral 250 denotes a frame memory that stores reconstructed image data. Reference numeral 240 denotes an image reconstruction unit that generates predicted image data by appropriately referencing the frame memory 250 based on the prediction information output from the prediction unit 220, and generates and outputs reconstructed image data from the predicted image data and the input prediction error.

[0040] Reference numeral 251 denotes an in-loop filter unit, which performs in-loop filtering such as deblocking filtering and sample adaptive offset on the reconstructed image acquired from the frame memory 250, and outputs the filtered image.

[0041] An entropy coding unit 260 encodes the quantization coefficients output from the transform / quantization unit 230 and the prediction information output from the prediction unit 220 to generate and output coded data.

[0042] A bitstream generation unit 270 generates header code data by encoding the outputs from the image analysis unit 200 and the RPR control information generation unit 210. Furthermore, the bitstream generation unit 270 combines the header code data with the code data output from the entropy encoding unit 260 to generate and output a bitstream.

[0043] Next, the image encoding operation in the compression encoding unit 130 will be described below. In this embodiment, moving image data is input in frame units. After receiving one frame of image data, the image analysis unit 200 calculates a field of view change value based on the image data. When an arbitrary frame is taken as a reference frame, the field of view change value is the ratio of the field of view of the reference frame to the frame to be encoded.

[0044] Next, the RPR control information generator 210 signals by setting sps_ref_pic_resampling_enabled_flag of the SPS (Sequence Parameter Set) to 1 to indicate that RPR is to be used.

[0045] Also, the number of pixels in the input frame is calculated, and the number of pixels in the vertical direction and the horizontal direction of luminance are stored as pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples of the PPS (Picture Parameter Set), respectively.

[0046] The prediction unit 220 cuts the image data input from the image analysis unit 200 into a plurality of blocks and performs prediction processing on a block-by-block basis. As a result of the prediction processing, prediction errors are generated and input to the transformation and quantization unit 230. The prediction unit 220 also generates prediction information and outputs it to the entropy coding unit 260 and the image reproduction unit 240. Here, the prediction processing performed by the prediction unit 220 and the prediction information output from the prediction unit 220 will be described in more detail.

[0047] In image coding technologies such as VVC, a prediction process is used to predict pixels of a block to be coded using pixels of a previously coded block in order to reduce the amount of data in the coded bitstream while maintaining the image quality of the reproduced image.

[0048] Prediction processes include intra-prediction, which uses pixels from blocks already coded in the same frame, and inter-prediction, which uses pixels from blocks in different coded frames. VVC also standardizes a technology called RPR, which allows decoding even when the resolution of the coded frame to be referenced and the frame to be coded are different.

[0049] Here, as an explanation of RPR, inter prediction in the case where the resolution of a reference encoded frame and the resolution of a current encoding frame are different will be further explained.

[0050] 3A and 3B are diagrams illustrating an example of coding order information generated by the coding order determination unit 160 according to the first embodiment. It is assumed that frames 301 to 306 as multiple image frames are arranged in chronological order.

[0051] 3B is a diagram showing an example of the configuration of coding order information 307. In the example shown in FIG. 3B, coding order information 307 includes the coding order, display order index, and picture type of each of frames 301 to 306.

[0052] For example, the picture type of frame 301 is an I-frame (I-picture (Intra-Picture)) that is composed only of intra-references, while frames 302 and 303 are B-frames (B-pictures (Bidirectional-Picture)) that include references to two frames.

[0053] Frames 304 to 306 are P frames (P pictures (Predictive-Pictures)) that include a reference to the previous frame.

[0054] 4 is a diagram showing an example of rearrangement of the video stream of the first embodiment in the encoding order, in which the compression encoding unit 130 rearranges frames 301 to 306 in the encoding order. The arrows are used to illustrate examples of reference frames.

[0055] 5 is a diagram for explaining an example of resolution conversion of a video stream rearranged in the encoding order of embodiment 1, showing an example in which the resolution is changed midway through the stream and RPR is used. When switching from frame 304 to frame 302, the image size is reduced to 2 / 3 both vertically and horizontally.

[0056] In terms of specific pixel count, for example, frames 301 and 304 have a pixel count of 2880 x 1620. Other frames have a pixel count of, for example, 1920 x 1080. Where a change in resolution occurs, the reference encoded frame is scaled (enlarged or reduced) to match the resolution of the frame to be encoded before inter prediction is performed.

[0057] Fig. 6 is a diagram showing an example of the display order of video streams using RPR according to the first embodiment, and shows the order in which streams compressed using the RPR of Fig. 5 are displayed. When switching from frame 301 to frame 302, the resolution is reduced to 2 / 3 both vertically and horizontally, and when switching from frame 303 to frame 304, the resolution is increased to 3 / 2 both vertically and horizontally.

[0058] Furthermore, when switching from frame 304 to frame 305, the resolution is reduced to 2 / 3 both vertically and horizontally, resulting in a stream in which frame 304 with a higher resolution is inserted into the reduced stream from frame 302.

[0059] An example of a method for scaling a reference encoded frame is shown below. For simplicity, only luminance values ​​will be explained. Color differences can be similarly scaled taking the number of samples into account, so their explanation will be omitted.

[0060] (Step 1) First, find the scaling ratios in the vertical and horizontal directions, scalingRatio[0] and scalingRatio[1], respectively. Hereinafter, scalingRatio[0] and scalingRatio[1] will be collectively referred to as scalingRatio[x].

[0061] scalingRatio[x] is determined by the ratio between the size of a scaling window (to be described later) of a reference encoded frame and the scaling window of the encoding target frame. In this embodiment, scalingRatio[x] is obtained as RPR control information from the RPR control information generator.

[0062] Scaling window is a technology standardized by VVC, which improves coding efficiency for video that involves changes in the angle of view, such as when zooming, by synchronizing the scaling window with the change in the angle of view.

[0063] In this embodiment, the scaling window is not explicitly set, so scalingRatio[x] is calculated by dividing the number of pixel samples of the video output by the decoder of the reference encoded frame by the number of pixel samples of the video output by the decoder of the current encoding frame.

[0064] (Step 2) Determines the interpolation filter used for scaling. For example, selects the coefficients of the interpolation filter depending on the value of scalingRatio[x].

[0065] FIG. 7 is a diagram showing an example of filter coefficient table 1 of the first embodiment. When scalingRatio[x] is equal to or greater than 1.75, for example, the coefficients in table 1 of FIG. 7 are used.

[0066] FIG. 8 is a diagram showing an example of filter coefficient table 2 in the first embodiment. When scalingRatio[x] is less than 1.75 times and is equal to or greater than 1.25 times, for example, the coefficients in table 2 in FIG. 8 are used.

[0067] 9 is a diagram showing an example of filter coefficient table 3 in embodiment 1, and when scalingRatio[x] is less than 1.25, for example, the coefficients in table 3 in Fig. 9 are used. In each table, the coefficients of the interpolation filter are determined by the sample position p to be calculated.

[0068] Tables 1 to 3 in Figures 7 to 9 are filters defined for brightness expansion / reduction in the VVC standard. Table 1 is created based on the cutoff frequency when the ratio is 2x, and Table 2 is created based on the cutoff frequency when the ratio is 1.5x.

[0069] p is an integer ranging from 0 to 15, and is the numerator value when the smallest sample unit is divided into 1 / 16th units. For example, if a certain sample point A and a certain sample point A+1 are divided into 16, the filter coefficients for the third sample point (A+3 / 16, p=3) are fL[3][i]=[-4, -1, 16, 29, 23, 7, -4, -2], referring to Table 1.

[0070] Hereinafter, when the filter coefficients obtained in this way are used in the horizontal direction, they will be written as fLH[p][i] (=fL[p][i]), and when they are used in the vertical direction, they will be written as fLV[p][i] (=fL[p][i]).

[0071] (Step 3) The reference image generated from the encoding frame to be referenced is resampled according to scalingRatio[x] so that it has the same resolution as the encoding frame. For example, the position of each pixel when the encoding frame is scaled by scalingRatio[x] is calculated with 1 / 16 pixel accuracy.

[0072] Also, the reference image is interpolated 16 times using the filter obtained in step 2. For example, if scalingRatio[x]≧1.25, a new sampling point (x3+px, y3+py) is obtained using the following equations 1 and 2.

[0073] Here, we use the notation that the coordinates of the target pixel in the reference image are (xi, yi), the coordinates of the adjacent pixel to the left are (x(i-1), yi), and the coordinates of the adjacent pixel below are (xi, y(i-1)).

[0074] Also, px and py are integers modulo 16 (divisor), and are indices of coordinates obtained by dividing the coordinates of adjacent pixels in the horizontal and vertical directions by 16, respectively, and L(x, y) represents the luminance value of coordinates (x, y). a is a normalization constant.

[0075]

number

[0076]

number

[0077] Here, fLH[p][i] and fLV[p][i] are generated from Table 2. Furthermore, yn is a coordinate value from y0 to y7. The samples generated in this manner are called upsampled images. This upsampled image is thinned out, and a reference image is generated after resampling to the same resolution as the frame to be coded from the reference image.

[0078] 10(A) to 10(C) are diagrams for explaining an example of a thinning method in embodiment 1, and Fig. 11(A) and 11(B) are diagrams for explaining an example of a thinning method for one line in embodiment 1. The thinning method is the same in both the vertical and horizontal directions, so for simplicity, only the horizontal direction will be explained.

[0079] Let 701 be a reference image, 702 be a frame to be coded, and they be composed of unit areas such as 703. It is assumed that one pixel value is defined for one unit area. In this case, the reference image 701 is an image of 6×6 pixels made up of unit areas, and the frame to be coded 702 is an image of 4×4 pixels made up of unit areas.

[0080] The origin is the upper left vertex of the entire image, each unit area is 1x1 in size, and the coordinates of each pixel value are the coordinates of the upper left vertex of the unit area. Here, the coordinate values ​​of each pixel in the encoding frame are multiplied by scalingRatio[x] (2 / 3 in the example of Figure 10).

[0081] Then, for the x coordinate of the encoding target frame 702 being (0, 1, 2, 3), the coordinate values ​​(0, 3 / 2, 6 / 2, 9 / 2) can be calculated and stored in, for example, a coordinate array H[x].

[0082] 10C, 704 is an image obtained by resampling the reference image to the same number of pixels as the frame to be encoded. The pixel values ​​of each unit area of ​​the resampled image 704 may be constructed using pixel values ​​from the upsampled image that correspond to the coordinates of the enlarged image.

[0083] An example of the construction method will now be described with reference to Fig. 11. 801 depicts only the unit area of ​​y = 0 of the image of 704. Similarly, 804 depicts only the unit area of ​​y = 0 of the reference image 701.

[0084] Here, the luminance value of the reference image 801 is Y[x], and the luminance value of the resampled image is Y'[x]. Note that x is the x-coordinate value of the luminance value. Then, to obtain the luminance value of coordinate position x of the resampled image, H[x] in the coordinate array 803 is referenced, the coordinate position of the reference image is found, and the luminance value of that coordinate position is obtained. When limited to the explanation of the x-coordinate only, it can be written as in the following Equation 3.

[0085]

number

[0086] For example, the coordinate of the brightness value to be found in unit area 802 is 1, so x=1 is set, the coordinate array H[1]=3 / 2 is found, and the value of Y[3 / 2] is obtained. 805 is a diagram showing the unit area of ​​coordinate value 1 in 804 interpolated from 16 areas using the filter.

[0087] Each of these 16 regions is 1 / 16 the size of the unit region, and the pixel values ​​of each region are determined by the filter. That is, Y[3 / 2] is the brightness value of the coordinate corresponding to 3 / 2 of 806 (=1 + 8 / 16).

[0088] In this way, after resampling at the same resolution, the reference image can be constructed by upsampling the reference image and thinning out pixel values ​​other than those corresponding to the enlarged pixel positions in the frame to be coded.

[0089] Inter-prediction is a process of predicting the pixels of a block to be coded by referring to the pixels of an already coded frame, or, if the number of pixels of the already coded frame differs from the number of pixels of the frame to be coded, by referring to the pixels of a resampled image constructed using the above method.

[0090] For simplicity, the coded frame and the resampled image are collectively referred to as a target image for inter prediction. For example, if there is no motion between the coded frame to be referenced and the target image for inter prediction, pixels of the target block for coding are predicted using pixels at the same positions in the target image for inter prediction.

[0091] In such a case, the motion vector (0, 0), which indicates no motion, is included in the prediction information. On the other hand, if there is motion between frames for the block to be coded, the motion vector (MVx, MVy) is included in the prediction information.

[0092] 12 is a flowchart showing a processing example of the encoding method of the first embodiment, and shows the control flow for GOP length frames in the compression encoding unit 130. Note that the CPU 101 or the like as a computer executes a computer program stored in memory, thereby sequentially performing the operations of the steps in the flowchart of FIG.

[0093] Step S100 is a step in which the coding order determination unit 160 acquires the coding order information 307 for the number of frames of the GOP length determined.

[0094] The next step S101 is a step of measuring the amount of buffering of the stream when transmitting from the communication unit 105 to the network for each unit time of the GOP length in order to estimate the communication bandwidth of the network.

[0095] In the next step S102, it is determined whether the network bandwidth is narrower than a predetermined value. That is, if the buffer amount measured in step S101 is increasing per unit time of the GOP length, it is determined that the network bandwidth is narrower than the predetermined value (Yes in step S102), and the process proceeds to step S103. Note that step S103 is an encoding step in which encoding is performed using RPR.

[0096] If the buffer amount measured in step S101 does not decrease or increase per unit time of the GOP length, it is determined that the network bandwidth is wider than a predetermined value (No in step S102), and the process proceeds to step S104. Step S104 is an encoding step that does not use RPR.

[0097] That is, when the network bandwidth is wider than a predetermined value, the difference frame is encoded by referring to image frames with different resolutions without converting the resolution. Here, steps S103 and S104 function as an encoding step (encoding means) in this embodiment.

[0098] Step S103 is a step in which RPR is used during encoding due to insufficient network bandwidth. Fig. 13 is a flowchart showing an example of the RPR encoding process of embodiment 1, and shows an example of a detailed flow of the encoding method using RPR in step S103. Note that the operation of each step in the flowchart of Fig. 13 is performed sequentially by a computer such as CPU 101 executing a computer program stored in memory.

[0099] The encoding without using RPR in step S104 in FIG. 12 is a simplified version of FIG. 13, so a description thereof will be omitted.

[0100] The detailed control of step S103 of encoding using RPR will be explained below with reference to Fig. 13. In step S200, a parameter count for counting frames up to the GOP length is set to 1, and a parameter last for counting the display order index of the encoding order information 307 is set to 1.

[0101] In the next step S201, the first frame in the coding order of the coding order information 307 is coded as an I frame. In the next step S202, it is determined whether count, which is a parameter for counting the number of frames, has reached the GOP length. If it is less than the GOP length, the process proceeds to step S203, and if it reaches the GOP length, the process flow in FIG. 13 ends.

[0102] In step S203, the display order index of the next frame is obtained in the coding order of the coding order information 307. Note that the display order index of the frame obtained at this time does not necessarily match the coding order.

[0103] In the following step S204, the total amount of code of the compressed stream that has been coded so far is calculated (accumulated), and in step S205, it is determined whether the total value of the amount of code (i.e., the accumulated amount of code of the coded image data) is greater than a predetermined threshold.

[0104] This predetermined threshold is set to a code amount that is a proportion of the maximum communication buffer amount of the communication unit 105, for example, 50%. Alternatively, if the effective bandwidth of the network can be estimated in step S102 of Fig. 12, it is set to, for example, 50% of the estimated effective bandwidth. If the total value of the code amount is greater than the predetermined threshold, the process proceeds to step S206, and if the total value of the code amount is equal to or less than the predetermined threshold (i.e., the cumulative amount of the code amount of the encoded image data is equal to or less than the predetermined threshold), the process proceeds to step S209.

[0105] In step S206, it is determined whether the coding order and the display order are interchanged. That is, it is determined whether the parameter "last", which counts the display order index, is smaller than the display order index. If the determination in step S206 is "Yes", it is determined that the coding order is the same as the display order, and the process proceeds to step S207. If the determination in step S206 is "No", the process proceeds to step S209.

[0106] In step S207, the parameter "last" for counting the display order index is substituted with the display order index, and updated. In the following step S208, resolution conversion is performed.

[0107] That is, the resolution of the image stored in RAM 103 is reduced using signal processing unit 120 in Fig. 1. In this way, in step S208, if the resolution of the predetermined image frame is different from the resolution of the reference image frame, the resolution of the reference image frame is converted to be the same as the resolution of the predetermined image frame.

[0108] In step S209, the compression encoding unit 130 encodes the difference frame from the image stored in the RAM 103. The picture type at this time follows the picture type in the encoding order information 307.

[0109] In the case of an image whose resolution has been converted in step S208, encoding is performed using RPR. Also, if the cumulative amount of code of the encoded image data is equal to or less than a predetermined threshold, image frames with different resolutions are referenced and differential frame encoding is performed without converting the resolution.

[0110] In the following step S210, count, which is a parameter for counting the number of frames, is incremented by 1, and the process returns to step S202.

[0111] In this way, in the first embodiment, when sequentially encoding a plurality of image frames, if the encoding order and display order of the image frames are the same, image frames with different resolutions are referenced and subjected to resolution conversion before being encoded as difference frames. On the other hand, if the encoding order and display order are different, image frames with different resolutions are referenced and encoded as difference frames without being subjected to resolution conversion.

[0112] Therefore, it is possible to realize an encoding device that can reduce screen disturbances that occur when switching resolutions and that can deliver stable streams.

[0113] (Embodiment 2) Fig. 14 is a flowchart showing an example of the RPR encoding process according to the second embodiment, and shows an example of a detailed flow of step S103 of encoding as an encoding method using RPR according to the second embodiment. Note that the operation of each step in the flowchart in Fig. 14 is performed sequentially by a computer such as a CPU 101 executing a computer program stored in a memory.

[0114] In step S300, count, which is a parameter for counting frames up to the GOP length, is set to 1. In the following step S301, the first frame in the coding order of the coding order information 307 is coded as an I frame.

[0115] In step S302, it is determined whether or not the count corresponding to the number of frames has reached the GOP length. If it has not reached the GOP length, the process proceeds to step S303, and if it has reached the GOP length, the process flow in FIG. 14 ends.

[0116] In step S303, the picture type of the next frame in the coding order information 307 is obtained, and in step S304, the total amount of code of the compressed stream that has been coded so far is calculated (accumulated).

[0117] In step S305, it is determined whether the total amount of code is greater than a predetermined threshold. This predetermined threshold is, for example, 50% of the maximum amount of code in the communication buffer of the communication unit 105. Alternatively, if the effective bandwidth of the network can be estimated in step S102 of Fig. 12, it is set to, for example, 50% of the effective bandwidth. If the total amount of code is greater than the predetermined threshold, the process proceeds to step S306, and if the total amount of code is equal to or less than the predetermined threshold, the process proceeds to step S308.

[0118] In step S306, the picture type of the coding order information 307 is determined. If it is a P frame, it is determined that only forward referencing is used and the coding order and display order are not reversed, and the process proceeds to step S307. If it is not a P frame (if it is a B frame), the process proceeds to step S308.

[0119] In step S307, the resolution of the image stored in the RAM 103 is reduced using the signal processing unit 120 in FIG. 7, and in step S308, the image stored in the RAM 103 is encoded into a difference frame by the compression encoding unit 130. The picture type at this time follows the picture type in the coding order information 307. If the image has undergone resolution conversion in step S307, the coding in step S308 uses RPR. In the following step S309, the count corresponding to the number of frames is incremented by 1, and the process returns to step S302.

[0120] In this way, in the second embodiment, when sequentially encoding a plurality of image frames, if the picture type of the image frames is a P frame, image frames with different resolutions are converted in resolution before being referenced for difference frame encoding, whereas if the picture type of the image frames is a B frame, image frames with different resolutions are referenced in resolution without being converted in resolution before being referenced for difference frame encoding.

[0121] (Embodiment 3) Fig. 15 is a flowchart showing an example of the RPR encoding process according to the third embodiment, and shows a detailed flow of step S103 of encoding, which is an encoding method using RPR, according to the third embodiment. Note that the operation of each step in the flowchart of Fig. 15 is performed sequentially by a computer such as a CPU 101 executing a computer program stored in a memory.

[0122] In step S400, the parameter count for counting frames up to the GOP length is set to 1, and the parameter last for counting the display order index of the coding order information 307 is set to 1.

[0123] In the next step S401, the first frame in the coding order of the coding order information 307 is coded as an I frame. In the next step S402, it is determined whether count, which is a parameter for counting the number of frames, has reached the GOP length. If it is less than the GOP length, the process proceeds to step S403, and if it reaches the GOP length, the process flow in FIG. 15 ends.

[0124] In step S403, the display order index of the next frame in the coding order of the coding order information 307 is obtained. Note that the display order index of the frame obtained at this time does not necessarily match the coding order.

[0125] In the following step S404, the total amount of code of the compressed stream that has been coded so far is calculated (accumulated), and in step S405, it is determined whether or not the total amount of code is greater than a predetermined threshold 1 (first threshold).

[0126] This predetermined threshold 1 (first threshold) is set to a code amount that is a proportion of the maximum communication buffer capacity of the communication unit 105, for example, 50%. Alternatively, if the effective bandwidth of the network can be estimated in step S102 of Fig. 12, it is set to, for example, 50% of the estimated effective bandwidth. If the total value of the code amount is greater than threshold 1, the process proceeds to step S406, and if the total value of the code amount is equal to or less than threshold 1, the process proceeds to step S413.

[0127] In step S406, it is determined whether the total value of the code amount is greater than a predetermined threshold 2 (second threshold). Threshold 2 is set to a ratio of the code amount to the maximum communication buffer amount of the communication unit 105, for example, 80%. Alternatively, if the effective bandwidth of the network can be estimated in step S102 of Fig. 12, threshold 2 is set to, for example, 80% of the effective bandwidth. If the total value of the code amount is equal to or less than threshold 2, the process proceeds to step S407, and if the total value of the code amount is greater than threshold 2, the process proceeds to step S410.

[0128] In step S407, it is determined whether the coding order and the display order are interchanged. That is, it is determined whether last, which is a parameter for counting the display order index, is smaller than the display order index. If the determination in step S407 is Yes, it is determined that the coding order is the same as the display order, and the process proceeds to step S408. If the determination in step S407 is No, the process proceeds to step S413.

[0129] In step S408, the display order index is substituted into last, which is a parameter for counting the display order index, and updated. In step S409, resolution conversion is performed. That is, the resolution of the image stored in RAM 103 is reduced using signal processing unit 120 in FIG. 1.

[0130] On the other hand, if the determination in step S406 is Yes, the process proceeds to step S410, where it is determined whether or not the parameter "last" for counting the display order index is smaller than the display order index. If the determination in step S410 is Yes, the process proceeds to step S411 to update last, which is a parameter for counting the display order index. If the determination in step S410 is No, the process proceeds to step S412.

[0131] In step S411, as in step S408, the display order index is substituted into last, which is a parameter for counting the display order index, to update it. In step S412, resolution conversion is performed.

[0132] That is, the resolution of the image stored in RAM 103 is reduced using signal processing unit 120 in Fig. 1. In this manner, in this embodiment, when the total value of the code amount exceeds threshold 2, resolution conversion is performed on the image stored in RAM 103 regardless of the display order index.

[0133] That is, when the cumulative amount of code of the encoded image data exceeds a second threshold that is higher than a predetermined threshold (first threshold), even if the encoding order is different from the display order, the image frames with different resolutions are resolution converted and then referenced to encode the difference frame.

[0134] In step S413, the image stored in the RAM 103 is subjected to difference frame encoding by the compression encoding unit 130. The picture type at this time conforms to the picture type of the encoding order information 307. If the image has undergone resolution conversion in step S409 or step S412, it is encoded using RPR in step S413. In step S414, the count corresponding to the number of frames is incremented by 1, and the process returns to step S402.

[0135] As a result, when a frame with a higher resolution is inserted in the middle of a reduced stream as shown in FIG. 6, compression encoding using RPR can be achieved while suppressing resolution conversion.

[0136] In the explanations of the first to third embodiments, the GOP length is used as a time unit for determining the encoding order by re-accumulating the accumulated amount of encoded data based on the GOP length. However, the encoding order may be determined in units of the reference range of B frames, so that resolution conversion is performed while determining the encoding order.

[0137] Although the GOP length is used as the cycle for determining the communication bandwidth, the code amount threshold may be updated in a cycle shorter than the GOP length, for example, every 100 msec. In this case, the RPR of the original resolution relative to the reduced resolution can be used in addition to the RPR of the reduced resolution relative to the original resolution.

[0138] As described above, according to the first to third embodiments, it is possible to improve the stability of video stream quality in a system that employs the RPR of the VVC encoding method and its successor standards.

[0139] The present invention has been described above in detail based on its preferred embodiments, but the present invention is not limited to the above embodiments, and various modifications and combinations of the above embodiments are possible based on the spirit of the present invention, and these are not excluded from the scope of the present invention.

[0140] The present invention also includes those that realize the functions of the above embodiments using, for example, at least one processor such as a CPU, memory, or circuit (for example, ASIC). Also, multiple processors may be used to perform distributed processing.

[0141] In order to realize part or all of the control in the above-described embodiments, a computer program that realizes the functions of the above-described embodiments may be supplied to an encoding processing device or the like via a network or various storage media. Then, a computer (or a CPU, MPU, or the like) in the encoding processing device or the like may read and execute the program. In this case, the program and the storage medium storing the program constitute the present invention. The present invention also includes the following combinations.

[0142] (Configuration 1) A coding device characterized by having a coding means for, when sequentially coding a plurality of image frames, if the coding order and display order of the image frames are the same, converting the resolution of the image frames with different resolutions before referring to them to code the difference frames, and if the coding order and display order are different, referencing the image frames with different resolutions without converting the resolution and coding the difference frames.

[0143] (Configuration 2) A coding device characterized by having a coding means for, when sequentially coding a plurality of image frames, if the picture type of the image frames is a P frame, converting the resolution of the image frames with different resolutions before referring to them to code the difference frame, and if the picture type of the image frames is a B frame, referencing the image frames with different resolutions without converting the resolution and coding the difference frame.

[0144] (Configuration 3) The encoding device according to configuration 1 or 2, wherein the encoding means performs encoding using RPR (Reference Picture Resampling).

[0145] (Configuration 4) The encoding means The encoding device according to any one of configurations 1 to 3, characterized in that, if the resolution of the predetermined image frame is different from the resolution of the reference image frame, the resolution of the reference image frame is converted to be the same as the resolution of the predetermined image frame.

[0146] (Configuration 5) The encoding device according to any one of configurations 1 to 4, characterized in that when the cumulative amount of code of the encoded image data is equal to or less than a predetermined threshold, the encoding means performs differential frame encoding by referring to the image frame with a different resolution without converting the resolution.

[0147] (Configuration 6) The encoding device according to claim 5, characterized in that when the accumulated amount exceeds a second threshold higher than the predetermined threshold, the encoding means converts the resolution of the image frames with different resolutions and then refers to them to encode the difference frame, even if the encoding order is different from the display order.

[0148] (Configuration 7) The encoding device according to configuration 5 or 6, wherein the encoding means re-accumulates the accumulated amount of encoded data based on a GOP length.

[0149] (Configuration 8) The encoding device described in any one of configurations 1 to 7, characterized in that when the network bandwidth is wider than a predetermined value, the encoding means encodes the difference frame by referring to the image frames with different resolutions without converting the resolution.

[0150] (Method 1) A coding method characterized by the fact that, when sequentially encoding a plurality of image frames, if the encoding order and display order of the image frames are the same, the image frames with different resolutions are subjected to resolution conversion before being referenced and differential frame encoding, and if the encoding order and display order are different, the image frames with different resolutions are referenced and differential frame encoding without resolution conversion.

[0151] (Method 2) A coding method characterized by the fact that, when sequentially encoding a plurality of image frames, if the picture type of the image frames is a P frame, the image frames with different resolutions are subjected to resolution conversion before being referenced and differential frame encoding is performed, and if the picture type of the image frames is a B frame, the image frames with different resolutions are referenced and differential frame encoding is performed without resolution conversion.

[0152] (Program) A computer program for controlling the encoding means of the encoding device according to any one of configurations 1 to 8 by a computer. [Explanation of symbols]

[0153] 100:Video distribution device 110: Imaging unit 120: Signal processing unit 130: Lens 140: Compression coding unit 150: Pan head 160: Encoding order determination unit 210: RPR control information generation unit

Claims

1. An encoding device characterized by having an encoding means for, when sequentially encoding a plurality of image frames, if the encoding order and display order of the image frames are the same, converting the resolution of the image frames having different resolutions before referring to them to encode the difference frame, and if the encoding order and display order are different, referencing the image frames having different resolutions without converting the resolution and encoding the difference frame.

2. a coding device comprising: coding means for, when sequentially coding a plurality of image frames, converting the resolution of the image frames having different resolutions before referring to the image frames to code the difference frames if the picture type of the image frames is a P frame; and, when the picture type of the image frames is a B frame, converting the resolution of the image frames having different resolutions before referring to the image frames to code the difference frames.

3. 2. The encoding device according to claim 1, wherein the encoding means performs encoding using RPR (Reference Picture Resampling).

4. The encoding means 2. The encoding device according to claim 1, wherein, if the resolution of the predetermined image frame is different from the resolution of the reference image frame, the resolution of the reference image frame is converted to be the same as the resolution of the predetermined image frame.

5. The encoding device according to claim 1, characterized in that, when the cumulative amount of code of the encoded image data is equal to or less than a predetermined threshold, the encoding means performs differential frame encoding by referring to the image frame having a different resolution without converting the resolution.

6. The encoding device according to claim 5, characterized in that, when the accumulated amount exceeds a second threshold value that is higher than the predetermined threshold value, the encoding means converts the resolution of the image frames having different resolutions and then refers to the image frames to encode the difference frame, even if the encoding order is different from the display order.

7. 6. The encoding device according to claim 5, wherein said encoding means re-accumulates said integrated amount of encoded data based on a GOP length.

8. 2. The encoding device according to claim 1, wherein, when a network bandwidth is greater than a predetermined value, the encoding means performs differential frame encoding by referring to the image frames having different resolutions without converting the resolutions.

9. An encoding method characterized by the fact that, when sequentially encoding a plurality of image frames, if the encoding order and display order of the image frames are the same, the image frames with different resolutions are subjected to resolution conversion before being referenced and differential frame encoding, and if the encoding order and display order are different, the image frames with different resolutions are referenced and differential frame encoding without resolution conversion.

10. An encoding method characterized by the fact that, when sequentially encoding a plurality of image frames, if the picture type of the image frames is a P frame, the image frames with different resolutions are subjected to resolution conversion before being referred to for difference frame encoding, and if the picture type of the image frames is a B frame, the image frames with different resolutions are subjected to resolution conversion before being referred to for difference frame encoding.

11. A computer program for controlling the encoding means of the encoding device according to any one of claims 1 to 8 by a computer.

Citation Information

Patent Citations

  • Video stream supply system and apparatus, and video stream receiving apparatus

    JP2007088539A

  • Image encoder

    JP2013214894A