Method and encoder for inter-encoding image frame
The two-pass encoding method addresses artifacts in high-compression regions by selectively reducing the compression of intra-coded blocks, enhancing perceived quality and reducing artifacts over time.
Patent Information
- Application Number
- JP2024194891
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-11-07
- Publication Date
- 2025-06-17
AI Technical Summary
Artifacts appear in regions of image frames with high compression levels due to the difference in perceived quality between intra-coded and inter-coded pixel blocks, which worsen over time.
A two-pass encoding method that identifies and selectively reduces the compression of intra-coded pixel blocks in highly compressed regions, increasing the likelihood of inter-coding these blocks and reducing artifacts.
The method improves the perceived quality in high-compression regions by reducing artifacts and ensuring that intra-coded blocks are coded with higher quality, even in subsequent frames.
Smart Images

Figure 2025090522000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video coding. In particular, the present invention relates to a method and an encoder for inter-coding image frames in a sequence of image frames.
Background Art
[0002] When encoding video image frames, it is common to apply spatially varying compression levels to each image frame in order to reduce the bitrate while maintaining the perceived quality levels in different regions of the image frame. Lower compression levels can be used in regions of the image frame that are of more interest to the viewer, while higher compression levels can be used for regions that are of less interest to the viewer. This can be achieved by using more bits to encode regions of interest in the image and fewer bits for regions of less interest. The determination of what constitutes the region of interest (ROI) can vary depending on the application example. For example, the level of motion or the level of image detail can be used to determine whether a particular region of the image should be considered an ROI.
[0003] When using methods such as those described above, the inventors have noticed problems related to artifacts that appear in regions of image frames having a high compression level under certain circumstances. These types of artifacts are characterized by a significant difference in the perceived quality between different areas of the affected region. These artifacts tend to worsen over time and increase until the next I-frame is encoded. Examples of cases where these problems have been observed are in the background regions of moving objects in an image, such as the rear lane of a car, a slight change in light across a wall, or an aurora borealis reflected in a lake.
[0004] U.S. Patent No. 10,425,642 discloses a method for improving picture quality when there is a large variance in quantization parameters applied to residual coefficients of coding units across an image frame. To reduce the variance, an average quantization parameter is calculated for coding units of an image frame in a first encoding pass and then used to determine an updated quantization parameter to be applied to the coding units in a second encoding pass.
[0005] U.S. Patent Application Publication No. 2015 / 0373328 discloses a variable bitrate system in which an encoder varies quantization parameters for each frame using a two-pass encoding scheme. A first-pass analysis of the entire frame sequence determines which frames are more complex, and a second-pass analysis changes the quantization parameters of the frames for more efficient coding in light of the first-pass analysis.
[0006] European Patent No. 2,132,938 discloses the use of a two-pass encoding scheme to meet bitrate constraints or utilize unused bandwidth. Encoding in the first coding pass changes the video coding mode of a video block, such as from skip mode to direct mode, and adjusts the quantization parameter, thereby being improved in the second coding pass.
[0007] U.S. Patent Application Publication No. 2020 / 0260083 discloses how the variation or adaptation of coding quantization parameters across a picture can be made more effective by basing the determination of the coding quantization parameter for a given block of a picture on the variance of the statistical sample value distribution of a high-pass filtered version of the given block.
[0008] U.S. Patent Application Publication No. 20220312021 discloses an analysis-modulated video compression approach that enables a coding process to dynamically adapt quantization based on the content of surveillance images. A two-pass encoding method can be applied, in which case, the first pass is used to derive quantization parameter values associated with foreground objects based on, for example, a target bit rate and the number and size of objects within a scene.
[0009] Therefore, there is a need for improvement in this regard. SUMMARY OF THE INVENTION
[0010] Therefore, in view of the above, an object of the present invention is to overcome or mitigate the above problems by providing an encoding method that improves the perceptual quality in high-compression regions of image frames.
[0011] The above object is achieved by the present invention as defined by the appended independent claims. Advantageous embodiments are defined by the appended dependent claims.
[0012] The inventors recognized that artifacts appearing in highly compressed regions of an image frame are caused by some pixel blocks being intra-coded while other adjacent pixel blocks are inter-coded. Pixel blocks that are inter-coded are temporally predicted from information in previously encoded frames within the video. Thanks to the temporal prediction, the remaining residuals to be encoded are typically small, which enables the pixel blocks to be encoded and perceived as being of high quality despite the high compression level. This is in contrast to intra-coding blocks that are spatially predicted only from information within the current frame. In that case, the residuals to be encoded are usually much larger, resulting in a relatively low perceived quality for the same high compression level. As a result, inter-coded blocks are perceived as having higher quality than intra-coded blocks. This quality difference causes artifacts to appear in the decoded video and is particularly prominent around the boundaries between intra-coded blocks and adjacent inter-coded blocks. The application of a deblocking filter in the decoder tends to further accentuate the artifacts by introducing ringing artifacts at these boundaries. Inter-coded blocks in future image frames of the video also reference these intra-coded blocks, worsening the problem over time.
[0013] To overcome or mitigate the above problems, a two-pass encoding method is proposed that identifies and selectively reduces the compression of intra-coded pixel blocks located in highly compressed regions of an image frame. The two-pass encoding method of the present invention uses the obtained compression level applied to each pixel block in the frame to perform a first encoding pass to identify intra-coded pixel blocks in an image frame that may cause the described artifacts, i.e., intra-coded pixel blocks whose compression level exceeds a compression level threshold. The compression level of the blocks identified as potentially problematic in the first encoding pass is reduced before performing the second encoding pass. Reducing the compression level of these identified blocks typically has the effect of increasing the probability that the encoder will choose to inter-code the blocks identified in the second encoding pass. Thus, since artifacts are caused by some intra-coded pixel blocks and other inter-coded pixel blocks, the probability of artifacts appearing after the second encoding pass is reduced. Further, if the identified blocks are ultimately also intra-coded in the second encoding pass, the compression level is still reduced, the perceived quality is improved, and the impact of any potential artifacts still present after the second encoding pass is reduced.
[0014] The term "pixel block" should be understood to mean a set of adjacent pixels grouped together. These pixel blocks form the units within the image frame in which the encoder operates when encoding the image frame. These pixel blocks may also be referred to as macroblocks, coding tree units, or coding units, depending on the coding standard used to encode the image. Pixel blocks may most often be squares consisting of, for example, 8×8, 16×16, or 32×32 pixels. It is also possible to group pixels into pixel blocks of other sizes and shapes.
[0015] The term "compression level of a pixel block" refers to the degree or level to which the image data within a pixel block is compressed during the encoding of the pixel block. When compression is achieved by quantization, as is often the case with transform-based codecs, the compression level can correspond to the level of quantization applied when encoding the pixel block. Such a quantization level may also be expressed as a quantization value, quantization parameter, quantization index, or step size, depending on the encoding standard used to encode the image frame in the first and second encoding passes.
[0016] The expressions "inter-encode an image frame in a first encoding pass" and "inter-encode an image in a second encoding pass" refer to the image frame being inter-coded twice, i.e., in two rounds or phases. Thus, the first encoding pass refers to the first time the image frame is inter-coded, and the second encoding pass refers to the second time the image frame is inter-coded. Thus, this method may also be referred to herein as a two-pass encoding method or scheme. The purpose of the first encoding pass is to identify and reduce the compression level of pixel blocks in the encoded image frame that may cause artifacts. The image frame is then encoded a second time in the second encoding pass using the reduced compression level for the identified pixel blocks to output an encoded video with reduced artifacts.
[0017] The present invention consists of three aspects: a method, an encoder, and a computer-readable storage medium. The second and third aspects may generally have the same features and advantages as the first aspect. It should be further noted that the present invention relates to all combinations of features unless explicitly stated otherwise.
[0018] The above and additional objects, features, and advantages of the present invention will be better understood from the following illustrative and non-limiting detailed description of embodiments of the present invention with reference to the accompanying drawings, in which like reference numerals are used for like elements.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4a
Figure 4b
Figure 4c
Figure 4d
Figure 5
Figure 6
Modes for Carrying Out the Invention
[0020] Hereinafter, the present invention will be more fully described with reference to the accompanying drawings showing embodiments of the present invention.
[0021] Figure 1 shows a sequence of image frames 100 that includes intra-coded image frames 102-1 and 102-2 and inter-coded image frames 104-1 to 104-6. Intra-coded image frames are encoded independently of all other image frames within the sequence of image frames, and utilize the spatial redundancy of the pixels within the image frame itself without referring to other image frames. These types of intra-coded image frames are generally referred to as intra-frames, I-frames, or key frames. Inter-coded frames utilize the temporal redundancy of the pixels between image frames and are encoded to be dependent on other frames within the sequence of image frames. These types of inter-coded frames are generally referred to as inter-frames or delta frames and can be forward prediction frames, P-frames, or bidirectional prediction frames, B-frames.
[0022] The present invention includes a method for inter-encoding image frames within a sequence of image frames, such as one or more of image frames 104-1 to 104-6 within sequence 100, using a two-pass encoding scheme, the steps of which are shown in Figure 3 and described in further detail below. The method can be executed each time an image frame within sequence 100 of image frames is inter-encoded. Alternatively, it may be executed for the selection of image frames within sequence 100 of image frames, where the selection is less than all image frames within sequence 100 of image frames to be inter-encoded. The selection of frames may be predetermined to include every nth image frame to be inter-encoded, where n is an integer greater than 1, or may be adaptively determined as described in further detail below with reference to Figures 5 and 6. Image frames that are to be inter-encoded and not included in the selection may be encoded using a first encoding pass, and the second encoding pass and steps associated with the second encoding pass may be skipped for these frames.
[0023] FIG. 2 shows an encoder 200 for inter-encoding image frames in a sequence of image frames, such as one or more of the image frames 104-1 to 104-6 of FIG. 1. The encoder 200 includes a circuit 202 configured to perform any method described herein. More specifically, the circuit 202 is configured to implement the functions of the encoder 200, here represented by a first-pass encoding function 204, a second-pass encoding function 206, a compression level acquisition function 208, a pixel block identification function 210, and a compression level reduction function 212. From the disclosure herein, it will be understood that the encoder 200 in some embodiments has more functions than those shown in FIG. 2. The arrows in FIG. 2 reflect the flow of data between the functional blocks of the encoder 200. For example, the compression level obtained by function 208 is input to the first-pass encoding function 204 along with the inter-encoded image frame 104-i. The output of the first-pass encoding function 204, i.e., the inter-encoded version of the image frame 104-i, is provided as an input to the pixel block identification function 210 together with the compression level obtained by the compression level acquisition function 208. Next, the identified pixel blocks are input to the compression level reduction function 212, which outputs the compression level to be used by the second-pass encoding function 206 when the image frame 104-i is inter-encoded a second time. The second-pass encoding function 206 finally outputs the inter-encoded version 106-i of the image frame 104-i.
[0024] Accordingly, the functions disclosed herein may be implemented using a circuit or processing circuit including a general-purpose processor, a dedicated processor, an integrated circuit, an ASIC ("application-specific integrated circuit"), conventional circuitry, and / or combinations thereof configured or programmed to perform the disclosed functions. In the present disclosure, a circuit is hardware that performs or is programmed to perform the described functions. The hardware may be any hardware disclosed herein or known in other ways that is programmed or configured to perform the described functions.
[0025] In a purely hardware implementation form, each of functions 204, 206, 208, 210, 212 can be dedicated and can have corresponding circuits specially designed to implement the functions. The circuits may be in the form of one or more application-specific integrated circuits or one or more integrated circuits such as one or more field programmable gate arrays. As an example, the first pass encoding function 204 can thus correspond, during use, to a circuit that inter-encodes an image frame in a first encoding pass using the compression level obtained for each pixel block of the image frame, and the pixel block identification function 210 can correspond, during use, to a circuit that identifies pixel blocks within an image frame that are intra-coded in a first encoding pass and for which the obtained compression level exceeds a compression level threshold.
[0026] In an implementation form that also includes software, circuit 202 can include a processor. Since the processor includes transistors and other circuits therein, it can be regarded as a processing circuit or a circuit. In this case, circuit 202 can be regarded as a combination of hardware and software, and the software is used to configure the hardware and / or the processor. More specifically, the processor is configured to operate in relation to a memory 218 and computer code stored on the memory 218. Each of functions 204, 206, 208, 210, 212 can correspond to a part of the computer code stored in the memory 218 that causes the encoder 200 to perform the functions when executed by the processor. Thus, the combination of the processor, the memory 218, and the computer code causes the functions 204, 206, 208, 210, 212 of the encoder 200 to occur.
[0027] In view of the above, the memory 218 can thus constitute a (non-transitory) computer-readable storage medium, such as a non-volatile memory, which, when executed by a computer, includes computer program code for causing the computer to execute any of the methods herein. Examples of non-volatile memory include read-only memory, flash memory, ferroelectric RAM, magnetic computer storage devices, optical disks, and the like.
[0028] It is also possible to have a combination of a hardware implementation form and a software implementation form, which means that some of the functions 204, 206, 208, 210, 212 are implemented by dedicated circuits and other functions are implemented in software, that is, in the form of computer code executed by a processor. For example, in one embodiment, the first path encoding function 204 and the second path encoding function 206 are implemented in hardware. In particular, they may be implemented by a hardware-implemented encoding unit 214 that executes both the first and second encoding paths. Thus, the first encoding path and the second encoding path can be executed by the same encoding unit 214, that is, by the same hardware, or alternatively, by the same combination of hardware and software. This is advantageous in that it is not necessary to include additional hardware (or software) to implement an additional second encoding path. Further, the compression level acquisition function 208, the pixel block identification function 210, and the compression level reduction function 212 can be implemented as software by, for example, a software-implemented control unit 216 executed by a processor.
[0029] Next, the operation of the encoder 200 when executing a method for inter-encoding image frames within a sequence of image frames, such as an image frame 104-i corresponding to one of the image frames 104-1 to 104-6 of the image frame sequence 100, will be described with reference to the flowchart of FIG. 3 and further with reference to FIGS. 1, 2, and 4a to 4d.
[0030] In step 02, the compression level of each pixel block of the image frame is obtained by the compression level acquisition function 208. The obtained compression level is higher for pixel blocks in some regions of the image frame than for pixel blocks in other regions of the image frame. For example, as shown in FIG. 4a, the image frame 104-i inter-coded by the encoder 200 depicts a scene including a tree and a lake with subtle light changes due to clouds passing in front of the sun. Also, on the path in front of the lake, two moving objects, a runner and a cyclist, are traveling. Some regions 402 of the image frame 104-i may contain information that is more interesting to the viewer than other regions 403. For example, a moving object such as a person running or riding a bicycle may be more interesting than the lake or tree depicted in the background. Regions 402 that may be more interesting to the user are generally referred to as regions of interest (ROI). Usually, it is more important to maintain a higher level of quality in the ROI than in the non-ROI. Therefore, the compression level can be higher for pixel blocks that are not within the region of interest of the image frame 104-i than for pixel blocks that are within the region of interest of the image frame 104-i. This is shown in FIG. 4b, which shows the compression levels obtained for different regions of the image frame 104-i. The higher compression level 404 shown in white is applied to pixel blocks within the background region 403 of the image frame 104-i, and the lower compression level 406 shown in black is applied to pixel blocks within the region of interest 402 that depicts the runner and the cyclist. It should be understood that the examples of FIGS. 4a and 4b are simplified in that they show only two different compression levels for illustrative purposes. In an actual example, a wide range of compression levels can be set for different pixel blocks within the image frame according to their relevance or importance.
[0031] To obtain the compression level for each pixel block, the compression level acquisition function 208 can apply any known algorithm suitable for this purpose. In particular, it can apply an algorithm that first determines the relevance of different regions within the image frame and then sets the compression level of the pixel block according to the relevance of the region where the pixel block is located. The relevance can be set such that pixel blocks in more relevant regions are assigned a lower compression level than pixel blocks in less relevant regions. What characteristics are considered relevant can vary between different applications. For example, in the case of a surveillance application, the relevance of a region can be set based on the level of detail within the region. Regions with a low level of detail, such as those that appear empty, are not particularly interesting for video surveillance and are therefore considered irrelevant. Regions with a medium level of detail, such as those depicting a person or a vehicle, are of great interest for video surveillance and are therefore considered to have high relevance. At the other end of the scale, regions with many small details, such as grass or foliage, are also not particularly interesting for video surveillance and are considered irrelevant. An example of an algorithm that can be used is described in the applicant's European Patent No. 3021583.
[0032] Some of the available algorithms for determining the relevance of different image regions operate on the luma channel of the image frame and are color-blind in the sense that they ignore the chroma channels. These algorithms cannot detect relevant movements of image details that exist in the chroma channels but not in the luma channels, for example, when there is an aurora reflection on the lake in Figure 4a. As a result, these algorithms set a high compression level in regions where the relevance of the content in the luma channel is low but the relevance of the content in the chroma channel is higher (e.g., including movement). It has been found that the high compression levels in these regions, combined with the fact that there is content (movement) in the chroma channel, which is costly to encode, trigger the appearance of the type of encoding artifacts that the method 300 aims to reduce.
[0033] In step 04, the encoding function 204 of the first path of the encoder 200 inter-encodes the image frame 104-i in the first encoding path using the compression level obtained for each pixel block of the image frame, and each pixel block is either inter-coded or intra-coded. For this purpose, the first path encoding function 204 can implement inter-encoding according to video encoding technologies such as H.264, H.265, AV1, VP9 without the need for further modification. The determination of whether to intra-code or inter-code a pixel block is made for each block in a manner known per se by the first path encoding function 204 for the purpose of minimizing the encoding cost of each pixel block. For example, the cost of inter-coding a pixel block can be calculated and compared with the calculated cost of intra-coding the same pixel block. If the cost of inter-coding a pixel block is lower than the cost of intra-coding the same pixel block, a decision is made to inter-code the pixel block. Otherwise, a decision is made to intra-code the pixel block. Further, since the costs for inter-coding and intra-coding a pixel block depend on the compression level, the decision to inter-code or intra-code a pixel block depends on the compression level of the pixel block. For example, reducing the compression level can increase the likelihood that a pixel block will be inter-coded. Therefore, even if the image frame 104-i is inter-coded, in the first encoding path of step S04, some of the pixel blocks are intra-coded and other pixel blocks are inter-coded, which may result in artifacts observed in the high compression level regions. In particular, in the example of FIG. 4a, the subtle light changes over the lake can be a potential source of such artifacts.
[0034] In step 06, the pixel block identification function 210 identifies pixel blocks in the image frame 104-i that are intra-coded in the first coding path and whose obtained compression level exceeds the compression level threshold. The compression level threshold is determined such that intra-coded pixel blocks within the non-ROI of the image frame can be identified by the pixel block identification function 210, while intra-coded pixel blocks within the ROI 402 may remain un-identified. A specific value can be assigned to the compression level threshold. The compression level threshold may be the same for all pixel blocks. However, since the appropriate threshold to use varies depending on the image content within the image frame, it may be difficult to set the compression level threshold with respect to a specific absolute value applied to the entire image frame. Therefore, it is preferable to set the compression level threshold for each pixel block. This can be achieved by each pixel block in the image frame having a respective compression level threshold set in relation to the compression level used when the spatially corresponding pixel block was last intra-coded in the previous image frame within the sequence. In this way, the compression level threshold is not set absolutely for the entire image frame, but rather for each pixel block, in relation to the compression level of the intra-coded spatially corresponding pixel block in the previous image frame. The compression level of the spatially corresponding pixel block is presumed to have similar image content as the pixel block and thus functions as a baseline level for setting the compression level threshold of the pixel block. A higher baseline level results in a higher compression level threshold for the pixel block and vice versa. For purposes of illustration, assume that this method is currently applied to image frame 104-3 of the sequence of image frames 100. The first pixel block in image frame 104-3 has spatially corresponding pixel blocks in each of the previously encoded image frames 104-2, 104-1, and 102-1.Furthermore, assume that the pixel block spatially corresponding to the first pixel block was intra-coded when image frames 102-1 and 104-1 were encoded, but was inter-coded when image frame 104-2 was encoded. Thus, since image frame 104-1 was encoded after image frame 104-1, the pixel block spatially corresponding to the first pixel block was last intra-coded in image frame 102-1. Thus, the compression level threshold for the first pixel block is set in relation to the compression level used when intra-coding the spatially corresponding pixel block within image frame 104-1. Similarly, the second pixel block within image frame 104-3 has second pixel blocks spatially corresponding to each of the previously encoded image frames 104-2, 104-1, and 102-1. In this case, assume that the spatially corresponding pixel blocks were intra-coded in intra-frame 102-1, but were not intra-coded in the inter-coded frames 104-1 and 104-2. Thus, the compression level threshold for the second pixel block is set in relation to the compression level used when intra-coding the spatially corresponding pixel block within the intra-coded frame 102-1. In this regard, the pixel block spatially corresponding to a particular pixel block generally means a pixel block that has the same spatial position as the particular pixel block, for example, with respect to pixel coordinates, but is within a different image frame.
[0035] In another example, each pixel block within an image frame has a respective compression level threshold that is set in relation to the compression level used when encoding the spatially corresponding pixel block within the most recently intra-coded frame. Thus, the compression level threshold for each pixel block within image frames 104-1 to 104-5 is set in relation to the compression level used when encoding the spatially corresponding pixel block within intra-frame 102-1. Further, the compression level threshold for each pixel block within image frame 104-6 is set in relation to the compression level used when encoding the spatially corresponding pixel block within intra-frame 102-2. This alternative is easier to implement at the cost of some reduced effectiveness in identifying potentially problematic intra-coding blocks.
[0036] Furthermore, the compression level threshold of a pixel block within an image frame may be set to a predetermined positive offset from the compression level used when the spatially corresponding pixel block was last intra-coded in the previous image frame of the sequence. Thus, the compression level threshold of a pixel block can be said to correspond to a particular relative increase in the compression level, or alternatively, a particular relative decrease in the perceived quality, compared to the spatially corresponding pixel block. It has been found that the same predetermined positive offset can be advantageously used for all pixel blocks. This positive offset may be greater than or equal to the minimum increase in the compression level of an intra-coded pixel block that results in a significant difference in image quality for a human viewer. The selection of the positive offset depends on various factors such as the coding standard used and the encoder implementation. However, as a guideline, a positive offset of approximately 6 steps in the quantization parameter for the H.264 or H.265 coding standard has been found to result in a slightly significant difference in image quality for some encoder implementations, while 10 - 15 steps have been found to result in a clearly visible difference in quality. For example, assuming a predefined positive offset of 6. The compression level threshold for each pixel block within an image frame is calculated by adding this positive offset to the compression level used when the spatially corresponding pixel block was last intra-coded. That is, if the compression level of the spatially corresponding pixel block in the previous frame within the image sequence was 31, the compression level threshold for the spatially corresponding pixel block in the current frame will be 37.
[0037] The compression level may also be predetermined regardless of any compression level used in the previous frame, and may be set to a fixed value such as 35 on a scale between 0 and 51, for example, which means that a pixel block having an obtained compression level exceeding this value is identified by the pixel block identification function 210 when it is also intra-coded by the first pass encoding function 204. In this example, the quantization parameter range of H.264 was used to set the scale. However, it is understood that other scales and thresholds may be used for other codecs.
[0038] As shown in FIG. 4c, three pixel blocks 408 are identified by the pixel block identification function 210, that is, they are intra-coded by the first pass encoding function 204 and have compression values exceeding the compression level threshold. The number 3 was chosen for ease of understanding, and it should be understood that in a real-world example, the number of identified intra-coded blocks can take both other, smaller values and larger values. A comparison with FIGS. 4a and 4c shows that these three pixel blocks 408 are located within the non-ROI of an image frame depicting a lake with slight light variations, and thus, since they are intra-coded blocks within a highly compressed region of the image frame, they have a high probability of causing the artifacts described herein. Steps S08 and S10 further illustrate how the effects of these artifacts can be reduced.
[0039] In step 08, the compression level reduction function 212 reduces the compression level of each identified pixel block. The reason for doing so is to reduce the impact of the artifacts that appear and increase the probability that the pixel block will be inter-coded by the second-pass coding function 206. How much the compression level is reduced varies between embodiments, and the choice of which approach to use in an actual implementation is a trade-off between artifact reduction and bitrate increase. However, it should be remembered that any reduction in the compression level, even the smallest possible, leads to a reduction in artifacts.
[0040] In some embodiments, the compression level is reduced by a fixed predetermined value, such as one step or several steps. In other embodiments, the compression level is reduced depending on the compression level threshold of each pixel block. More specifically, the compression level of each identified pixel block can be reduced to a value below the compression level threshold. Reducing the compression level to such a value has been found to lead to an acceptable trade-off between artifact reduction and bitrate increase in many situations. For example, assume that the compression level threshold for a pixel block within an image frame is set to 34. The pixel block is identified by the pixel block identification function 210 and has an acquired compression level of 42, that is, a compression level 8 steps above the set threshold. The compression level reduction function 212 reduces the compression level of the identified pixel block to 34 and passes the new compression level value to the second-pass coding function 206 for further processing. This is shown in FIG. 4d, where the reduced compression levels 410 of the three identified pixel blocks 408 are shown in a grid pattern compared to the compression levels of the other pixel blocks within the image frame, with those for pixel blocks within the non-ROI of the image frame shown in white and those for pixel blocks within the ROI of the image frame shown in black.
[0041] Furthermore, referring to the embodiment described above with respect to step S06, the compression level of each identified pixel block may be reduced to the compression level used when the spatially corresponding pixel block was last intra-coded within the previous image frame in the sequence. Thus, in this case, the compression level not only decreases below the compression level threshold, but is also reduced to be equal to the compression level of the last intra-coded spatially corresponding pixel block. This sacrifices some higher bitrate, but it has been found to further reduce artifacts associated with the identified pixel blocks.
[0042] In step 10, the encoding function 206 of the second pass uses a reduced compression level for each identified pixel block 408 and the obtained compression level for each remaining pixel block to inter-encode the image frame 104-1 in the second encoding pass. The second pass encoding function 206 generally operates in the same manner as the first pass encoding function 204, i.e., uses the same version of the encoder implementation as the first pass encoding function 204. This can be achieved by the first pass encoding function 204 and the second pass encoding function 206 using the same or identical hardware or software having the same settings for performing inter-encoding in steps S04 and S10. In some embodiments, the first and second encoding passes are executed by the same encoding unit 214, i.e., they are executed by the same encoding hardware and / or encoding software. In other embodiments, the first and second encoding passes are executed by different but identical encoding units, such as by two identical encoder cores, or by two identical instances of encoding hardware or encoding software. Therefore, it should be understood that all features of the first pass encoding function 204 referred to herein generally also apply to the second pass encoding function 206. The main difference between the first encoding pass and the second encoding pass is the compression level applied to each pixel block within the image frame 104-1. In the first encoding pass, the original compression level obtained by the compression level acquisition function 208 is applied to each pixel block within the image frame. In the second encoding pass, as shown in FIG. 4d, the reduced compression level 410 is applied to the pixel block 408 identified by the pixel block identification function 210. As shown in FIGS. 4b and 4d, any pixel block not identified by the pixel block identification function 210 is encoded by the second pass encoding function 206 with the original compression levels 404, 406 shown in black and white, i.e., the compression levels of these non-identified pixel blocks remain unchanged.
[0043] Since the compression level of the identified pixel block is reduced, the appearance of potentially remaining artifacts is reduced, and the overall perceived quality of the image frame is improved. In particular, the difference between the intra-coded block and the inter-coded block in the region of higher compression is less noticeable because the identified intra-coded block is coded with higher quality in the second coding pass. As described above, the reduced compression level also increases the probability that the identified pixel block is inter-coded by the second pass coding function 206, further reducing the impact of any artifacts that may have been present in the image frame prior to coding.
[0044] In some embodiments, the inter-coding in the first coding pass in step S04 operates on the image frame 104-1 with a lower resolution than the inter-coding in the second coding pass in step S10. Specifically, the second coding pass in step S10 can operate at the original (full) resolution of the image frame 104-1, and the first coding pass S04 can operate at a lower resolution of the image frame 104-1. For this purpose, the encoder 200 can further implement a function of providing a low-resolution version of the image frame 104-1 for use in the first coding pass, for example, by downsampling the image frame 104-1. As an example, the lower resolution can correspond to half or a quarter of the original resolution. The advantage of this embodiment is that processing power is saved in the first coding pass because the resolution of the image frame 104-1 is lower. This advantage comes with the risk that, although it has been found that most problematic regions can be identified even at a lower resolution, some regions with artifacts may not be identifiable at a lower resolution.
[0045] In the embodiment implementing step S03, it should be noted that a pixel block in the low-resolution image frame 104-1 can correspond to a plurality of pixel blocks in the image frame at the original resolution. As an example, when the downscaling factor is 4, a 16×16 pixel block in the low-resolution version of the image frame 104-1 can correspond to four 16×16 pixel blocks in the original image frame 104-1. The encoder 200 tracks this correspondence between the pixel blocks in the low-resolution version of the image frame 104-1 and the pixel blocks in the original image frame 104-1. Thereby, the encoder 200 can map the pixel blocks identified in the low-resolution version of the image frame 104-1 to the pixel blocks in the original image frame 104-1 in step S06. Then, these pixel blocks in the original image frame 104-1 can have their compression levels reduced in step S08 as described above.
[0046] As described in connection with FIG. 1, the two-pass inter-coding method need not be performed for all the inter-coded image frames 104-1 to 104-6, but rather for a subset thereof. For the remaining inter-coded image frames, the encoder 200 may instead apply a single-pass inter-coding method, for example, by applying steps S02 and S04 of the method of FIG. 3. The advantage of not applying a second coding pass for each inter-coded image frame is that processing power is saved. However, if the second coding pass is not applied frequently enough, there is a risk of artifacts in the high-compression regions. Therefore, the second coding pass should preferably be applied only when it is necessary to reduce artifacts in the high-compression regions. Two embodiments are described below in which it is adaptively determined for which of the inter-coded frames 104-1 to 104-6 the second coding pass should be applied.
[0047] The first of these two embodiments relates to a method 500 shown in the flowchart of FIG. 5. In addition to the steps of method 300 described in connection with FIG. 3, method 500 further includes a step S07a of determining how often to execute method 500 when inter-coding future image frames in a sequence of image frames based on the number of identified pixel blocks within the currently inter-coded image frame. The frequency can refer to at what frame interval method 500 should be repeated. For example, assume that when encoder 200 first executes method 500 when inter-coding image frame 104-1, and in step S07a, it is determined that method 500 should be executed at a frame interval of 3. Next, encoder 200 executes method 500 when next inter-coding image frame 104-4. When inter-coding intermediate frames 104-2 and 104-3, encoder 200 can apply a single-pass coding scheme, for example, by applying steps S02 and S04 instead.
[0048] In step S07a, the more pixel blocks are identified, the more preferably the method 500 is executed more frequently (e.g., at a shorter frame interval) than when fewer pixel blocks are identified. In this way, the second encoding path is executed less frequently when there are fewer artifacts, as measured by the number of identified pixel blocks, and is executed at a shorter frame interval when the number of identified pixel blocks indicates that there are more artifacts. Thereby, the increase in processing power caused by the second encoding path can be balanced against the need for the second encoding path. Further, since step S07a is repeated each time the method 500 is applied to an image frame, the frame interval is adapted to the current presence of artifacts within the image frame. Thus, the frame interval typically becomes longer during periods of video sequences with fewer artifacts after the first encoding path S04 and becomes shorter during periods with more artifacts after the first encoding path S04. The number of identified pixel blocks may be given as an absolute number or as a percentage. In the latter case, the percentage can correspond to the ratio between the number of identified pixel blocks and the total number of pixel blocks in the frame. Alternatively, it can correspond to the ratio between the number of identified pixel blocks and the number of pixel blocks whose obtained compression level exceeds a compression level threshold.
[0049] To implement step S07a, the encoder 200 can use a look-up table that associates the number of identified pixel blocks with the frame interval at which the method 500 should be repeated. Such a look-up table may be pre-determined, and the values in the table may be set so that a desired trade-off between processing power and artifacts is achieved.
[0050] In this embodiment, similar to what was described in relation to step S03 of FIG. 3, the encoder 200 preferably operates at a resolution lower than the full resolution of the image frame in the first encoding path of step S04 in order to save processing power.
[0051] A second adaptive embodiment relates to method 600 shown in the flowchart of FIG. 6. In addition to the steps of method 300 described in relation to FIG. 3, method 600 further includes step S07b where the encoder 200 checks whether the number of identified pixel blocks is greater than a pixel block threshold. If this condition is met, the encoder 200 proceeds to execute steps S08 and S10. Otherwise, it proceeds to step S12 to encode the next image frame in the sequence of image frames. In this method, the steps of obtaining a compression level S02, inter-encoding the image frame in the first encoding path S04, and identifying pixel blocks S06 are executed each time an image frame is inter-encoded, but the steps of reducing the compression level S08 and encoding the image frame in the second encoding path S10 are executed only if the number of identified pixel blocks in the image frame is greater than the pixel block threshold. The check performed in step S07b can be regarded as a way to check whether a second encoding path is needed. More specifically, if the number of potentially problematic pixel blocks after the first encoding path S04 is large enough, i.e., higher than the pixel block threshold, processing power is spent applying an additional second encoding path using a lower compression level in the problematic blocks. Otherwise, the result after the first encoding path S04 is considered good enough and the result of the first encoding path S04 becomes the output of the encoder 200 for this frame. Thus, method 600 can be applied to each image frame that is inter-coded, but for some of them, the second encoding path of step S10 is executed only as needed. Also in this case, the number of identified pixel blocks may be given as an absolute number or as a percentage.
[0052] In this embodiment, the inter-coded image frames are preferably encoded at full resolution in the first encoding path of step S04 and the second encoding path of step S10. The reason is that even when a negative result is reached in step S07b and the encoder 200 outputs the result from the first encoding path of step S04, it is desirable for the encoder 200 to provide an image frame encoded at full resolution.
[0053] The pixel block threshold may be determined in advance, for example, by applying a test procedure in which different pixel block thresholds are tested for one or more video test sequences and the one that gives a desired trade-off between artifacts and processing power is selected. Also, it is possible to adjust the pixel block threshold while the method 600 is being executed on a video sequence, for example, so that a desired percentage of the inter-coded image frames undergo the second encoding path.
[0054] In the above, only inter-coding has been described. However, it should be understood that the encoder 200 of FIG. 2 can be further configured to intra-code the intra-coded image frames 102-1 and 102-2. When performing intra-coding, the encoder 200 can operate in a conventional manner, and thus intra-coding is not described in more detail herein.
[0055] Those skilled in the art will understand that the above-described embodiments can be modified in many ways and the advantages of the present invention as shown in the above embodiments can still be used. Therefore, the present invention should not be limited to the disclosed embodiments but should be defined only by the appended claims. Further, as will be understood by those skilled in the art, the disclosed embodiments may be combined.
Claims
1. 1. A method for inter-coding image frames in a sequence of image frames, comprising: obtaining a compression level for each pixel block of the image frame, the compression level corresponding to a quantization level and being higher for pixel blocks in some regions of the image frame than for pixel blocks in other regions of the image frame; inter-coding the image frame in a first encoding pass using the obtained compression level for each pixel block of the image frame, each pixel block being either inter-coded or intra-coded; identifying pixel blocks in the image frame that were intra-coded in the first encoding pass and for which the obtained compression level exceeds a compression level threshold; reducing said compression level for each identified pixel block; inter-encoding the image frame in a second encoding pass using the reduced compression level for each identified pixel block and using the obtained compression level for each remaining pixel block; A method comprising:
2. The method of claim 1 , wherein the compression level of each identified pixel block is reduced to a value below the compression level threshold.
3. 2. The method of claim 1, wherein each pixel block in the image frame has a respective compression level threshold that is set relative to the compression level used when a spatially corresponding pixel block having the same spatial location as the pixel block was last intra-coded in a previous image frame in the sequence.
4. 4. The method of claim 3, wherein the compression level threshold for a pixel block in the image frame is set to a predetermined positive offset corresponding to a relative increase in compression level from the compression level used when a spatially corresponding pixel block was last intra-coded in a previous image frame of the sequence.
5. 5. The method of claim 4, wherein the compression level of each identified pixel block is reduced to the compression level used when a spatially corresponding pixel block was last intra-coded in a previous image frame in the sequence.
6. The method of claim 1 , wherein the method is performed each time an image frame in the sequence of image frames is inter-coded.
7. The method of claim 1 , wherein the method is performed for a selection of image frames in the sequence of image frames, the selection being less than every image frame in the sequence of image frames that is inter-coded.
8. 8. The method of claim 7, further comprising determining a frequency with which the method is to be performed when inter-coding future image frames in the sequence of image frames based on a number of identified pixel blocks in a currently inter-coded image frame, wherein a greater number of identified pixel blocks causes the method to be performed more frequently than a smaller number of identified pixel blocks.
9. The method of claim 1 , wherein the inter-coding in the first coding pass operates on a lower resolution of the image frames than the inter-coding in the second coding pass.
10. 8. The method of claim 7, wherein the steps of obtaining a compression level, inter-coding the image frame in a first encoding pass, and identifying pixel blocks are performed every time an image frame is inter-coded, but the steps of reducing the compression level and encoding the image frame in a second encoding pass are performed only on the condition that a number of identified pixel blocks in the image frame is higher than a pixel block threshold.
11. The method of claim 1 , wherein the compression level is higher for pixel blocks that are not within a region of interest of the image frame than for pixel blocks that are within a region of interest of the image frame.
12. The method of claim 1 , wherein the first encoding pass and the second encoding pass are performed by a same encoding unit.
13. An encoder for inter-encoding image frames in a sequence of image frames, comprising circuitry configured to perform the method of claim 1.
14. A computer readable storage medium containing computer program code which, when executed by a computer, causes the computer to perform the method of claim 1.