Method and encoder for inter-frame encoding of image frames

Through the two-pass encoding scheme, the compression level of intra-coded pixel blocks in the high-compression area is identified and reduced, and the problem of artifacts in the high-compression level area in video encoding is solved, and the perceived quality is improved.

CN120021248APending Publication Date: 2025-05-20AXIS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411609115.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-11-12
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

In video encoding, artifacts are prone to occur in image frame areas with high compression levels, resulting in differences in perceived quality, especially at the boundary between intra-coded blocks and inter-coded blocks.

Method used

Using a two-pass encoding scheme, the compression level of intra-coded pixel blocks in the highly compressed area of ​​the image frame is identified and reduced through the first encoding process, and the probability of inter-frame encoding in the second encoding process is increased, thereby reducing the occurrence of artifacts.

Benefits of technology

The artifacts in the high-compression area are effectively reduced, the perceived quality of the image frame is improved, and the impact of potential artifacts that still exist after the second encoding process is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120021248A_ABST
    Figure CN120021248A_ABST
Patent Text Reader

Abstract

The invention relates to a method for inter-encoding an image frame and an encoder. Specifically, a method and an encoder for inter-encoding image frames in a sequence of image frames are provided. The method comprises obtaining (S02) a compression level for each pixel block of the image frame, inter-encoding the image frame using the obtained compression levels in a first encoding process (S04) and identifying (S06) pixel blocks in the image frame that are intra-encoded in the first encoding process and the obtained compression levels exceed a compression level threshold. The method further comprises reducing (S08) a compression level of the identified block of pixels, and inter-encoding the image frame in a second encoding process (S10) using the reduced compression level of the identified block of pixels and the obtained compression level of each remaining block of pixels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video coding. In particular, the present invention relates to methods and encoders for inter-frame coding of image frames in a sequence of image frames. Background Art

[0002] When encoding the image frames of a video, a spatially varying compression level is typically applied to each image frame in order to reduce the bit rate while maintaining a level of perceived quality in different regions of the image frame. A lower compression level may be used in regions of the image frame that are of more interest to the viewer, while a higher compression level may be used for regions that are of less interest to the viewer. This can be achieved by using more bits to encode the regions of interest of the image and fewer bits to encode the less interesting regions. The composition of the region of interest (ROI) can vary depending on the application. For example, the level of movement or the level of image detail can determine whether a particular region of the image should be considered an ROI.

[0003] When using the method as described above, the inventors noticed the problem of artifacts appearing in regions of image frames with a high compression level in certain cases. These types of artifacts are characterized by a significant difference in perceived quality between different zones of the affected region. These artifacts tend to deteriorate and increase over time until the next I-frame is encoded. Examples where these problems have been observed are in the background regions of moving objects in an image, such as the road behind a car, subtle light variations on a wall, or the aurora reflected on a lake surface.

[0004] US10425642B1 discloses a method for improving image quality when there is a large variance in the quantization parameters of the residual coefficients of the coding units applied to an image frame. To reduce the variance, an average quantization parameter is calculated for the coding units of the image frame in a first coding process, and then this average quantization parameter is used to determine the updated quantization parameters to be applied to the coding units in a second coding process.

[0005] US20150373328A1 discloses a variable bit rate system in which the encoder changes the quantization parameter frame by frame using a two-pass coding scheme. A first pass analysis of the entire frame sequence determines which frames are more complex, and a second pass analysis changes the quantization parameters of the frames based on the first pass analysis for more efficient coding.

[0006] EP2132938B1 discloses the use of a two-pass coding scheme for the purpose of meeting bit rate constraints or utilizing unused bandwidth. The coding of a first coding process is refined in a second coding process by changing the video coding mode of the video blocks (such as from skip mode to direct mode) and adjusting the quantization parameter.

[0007] US200260083A1 discloses how to determine the coding quantization parameters of a predetermined block of an image by the dispersion of the statistical sample value distribution based on a high-pass filtered version of the predetermined block, so that the change or adaptation of the coding quantization parameters on the image can be made more effective.

[0008] US20020312021A1 discloses an analysis modulation video compression method, which allows the coding process to dynamically adjust quantization based on the content of the monitored image. A two-pass coding scheme can be applied, in which the first pass is used to derive the quantization parameter values associated with the foreground objects based on, for example, the target bit rate and the number and size of the objects in the scene.

[0009] Therefore, there is a need for improvement in this regard. Summary of the Invention

[0010] In view of the above, the object of the present invention is to overcome or mitigate the above problems by providing an encoding method that improves the perceptual quality in the highly compressed regions of an image frame.

[0011] The above object is achieved by the present invention as defined in the appended independent claims. Advantageous embodiments are defined by the appended dependent claims.

[0012] The inventors have recognized that artifacts appearing in the highly compressed regions of an image frame are caused by some pixel blocks being intra-coded while other adjacent pixel blocks are inter-coded. The inter-coded pixel blocks are temporally predicted based on the information in the previously encoded frames in the video. Thanks to the temporal prediction, the remaining residuals to be encoded are usually small, which allows the pixel blocks to be encoded and perceived as high quality despite the high compression level. This is in contrast to the intra-coded blocks that are spatially predicted only based on the information in the current frame. In that case, the residuals to be encoded are usually much larger, resulting in a relatively low perceptual quality at the same high compression level. As a result, the inter-coded blocks are perceived as having higher quality than the intra-coded blocks. This difference in quality causes artifacts to appear in the decoded video and will be particularly obvious around the boundaries between the intra-coded blocks and the neighboring inter-coded blocks. Applying a deblocking filter at the decoder easily further accentuates the artifacts by introducing ringing artifacts at these boundaries. The inter-coded blocks in the future image frames of the video will also refer to these intra-coded blocks, causing the problem to deteriorate over time.

[0013] To overcome or mitigate the above problems, a two-pass encoding scheme is proposed to identify and selectively reduce the compression of intra-coded pixel blocks located in highly compressed regions of an image frame. The two-pass encoding scheme of the present invention includes the steps of performing a first encoding process using the obtained compression levels applied to each pixel block in the frame to identify intra-coded pixel blocks in the image frame that may cause the described artifacts to appear, i.e., intra-coded pixel blocks with a compression level exceeding a compression level threshold. Before running the second encoding process, reduce the compression levels of the blocks that have been identified as potentially problematic in the first encoding process. Reducing the compression levels for these identified blocks generally has the effect of increasing the probability that the encoder will select to inter-code the identified blocks during the second encoding process. Thus, since artifacts are caused by some pixel blocks being intra-coded while other pixel blocks are inter-coded, the probability of artifacts appearing after the second encoding process is reduced. Further, if the identified blocks are ultimately also intra-coded during the second encoding process, the compression levels will still be reduced and the perceived quality will be improved, reducing the impact of any potential artifacts that still remain after the second encoding process.

[0014] The term "pixel block" should be understood to mean a group of adjacent pixels that have been grouped together. These pixel blocks form units in the image frame, and the encoder operates on the pixel blocks when encoding the image frame. Depending on the encoding standard used to encode the image, these pixel blocks may also be represented as macroblocks, coding tree units, or coding units. Pixel blocks can in most cases be square, for example, including 8×8, 16×16, or 32×32 pixels. Pixels can also be grouped into pixel blocks of other sizes and shapes.

[0015] The term "compression level of a pixel block" refers to the degree or level to which the image data in the pixel block is compressed during the encoding of the pixel block. When compression is achieved through quantization, as is typically the case with transform-based codecs, the compression level can correspond to the quantization level applied when encoding the pixel block. Depending on the encoding standard used to encode the image frame during the first and second encoding processes, such a quantization level may also be represented as a quantization value, quantization parameter, quantization index, or step size.

[0016] The terms "inter-frame encoding of an image frame during a first encoding process" and "inter-frame encoding of an image frame during a second encoding process" mean that the image frame is inter-frame encoded twice, i.e., during two rounds or two stages. The first encoding process thus refers to the first time the image frame is inter-frame encoded, and the second encoding process refers to the second time the image frame is inter-frame encoded. Thus, the method is sometimes also referred to herein as a two-pass encoding method or a two-pass encoding scheme. The purpose of the first encoding process is to identify and reduce the compression level of pixel blocks in the encoded image frame that may generate artifacts. Then, during the second encoding process, the image frame is secondarily encoded using the reduced compression level for the identified pixel blocks to output an encoded video in which the artifacts are alleviated.

[0017] The present invention includes three aspects; a method, an encoder, and a computer-readable storage medium. The second and third aspects generally may have the same features and advantages as the first aspect. Further note that the present invention relates to all combinations of features unless otherwise explicitly stated. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The above and other objects, features, and advantages of the present invention will be better understood from the following illustrative and non-limiting detailed description of embodiments of the present invention with reference to the accompanying drawings, in which like reference numerals will be used for like elements, wherein:

[0019] Figure 1 Shows an image frame sequence including both intra-frame encoded image frames and inter-frame encoded image frames.

[0020] Figure 2 Schematically illustrates an encoder for inter-frame encoding an image frame according to an embodiment.

[0021] Figure 3 Is a flowchart of a method for inter-frame encoding an image frame in an image frame sequence according to an embodiment.

[0022] Figure 4a Shows an image frame of a monitoring scene including two regions of interest.

[0023] Figure 4b Shows applied to Figure 4a The compression levels of different regions of the image frame.

[0024] Figure 4c Shows Figure 4a Three identified pixel blocks in the image frame of.

[0025] Figure 4d Shows for Figure 4c The updated compression levels of the three identified pixel blocks of.

[0026] Figure 5Flowchart of a method for inter - frame encoding of image frames in an image frame sequence, the method adaptively determining which inter - frame encoded frames should be applied with a second encoding process.

[0027] Figure 6 Flowchart of a method for inter - frame encoding of image frames in an image frame sequence, the method adaptively determining which inter - frame encoded frames should be applied with a second encoding process. Detailed implementation

[0028] The present invention will now be described more fully hereinafter with reference to the accompanying drawings which show embodiments of the invention.

[0029] Figure 1 Shows a sequence of image frame 100 including image frames 102 - 1 and 102 - 2 to be intra - frame encoded and image frames 104 - 1 to 104 - 6 to be inter - frame encoded. Using the spatial redundancy of the pixels within the image frame itself without referring to any other image frames, the intra - frame encoded image frames are encoded independently of all other image frames in the image frame sequence. These types of intra - frame encoded image frames are commonly referred to as intra - frames, I - frames or key frames. Using the temporal redundancy of the pixels between the image frames, the inter - frame encoded frames are encoded as being dependent on other frames in the image frame sequence. These types of inter - frame encoded frames are commonly referred to as inter - frames or delta frames and can be forward - predicted frames (P - frames) or bidirectional - predicted frames (B - frames).

[0030] The present invention includes a method for inter - frame encoding of image frames (such as one or more of image frames 104 - 1 to 104 - 6 in sequence 100) in an image frame sequence using a two - pass encoding scheme, the steps of the method being illustrated in Figure 3 and will be described in further detail below. Each time an image frame among image frames 104 - 1 to 104 - 6 in image frame sequence 100 is inter - frame encoded, this method can be executed. Alternatively, this method can be executed for a selection of the image frames in image frame sequence 100, the selection being less than all the image frames to be inter - frame encoded in the image frame sequence. The selection of frames can be predetermined, such as including every nth image frame to be inter - frame encoded, where n is an integer greater than 1, or can be adaptively determined as will be described in more detail with reference to Figure 5 and Figure 6 Image frames to be inter - frame encoded and not included in the selection can be encoded using a first encoding process, and for these frames, the second encoding process and steps related to the second encoding process can be skipped.

[0031] Figure 2 Shows for image frames in an image frame sequence (such as Figure 1An encoder 200 that performs inter-frame encoding on one or more of image frames 104-1 to 104-6). The encoder 200 includes a circuit 202 that is configured to perform any of the methods described herein. More specifically, the circuit 202 is configured to implement the functions of the encoder 200, which are illustrated herein by a first-pass encoding function 204, a second-pass encoding function 206, a compression level obtaining function 208, a pixel block identification function 210, and a compression level reduction function 212. It will be understood from the disclosure herein that in some embodiments, the encoder 200 has more functions than Figure 2 illustrated therein. Figure 2 The arrows in illustrate the data flow between the functional blocks of the encoder 200. For example, the compression level obtained by function 208 is input to the first-pass encoding function 204 together with the image frame 104-i to be inter-frame encoded. The output of the first-pass encoding function 204 (i.e., the inter-frame encoded version of the image frame 104-i) is provided as an input to the pixel block identification function 210 together with the compression level obtained by the compression level obtaining function 208. Then, the identified pixel blocks are input to the compression level reduction function 212, which in turn outputs the compression level to be used by the second-pass encoding function 206 when performing inter-frame encoding on the image frame 104-i for the second time. The second-pass encoding function 206 finally outputs the inter-frame encoded version 106-i of the image frame 104-i.

[0032] Thus, the functions disclosed herein can be implemented using a circuit or a processing circuit that includes a general-purpose processor, a dedicated processor, an integrated circuit, an ASIC ("application-specific integrated circuit"), conventional circuitry, and / or a combination thereof, which are configured or programmed to perform the disclosed functions. In the present disclosure, a circuit is hardware that performs or is programmed to perform the recited functions. The hardware can be any hardware disclosed herein or otherwise known that is programmed or configured to perform the recited functions.

[0033] In a pure hardware implementation, each of the functions 204, 206, 208, 210, 212 can have a corresponding circuit that is dedicated and specifically designed to implement that function. The circuit can be in the form of one or more integrated circuits, such as one or more application-specific integrated circuits or one or more field-programmable gate arrays. By way of example, the first-pass encoding function 204 can thus correspond to a circuit that performs inter-frame encoding on an image frame during a first encoding process using the compression level obtained for each pixel block of the image frame, and the pixel block identification function 210 can correspond to a circuit that identifies, during use, pixel blocks in the image frame that are intra-frame encoded during the first encoding process and for which the obtained compression level exceeds a compression level threshold.

[0034] In implementations that also include software, circuit 202 can include a processor. Processors are considered processing circuits or circuits because they include transistors and other circuits therein. In this case, circuit 202 can be regarded as a combination of hardware and software, with the software being used to configure the hardware and / or the processor. More specifically, the processor is configured to operate in association with memory 218 and the computer code stored on memory 218. Functions 204, 206, 208, 210, 212 can each correspond to a portion of the computer code stored in memory 218, which, when executed by the processor, causes encoder 200 to perform that function. Thus, the combination of the processor, memory 218, and the computer code causes functions 204, 206, 208, 210, 212 of encoder 200 to occur.

[0035] In view of the above, memory 218 can thus constitute a (non-transitory) computer-readable storage medium, such as a non-volatile memory, that includes computer program code which, when executed by a computer, causes the computer to perform any of the methods herein. Examples of non-volatile memories include read-only memories, flash memories, ferroelectric RAMs, magnetic computer storage devices, and optical discs, among others.

[0036] It should be understood that there can also be a combination of hardware and software implementations, meaning that some of functions 204, 206, 208, 210, 212 are implemented by dedicated circuits and other functions are implemented in software, i.e., in the form of computer code executed by a processor. For example, in one implementation, the first-pass encoding function 204 and the second-pass encoding function 206 are implemented in hardware. In particular, they can be implemented by an encoding unit 214 implemented in hardware that performs both the first encoding process and the second encoding process. The first encoding process and the second encoding process can thus be executed by the same encoding unit 214, i.e., by the same hardware, or alternatively by the same combination of hardware and software. This is advantageous because no additional hardware (or software) is required to implement an additional second encoding process. Further, the compression level obtaining function 208, the pixel block identification function 210, and the compression level reduction function 212 can be implemented as software, such as a control unit 216 implemented in software executed by a processor.

[0037] Reference will now be made Figure 3 to the flowchart of Figure 1 and further reference will be made Figure 2 to Figures 4a to 4d describe the operation of encoder 200 when performing a method for inter-frame encoding of an image frame (such as an image frame 104-i corresponding to one of image frames 104-1 to 104-6 of image frame sequence 100) in a sequence of image frames.

[0038] In step 02, the compression level of each pixel block of the image frame is obtained by the compression level acquisition function 208. For pixel blocks in some regions of the image frame, the obtained compression level is higher than that of pixel blocks in other regions of the image frame. For example, as Figure 4a shown in, the image frame 104-i to be inter-frame encoded by the encoder 200 depicts a scene including trees and a lake, and there are subtle light changes on the lake surface due to clouds passing in front of the sun. Further, two moving objects, a runner and a cyclist, are moving on the path in front of the lake. Some regions 402 of the image frame 104-i may include information that is more interesting to the viewer than other regions 403. For example, moving objects such as the runner or the cyclist may be more interesting than the lake or trees depicted in the background. Those regions 402 that the user may be more interested in are generally referred to as regions of interest (ROIs). It is usually more important to maintain a high level of quality in the ROI than in the non-ROI. Therefore, the compression level of pixel blocks not within the region of interest of the image frame 104-i can be higher than that of pixel blocks within the region of interest 402 of the image frame 104-i. This is illustrated in Figure 4b which shows the compression levels obtained for different regions of the image frame 104-i. The higher compression level 404 shown in white has been applied to the pixel blocks in the background region 403 of the image frame 104-i, while the lower compression level 406 shown in black has been applied to the pixel blocks in the region of interest 402 depicting the runner and the cyclist. It should be understood that Figure 4a and Figure 4b the examples are simplified because only two different compression levels are shown for the purpose of illustration. In real-world examples, a wide range of compression levels can be set depending on the correlation or importance of different pixel blocks in the image frame.

[0039] To obtain the compression level for each pixel block, the compression level obtaining function 208 may apply any known algorithm suitable for this purpose. In particular, an algorithm may be applied that first determines the correlation of different regions in the image frame and then sets the compression level for the pixel block depending on the correlation of the region in which the pixel block is located. The correlation may be set such that pixel blocks in more correlated regions are assigned a lower compression level than pixel blocks in less correlated regions. Which attributes are considered relevant may vary between different applications. For example, for a surveillance application, the correlation of a region may be set based on the level of detail in the region. Regions with a low level of detail (such as the sky) are not of particular interest for video surveillance and are therefore considered irrelevant. Regions with a medium level of detail (such as regions depicting people or vehicles) are of great interest for video surveillance and are therefore considered to have a high correlation. On the other hand, regions with many small details (such as grass or leaves) are not of particular interest for video surveillance and are considered irrelevant. Examples of algorithms that may be used are described in the applicant's patent EP3021583B1.

[0040] Some of the available algorithms for determining the correlation of different image regions are color blind because they operate on the luminance channel of the image frame but ignore the chrominance channel. Such algorithms, for example, will not be able to detect relevant motion of image details such as an aurora reflection on a lake surface in the chrominance channel but not in the luminance channel as shown in Figure 4a As a result, these algorithms will set a high compression level in regions such as those where the content correlation in the luminance channel containing the motion is low but the content correlation in the chrominance channel is high. It has been found that the high compression level in these regions, combined with the fact that there is content (motion) with high encoding cost in the chrominance channel, triggers the appearance of the type of encoding artifacts that method 300 is designed to mitigate.

[0041] In step 04, the first-pass encoding function 204 of the encoder 200 performs inter-frame encoding on the image frame 104-i using the obtained compression level of each pixel block of the image frame during the first encoding process, where each pixel block is either inter-frame encoded or intra-frame encoded. For this purpose, the first-pass encoding function 204 can implement inter-frame encoding according to video coding techniques such as H.264, H.265, AV1, VP9 without further modification. The first-pass encoding function 204 makes a decision on whether to perform intra-frame encoding or inter-frame encoding on the pixel blocks on a block-by-block basis in a manner known per se, with the aim of minimizing the encoding cost of each pixel block. For example, the cost of performing inter-frame encoding on a pixel block can be calculated and compared with the calculated cost of performing intra-frame encoding on the same pixel block. If the cost of performing inter-frame encoding on a pixel block is lower than the cost of performing intra-frame encoding on the same pixel block, a decision is made to perform inter-frame encoding on the pixel block. Otherwise, a decision is made to perform intra-frame encoding on the pixel block. Further, since the costs of inter-frame encoding and intra-frame encoding of pixel blocks depend on the compression level, the decision on whether to perform inter-frame encoding or intra-frame encoding on a pixel block also depends on the compression level of the pixel block. For example, reducing the compression level may make a pixel block more likely to be inter-frame encoded. Accordingly, even if the image frame 104-i is to be inter-frame encoded, during the first encoding process of step S04, some pixel blocks will be intra-frame encoded and other pixel blocks will be inter-frame encoded, which in turn may produce the observed artifacts in areas of high compression levels. In particular, in the Figure 4a example of, the subtle light variations on the lake surface may be a potential source of such artifacts.

[0042] In step 06, the pixel block recognition function 210 recognizes pixel blocks in the image frame 104-i that are intra-coded in the first coding process and have a compression level exceeding the compression level threshold. The compression level threshold is determined such that intra-coded pixel blocks in the non-ROI of the image frame can be recognized by the pixel block recognition function 210, but intra-coded pixel blocks in the ROI 402 can remain unrecognized. The compression level threshold can be assigned a specific value. The compression level threshold can be the same for all pixel blocks. However, since the appropriate threshold to be used varies with the image content within the image frame, it may be difficult to set the compression level threshold based on an absolute specific value applied to the entire image frame. Therefore, preferably, the compression level threshold is set for each pixel block individually. This can be done by having each pixel block in the image frame have a corresponding compression level threshold, the setting of which is related to the compression level used when the spatially corresponding pixel block in the previous image frame in the sequence was last intra-coded. In this way, the compression level threshold is not set absolutely for the entire image frame, but is set for each pixel block individually, and the setting of the compression level is related to the compression level of the spatially corresponding pixel block that was intra-coded in the previous image frame. The compression level of the spatially corresponding pixel block (which may have similar image content to the pixel block) thus serves as a baseline level for setting the compression level threshold of the pixel block. A higher baseline level results in a higher compression level threshold for the pixel block, and vice versa. For illustration, assume that the method is currently applied to the image frame 104-3 of the image frame sequence 100. The first pixel block in the image frame 104-3 has a spatially corresponding pixel block in each of the previously encoded image frames 104-2, 104-1, and 102-1. Further assume that when the image frames 102-1 and 104-1 were encoded, the pixel block spatially corresponding to the first pixel block was intra-coded, but when the image frame 104-2 was encoded, the pixel block spatially corresponding to the first pixel block was inter-coded. Accordingly, since the image frame 104-1 was encoded later than the image frame 102-1, the pixel block spatially corresponding to the first pixel block was last intra-coded in the image frame 104-1. Therefore, the setting of the compression level threshold for the first pixel block is related to the compression level used when the spatially corresponding pixel block in the image frame 104-1 was intra-coded. Similarly, the second pixel block in the image frame 104-3 has a spatially corresponding second pixel block in each of the previously encoded image frames 104-2, 104-1, and 102-1. In this case, assume that the spatially corresponding pixel block was intra-coded in the intra-frame 102-1 but not in the inter-coded frames 104-1 and 104-2. Therefore, the setting of the compression level threshold for the second pixel block is related to the compression level used when the spatially corresponding pixel block in the intra-coded frame 102-1 was intra-coded.In this regard, a spatially corresponding pixel block of a particular pixel block generally means a pixel block that has the same spatial position as the particular pixel block (e.g., in terms of pixel coordinates) but in another image frame.

[0043] In another example, each pixel block in an image frame has a corresponding compression level threshold, and the setting of the corresponding compression level threshold is related to the compression level used when encoding a spatially corresponding pixel block in the last intra-coded frame. Thus, the setting of the compression level threshold for each pixel block in image frames 104-1 to 104-5 is related to the compression level used when encoding a spatially corresponding pixel block in intra-frame 102-1. Further, the setting of the compression level threshold for each pixel block in image frame 104-6 is related to the compression level used when encoding a spatially corresponding pixel block in intra-frame 102-2. This alternative is easier to implement, but at the cost of slightly lower efficiency in identifying potentially problematic intra-coded blocks.

[0044] Further, the compression level threshold of a pixel block in an image frame can be set to have a predefined positive offset compared to the compression level used when last intra-coding the spatially corresponding pixel block in the previous image frame of the sequence. Thus, compared to the spatially corresponding pixel block, the compression level threshold of the pixel block can be said to correspond to a specific relative increase in the compression level, or, from a different perspective, to a specific relative decrease in the perceived quality. It has been found that the same predefined positive offset can be advantageously used for all pixel blocks. The positive offset can be equal to or greater than the minimum increase in the image quality of a human viewer caused by the compression level of an intra-coded pixel block. The choice of the positive offset will depend on different factors such as the coding standard used and the encoder implementation. However, as a guide, for some encoder implementations, it has been found that a positive offset of about 6 levels of the quantization parameter of the H.264 or H.265 coding standard results in a significant difference in image quality, while a positive offset of 10-15 levels results in a visibly noticeable quality difference. For example, assume a predefined positive offset of 6. The compression level threshold for each pixel block in the image frame is calculated by adding the positive offset to the compression level used when last intra-coding the spatially corresponding pixel block, i.e., if the compression level of the spatially corresponding pixel block in the previous frame in the image sequence is 31, then the compression level threshold of the spatially corresponding pixel block in the current frame will be 37.

[0045] The compression level can also be predetermined, regardless of any compression level used in previous frames. For example, it can be set to a fixed value such as 35 on a scale between 0 and 51, which means that if a pixel block is also intra-coded by the first-pass encoding function 204, the pixel block with a compression level exceeding this value will be recognized by the pixel block recognition function 210. In this example, the quantization parameter range of H.264 is used to set the scale. However, it should be understood that other scales and thresholds can be used for other codecs.

[0046] As Figure 4c illustrated, three pixel blocks 408 have been recognized by the pixel block recognition function 210, that is, they have been intra-coded by the first-pass encoding function 204 and have a compression value exceeding the compression level threshold. The number 3 is chosen for ease of understanding, and it should be understood that in a real-world example, the number of recognized intra-coded blocks can take other smaller or larger values. A comparison with Figure 4a and Figure 4c shows that these three pixel blocks 408 are located in the non-ROI of the image frame depicting the lake surface, where there are subtle light variations on the lake surface, and thus are very likely to cause the artifacts described herein because they are intra-coded blocks in the highly compressed region of the image frame. Steps S08 and S10 will further describe how to mitigate the effects of these artifacts.

[0047] In step 08, the compression level reduction function 212 reduces the compression level of each recognized pixel block. The reason for doing this is to reduce the impact of artifacts and increase the probability that the pixel block will be inter-coded by the second-pass encoding function 206. How much the compression level is reduced varies between implementations, and the choice of which method to use in actual implementation will be a trade-off between artifact mitigation and bitrate increase. However, it should be remembered that any reduction in the compression level, even the smallest possible reduction, will result in a reduction in artifacts.

[0048] In some implementations, the compression level is reduced by a fixed predetermined value, such as reducing one level or several levels. In other implementations, the compression level is reduced depending on the compression level threshold of each pixel block. More specifically, the compression level of each recognized pixel block can be reduced to a value equal to or lower than the compression level threshold. It has been found that in many cases, reducing the compression level to such a value can achieve an acceptable trade-off between artifact mitigation and bitrate increase. For example, assume that the compression level threshold of a pixel block in an image frame is set to 34. The pixel block is recognized by the pixel block recognition function 210 and has an obtained compression level of 42, that is, 8 levels higher than the set threshold. The compression level reduction function 212 reduces the compression level of the recognized pixel block to 34 and passes the new compression level value to the second-pass encoding function 206 for further processing. This is in Figure 4dIn the figure, the reduced compression levels 410 of the three identified pixel blocks 408 are shown in a grid pattern, compared to the compression levels of other pixel blocks in the image frame. Pixel blocks in the non-ROI of the image frame are shown in white, and pixel blocks in the ROI of the image frame are shown in black.

[0049] Furthermore, referring to the embodiment discussed above with respect to step S06, the compression level of each identified pixel block can be reduced to the compression level used when the spatially corresponding pixel block in the previous image frame in the sequence was last intra-coded. In this case, the compression level is thus reduced not only to be equal to or lower than the compression level threshold, but also to be equal to the compression level of the last intra-coded spatially corresponding pixel block. It has been found that this can even further reduce the artifacts associated with the identified pixel blocks, although at the cost of a slightly higher bit rate.

[0050] In step 10, the second-pass encoding function 206 performs inter-frame encoding on the image frame 104-1 using the reduced compression levels of each identified pixel block 408 and the obtained compression levels of each remaining pixel block. The second-pass encoding function 206 generally operates in the same manner as the first-pass encoding function 204, i.e., it is implemented using the same version of the encoder as the first-pass encoder function 204. This can be achieved by the first-pass encoding function 204 and the second-pass encoding function 206 using the same or similar hardware or software with the same settings for performing inter-frame encoding in steps S04 and S10. In some embodiments, the first encoding process and the second encoding process are performed by the same encoding unit 214, i.e., they are performed by the same encoding hardware and / or encoding software. In other embodiments, the first encoding process and the second encoding process are performed by different but similar encoding units, such as two similar instances of encoding hardware or encoding software, such as through two similar encoder cores. Therefore, all features of the first-pass encoding function 204 mentioned herein should be understood to generally also apply to the second-pass encoding function 206. The main difference between the first encoding process and the second encoding process lies in the compression levels applied to each pixel block in the image frame 104-1. In the first encoding process, the original compression levels obtained by the compression level obtaining function 208 are applied to each pixel block in the image frame. In the second encoding process, as Figure 4d shown in the figure, the reduced compression levels 410 are applied to the pixel blocks 408 identified by the identifying pixel block function 210. As Figure 4b and Figure 4d shown in the figure, any pixel block not identified by the identifying pixel block function 210 will be encoded by the second-pass encoding function 206 with the original compression levels 404, 406 shown in black and white, i.e., the compression levels of these un-identified pixel blocks will remain unchanged.

[0051] Since the compression level of the identified pixel blocks is reduced, the occurrence of any potential remaining artifacts will be reduced, and the overall perceived quality of the image frame will be improved. In particular, the difference between the intra-coded blocks and the inter-coded blocks in the regions of higher compression will be less obvious, because the identified intra-coded blocks will be coded with higher quality in the second coding process. As previously mentioned, the reduced compression level also increases the probability that the identified pixel blocks will be inter-coded by the second pass coding function 206, further reducing the impact of any artifacts that may be present in the image frame before coding.

[0052] In some embodiments, the inter-coding in the first coding process in step S04 operates at a lower resolution of the image frame 104-1 than the inter-coding in the second coding process in step S10. In particular, the second coding process in step S10 can operate on the original (full) resolution of the image frame 104-1, while the first coding process S04 can operate on a lower resolution of the image frame 104-1. For this purpose, the encoder 200 can further implement a function such as downsampling the image frame 104-1 to provide a lower resolution version of the image frame 104-1 for use in the first coding process. For example, the lower resolution can correspond to half or a quarter of the original resolution. The advantage of this embodiment is that, due to the lower resolution of the image frame 104-1, processing power is saved in the first coding process. This advantage brings the risk that some regions may not be able to identify artifacts at the lower resolution, although it has been found that, in practice, most of the problematic regions are also identifiable at the lower resolution.

[0053] For the embodiments implementing step S03, note that the pixel blocks at the lower resolution in the image frame 104-1 can correspond to several pixel blocks at the original resolution in the image frame. For example, in the case of a reduction factor of 4, a 16x16 pixel block in the lower resolution version of the image frame 104-1 can correspond to four 16x16 pixel blocks in the original image frame 104-1. The encoder 200 tracks this correspondence between the pixel blocks in the lower resolution version of the image frame 104-1 and the pixel modules in the original image frame 104-1. This allows the encoder 200 to map the pixel blocks identified in the lower resolution version of the image frame 104-1 to the pixel blocks in the original image frame 104-1 in step S06. These pixel blocks in the original image frame 104-1 can then have their compression level reduced in step S08 as explained above.

[0054] As combined with Figure 1As mentioned, the two-pass inter-frame encoding scheme is not necessarily performed on all of the image frames 104-1 to 104-6 to be inter-frame encoded, but rather on a subset thereof. For the remaining image frames to be inter-frame encoded, the encoder 200 may instead (e.g., by applying Figure 3 the steps S02 and S04 of the method) apply a single-pass inter-frame encoding scheme. The advantage of not applying the second encoding process to all of the image frames to be inter-frame encoded is that processing power is saved. However, if the second encoding process is not applied frequently enough, there is a risk of artifacts appearing in high-compression regions. Therefore, the second encoding process should preferably be applied only when it is necessary to reduce artifacts in high-compression regions. In the following, two embodiments will be described in which it is adaptively determined which of the inter-frame encoded frames 104-1 to 104-6 should be applied the second encoding process.

[0055] The first of these two embodiments relates to Figure 5 the method 500 illustrated in the flowchart of. In addition to the steps of the method 300 described in conjunction with Figure 3 the method 500 further includes step S07a of determining, based on the number of identified pixel blocks in the currently inter-frame encoded image frame, how frequently the method 500 should be performed when inter-frame encoding future image frames in the image frame sequence. The frequency may refer to at which frame interval the method 500 should be repeated. For example, assume that the encoder 200 first performs the method 500 when inter-frame encoding the image frame 104-1, and determines in step S07a that the method 500 should be performed at three-frame intervals. Then, the encoder 200 will perform the method 500 the next time it inter-frame encodes the image frame 104-4. When inter-frame encoding the intermediate frames 104-2 and 104-3, the encoder 200 may instead (e.g., by applying steps S02 and S04) apply a single-pass encoding scheme.

[0056] In step S07a, a higher number of identified pixel blocks preferably causes method 500 to be performed more frequently, such as with a lower frame interval, compared to a lower number of identified pixel blocks. Thus, when the presence of artifacts (as measured by the number of identified pixel blocks) is low, the second encoding process runs less frequently, and when the number of identified pixel blocks indicates a high presence of artifacts, the second round of the encoding process runs with a shorter frame interval. This allows the increase in processing power caused by the second encoding process to be balanced with the demand for the second encoding process. Further, since step S07a is repeated each time method 500 is applied to an image frame, the frame interval is adapted to the artifacts currently present in the image frame. Thus, when the presence of artifacts is low after the first encoding process S04, the frame interval will generally be longer during a time period of the video sequence, and when the presence of artifacts is high after the first encoding process S03, the frame interval will be shorter during the time period. The number of identified pixel blocks can be given as an absolute number or as a percentage. In the latter case, the percentage can correspond to the ratio between the number of identified pixel blocks in the frame and the total number of pixel blocks. Alternatively, it can correspond to the ratio between the number of identified pixel blocks and the number of pixel blocks for which the obtained compression level exceeds a compression level threshold.

[0057] To implement step S07a, encoder 200 can use a look-up table that associates the number of identified pixel blocks with the frame interval at which method 500 should be repeated. Such a look-up table can be predetermined, and the values in the table can be set such that a desired trade-off between processing power and artifacts is achieved.

[0058] In this embodiment, similar to that described in Figure 3 step S03 in conjunction with

[0059] For reasons of saving processing power, encoder 200 preferably operates at a resolution lower than the full resolution of the image frame during the first encoding process of step S04. Figure 6 The second adaptive embodiment relates to Figure 3In addition to the steps of method 300 described, method 600 further includes step S07b where encoder 200 checks whether the number of identified pixel blocks is higher than a pixel block threshold. If this condition is met, encoder 200 proceeds with steps S08 and S10. Otherwise, step S12 is continued to encode the next image frame in the sequence of image frames. By this method, each time an image frame is inter-frame encoded, steps S02 of obtaining a compression level, step S04 of inter-frame encoding the image frame in a first encoding process, and step S06 of identifying pixel blocks are performed, but steps S08 of reducing the compression level and step S10 of encoding the image frame in a second encoding process are only performed if the number of identified pixel blocks in the image frame is higher than the pixel block threshold. The check performed in step S07b can be regarded as a way of checking whether a second encoding process is needed. More specifically, if the number of potentially problematic pixel blocks after the first encoding process S04 is large enough, i.e., higher than the pixel block threshold, processing power will be used to apply an additional second encoding process with a reduced compression level in the problematic blocks. Otherwise, the result after the first encoding process S04 is considered good enough, and the result of the first encoding process S04 will be the output of encoder 200 for that frame. Method 600 can thus be applied to each image frame to be inter-frame encoded, but for some of these image frames, the second encoding process of step S10 is only performed on an as-needed basis. Also in this case, the number of identified pixel blocks can be given as an absolute number or as a percentage.

[0060] In this embodiment, the image frames to be inter-frame encoded are preferably encoded at full resolution in the first encoding process of step S04 and in the second encoding process of step S10. The reason is that in the case where a negative result is reached in step S07b and encoder 200 outputs the result of the first encoding process from step S04, one also wants encoder 200 to provide the encoded image frame at full resolution.

[0061] The pixel block threshold can be pre-determined, for example, by applying a test program where different pixel block thresholds are tested for one or more video test sequences, and the threshold that gives a desired trade-off between artifacts and processing power is selected. The pixel block threshold can also be adjusted while method 600 is running on a video sequence, such that, for example, a desired proportion of the image frames to be inter-frame encoded undergoes the second encoding process.

[0062] Only inter-frame encoding has been described above. However, it should be understood that Figure 2 encoder 200 can further be configured to intra-frame encode image frames 102-1 and 102-2 to be intra-frame encoded. When performing intra-frame encoding, encoder 200 can operate in a conventional manner, and thus intra-frame encoding is not described in more detail herein.

[0063] It will be recognized that those skilled in the art can modify the above-described embodiments in various ways and still utilize the advantages of the present invention shown in the above-described embodiments. Accordingly, the present invention should not be limited to the embodiments shown, but should be defined only by the appended claims. Additionally, as will be understood by those skilled in the art, the embodiments shown can be combined.

Claims

1. A method for inter-frame coding of image frames in an image frame sequence, comprising: obtaining a compression level for each pixel block of the image frame, wherein the compression level corresponds to a quantization level and the compression level of pixel blocks in some areas of the image frame is higher than the compression level of pixel blocks in other areas of the image frame, inter-coding the image frame using the obtained compression level of each pixel block of the image frame in a first encoding process, wherein each pixel block is inter-coded or intra-coded, identifying a pixel block in the image frame that has been intra-coded in the first coding process and for which the obtained compression level exceeds a compression level threshold, reducing the compression level for each identified pixel block, and The image frame is inter-coded in a second encoding process using the reduced compression level for each identified pixel block and using the obtained compression level for each remaining pixel block.

2. The method according to claim 1, wherein: The compression level of each identified pixel block is reduced to a value equal to or below the compression level threshold.

3. The method according to claim 1, wherein: Each pixel block in the image frame has a corresponding compression level threshold, and the setting of the corresponding compression level threshold is related to the compression level used when the spatially corresponding pixel block in the previous image frame in the sequence was last intra-frame encoded, and the spatially corresponding pixel block has the same spatial position as the pixel block.

4. The method according to claim 3, wherein: The compression level threshold for a pixel block in the image frame is set to have a predefined positive offset from the compression level used when a spatially corresponding pixel block in a previous image frame of the sequence was last intra-frame encoded, and the predefined positive offset corresponds to a relative increase in the compression level.

5. The method according to claim 4, wherein: The compression level for each identified pixel block is reduced to the compression level used when the spatially corresponding pixel block in a previous image frame in the sequence was last intra-coded.

6. The method according to claim 1, wherein: The method is performed each time an image frame in the sequence of image frames is inter-frame encoded.

7. The method according to claim 1, wherein: The method is performed on a selection of image frames in the sequence of image frames, the selection being less than all image frames in the sequence of image frames to be inter-coded.

8. The method according to claim 7, further comprising: Based on the number of identified pixel blocks in the current inter-frame encoded image frame, determine how often the method is performed when inter-frame encoding future image frames in the image frame sequence, wherein a higher number of identified pixel blocks causes the method to be performed more frequently than performing the method at a lower number of identified pixel blocks.

9. The method according to claim 1, wherein: The inter-frame encoding in the first encoding process operates at a lower resolution of the image frame than the inter-frame encoding in the second encoding process.

10. The method according to claim 7, wherein: The steps of obtaining a compression level, inter-coding the image frame in a first encoding process and identifying pixel blocks are performed each time an image frame is inter-coded, but the steps of reducing the compression level and encoding the image frame in a second encoding process are performed only if the number of identified pixel blocks in the image frame is above a pixel block threshold.

11. The method according to claim 1, wherein: The compression level of pixel blocks that are not within the region of interest of the image frame is higher than the compression level of pixel blocks that are within the region of interest of the image frame.

12. The method according to claim 1, wherein: The first encoding process and the second encoding process are performed by the same encoding unit.

13. An encoder for inter-coding an image frame in a sequence of image frames, comprising circuitry configured to perform the method of claim 1.

14. A computer readable storage medium comprising computer program code which, when executed by a computer, causes the computer to perform the method according to claim 1.

Citation Information

Patent Citations

  • Efficient video block mode changes in second pass video coding

    EP2132938B1

  • Method of identifying relevant areas in digital images, method of encoding digital images, and encoder system

    EP3021583B1

  • Noisy media content encoding

    US10425642B1

  • Content adaptive bitrate and quality control by using frame hierarchy sensitive quantization for high efficiency next generation video coding

    US20150373328A1