Method, device, and recording medium for encoding / decoding image for machine
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ELECTRONICS & TELECOMM RES INST
- Filing Date
- 2026-01-21
- Publication Date
- 2026-07-30
Smart Images

Figure KR2026001263_30072026_PF_FP_ABST
Abstract
Description
Method, device, and recording medium for encoding / decoding images for a machine
[0001] The present disclosure relates to a method, apparatus, and recording medium for encoding / decoding images for a machine.
[0002] Traditionally, video encoding and decoding technologies have achieved improved compression efficiency and image quality by considering the human visual system. However, in the future, video encoding and decoding technologies are expected to be widely utilized not only for human vision but also in machine vision fields such as surveillance, intelligent transportation, smart cities, and intelligent factories.
[0003] Accordingly, there is a need to develop image encoding / decoding technology capable of achieving high-efficiency compression and recognition accuracy by simultaneously considering both human vision and machine vision.
[0004] The present disclosure aims to resolve the problem of image quality degradation caused by bit cutting by proposing a method to improve the Y component of a video frame after bit cutting.
[0005] A method, apparatus, and recording medium for encoding / decoding an image for a machine may include: decoding bit depth shift related information from a bitstream; performing a bit depth shift for a luminance sample based on the bit depth shift related information; and performing an enhancement technique based on the bit depth shift related information.
[0006] In a method, apparatus, and recording medium for encoding / decoding images for a machine, the bit depth shift related information may include a first flag indicating whether the bit depth shift is enabled.
[0007] In a method, apparatus, and recording medium for encoding / decoding images for a machine, the bit depth shift related information includes shift degree information indicating the degree of shift of an input luminance sample in the bit depth shift, and the shift degree information may be decoded in response to the bit depth shift related information indicating that the bit depth shift is activated.
[0008] In a method, apparatus, and recording medium for encoding / decoding images for a machine, the bit depth shift related information may include at least one of a second flag indicating whether the enhancement technique is performed or a first index information indicating the type of the enhancement technique.
[0009] In a method, apparatus, and recording medium for encoding / decoding images for a machine, the bit depth shift related information may include a second index information indicating whether the enhancement technique is performed and the type of the enhancement technique.
[0010] In a method, apparatus, and recording medium for encoding / decoding images for a machine, the enhancement technique may be at least one of a Contrast Limited Adaptive Histogram Equalization (CLAHE) technique, an Unsharp filter, or a Reserved technique.
[0011] By applying the method of enhancing the Y component of a video frame after bit cutting of the present disclosure, it is possible to solve the problem of image quality degradation that may occur due to a reduction in bit depth, particularly in high-resolution content where restoration is omitted.
[0012] FIG. 1 is a block diagram of an image encoder according to one embodiment of the present disclosure.
[0013] FIG. 2 is a block diagram of an image decoder according to one embodiment of the present disclosure.
[0014] Figure 3 illustrates the syntax structure for bit cutting.
[0015] Figures 4 and 5 illustrate experimental results using CLAHE.
[0016] Figures 6 and 7 illustrate experimental results using an Unsharp Filter.
[0017] FIG. 8 illustrates a flowchart of an image decoding method for a machine of the present disclosure.
[0018] FIG. 9 illustrates a flowchart of an image encoding method for a machine according to the present disclosure.
[0019] FIG. 10 illustrates an example of the CLAHE technique.
[0020] FIG. 11 illustrates an example of the application of an unsharp filter.
[0021] Figures 12 and 13 illustrate drawings related to the Reserved technique.
[0022] Figure 14 illustrates an example of parameters related to the external enhancement filter of the syntax structure for bit cutting.
[0023] Figure 15 illustrates an example in which several enhancement techniques (enhancement filters) of the syntax structure for bit cutting are applied.
[0024] The present disclosure is subject to various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the present disclosure. Similar reference numerals in the drawings refer to the same or similar functions across various aspects. The shapes and sizes of elements in the drawings may be exaggerated for clearer explanation. The detailed description of exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments as examples. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that various embodiments are different but need not be mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the present disclosure in relation to one embodiment. It should also be understood that the location or arrangement of individual components within each disclosed embodiment may be changed without departing from the spirit and scope of the embodiment. Accordingly, the following detailed description is not intended to be taken in a limiting sense, and the scope of the exemplary embodiments is limited only by the appended claims, together with all equivalents to those claimed therein, provided they are properly described.
[0025] In this disclosure, terms such as first, second, etc. may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of this disclosure, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.
[0026] Where it is stated that any component of the present disclosure is "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, or that there may be other components in between. On the other hand, where it is stated that a component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0027] The components shown in the embodiments of the present disclosure are depicted independently to represent different characteristic functions and do not imply that each component consists of separate hardware or a single software unit. That is, each component is listed and included as a separate component for convenience of explanation; however, at least two of the components may be combined to form a single component, or a single component may be divided into multiple components to perform a function, and such integrated and separated embodiments of each component are included within the scope of the rights of the present disclosure as long as they do not depart from the essence of the present disclosure.
[0028] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit this disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this disclosure, terms such as "comprising" or "having" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof. That is, the description in this disclosure that a specific configuration "comprising" does not exclude configurations other than that configuration, but means that additional configurations may be included within the scope of the practice or technical concept of this disclosure.
[0029] Some components of the present disclosure may not be essential components performing an essential function in the present disclosure, but may be optional components merely for enhancing performance. The present disclosure may be implemented by including only the components essential to embody the essence of the present disclosure, excluding components used merely for enhancing performance, and a structure including only the essential components, excluding optional components used merely for enhancing performance, is also included within the scope of the rights of the present disclosure.
[0030] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In describing the embodiments of this specification, if it is determined that a detailed description of related known configurations or functions may obscure the gist of this specification, such detailed description is omitted, and the same reference numerals are used for identical components in the drawings, and redundant descriptions of identical components are omitted.
[0031] In the present disclosure, unlike the existing VCM structure, functionality can be extended by defining a class of a VCM-dedicated NAL unit type.
[0032] In addition, the present disclosure discloses an improved bitstream of VCM, an NAL sample stream structure, and an NAL unit type class structure.
[0033] In addition, the present disclosure optimizes the size of the VCM bitstream by separating the SRD and PRD parameters within the RSD sample stream into NAL formats for time, space, ROI, and bit depth, and by removing and reconstructing redundant parameters.
[0034] FIG. 1 is a block diagram of an image encoder according to one embodiment of the present disclosure.
[0035] Referring to FIG. 1, the image encoder may include a preprocessing unit (110) and an image encoding unit (120).
[0036] The preprocessing unit (110) performs a preprocessing process to convert the input original images into images suitable for image encoding. At this time, the images input to the preprocessing unit (110) may be color or black-and-white images following the YUV or YCbCr format.
[0037] The preprocessing unit (110) may include at least one of a temporal resampling unit (112), a spatial resampling unit (114), or a region of interest-based processing unit (116).
[0038] The temporal resampling unit (112) resamples the images temporally. Only the resampled images can be selected as targets for image encoding. That is, through temporal resampling, encoding for some of the images input to the preprocessing unit (110) can be omitted. For example, odd-numbered images of a 60fps (frames per second) image can be omitted to convert the 60fps image into a 30fps image. Alternatively, images of a specific output order can be omitted by considering the temporal overlap between the images.
[0039] The spatial resampling unit (114) spatially resamples the image. Through spatial resampling, the size and / or spatial resolution of the image may be reduced. For example, an image with a resolution of 1920x1080 can be converted into an image of 960x540 or 480x270, etc.
[0040] The region of interest-based processing unit (116) sets a region of interest within the image so that image encoding / decoding is performed primarily on information important for machine inference tasks. The region of interest-based processing unit (116) can remove background areas excluding the set region of interest, or adjust the size and / or position of the region of interest within the image so that the region of interest is encoded / decoded with high quality.
[0041] The video encoding unit (120) can encode the video output from the preprocessing unit (110). Meanwhile, the video encoding unit (120) can encode the video by utilizing a conventional codec technology or a codec technology modified for VCM (Video Coding for Machine) based on a conventional codec technology. For example, the video encoding unit (120) can encode the video based on HEVC, VVC, or AV1. A bitstream is generated as a result of the video encoding, and the generated bitstream can be transmitted to a video decoder.
[0042] FIG. 2 is a block diagram of an image decoder according to one embodiment of the present disclosure.
[0043] Referring to FIG. 2, the image decoder may include an image decoder (210) and a post-processing unit (220).
[0044] The image decoding unit (210) decodes the bitstream received from the image encoder (110) to generate a decoded (or reconstructed) image. The image decoding unit (210) can decode the bitstream based on the codec technology used in the image encoding unit (120).
[0045] The post-processing unit (220) performs post-processing on the decoded image. Through post-processing, the size and frame rate of the image can be restored to match the original image.
[0046] The post-processing unit (220) may include at least one of a post-filtering unit (222), a region of interest-based restoration unit (224), a spatial restoration unit (226), or a temporal restoration unit (218).
[0047] The post-filtering unit (222) applies filtering to reduce the restoration error of the decoded image. For example, the post-filtering unit (222) may apply an in-loop filter to the decoded image. The in-loop filter may include at least one of a deblocking filter, a sample adaptive offset filter, an LMCS (Luma mapping chroma scaling), or an adaptive loop filter.
[0048] In the region of interest-based restoration unit (224), an image of the same size as the original image is obtained based on the region of interest information. For example, if a cropped image containing the region of interest is encoded, the decoded image will have a different size from the original image. Accordingly, the region of interest-based restoration unit can adjust the retargeted image to the original size. Here, the retargeted image may represent the decoded image or an image that has been upscaled through the spatial restoration unit (226). Alternatively, if the size or location of the region of interest within the image to be encoded is adjusted, the region of interest-based restoration unit (224) can adjust the location and size of the region of interest within the retargeted image to match the original image.
[0049] In the spatial restoration unit (226), upscaling is performed on the decoded image. Through upscaling, the decoded image can be restored to the same size and / or spatial resolution as the original image.
[0050] In the temporal restoration unit (228), an image of a temporal location where encoding / decoding was omitted is restored by temporal resampling. Specifically, the temporal restoration unit (228) can generate an image of a temporal location where encoding / decoding was omitted through interpolation between decoded images.
[0051] Meanwhile, additional information may be encoded and signaled to perform reverse processing on the image processing performed in the preprocessing unit (110). In the postprocessing unit (220), postprocessing of the decoded image may be performed based on the additional information to generate an image for machine inference to be performed on a machine. Meanwhile, the additional information may be referred to as 'metadata'.
[0052] Metadata may include at least one of temporal resampling information, spatial resampling information, or region of interest processing information.
[0053] Temporal resampling information may include at least one of a flag indicating whether temporal resampling has been performed or information indicating a temporal resampling rate.
[0054] For example, the above flag being 1 indicates that temporal resampling has been performed. In this case, information indicating the temporal resampling rate may be additionally encoded / decoded. When temporal resampling is performed, fewer images than the number of original images may be encoded / decoded. In the image decoder, images that have been omitted from encoding / decoding can be restored through temporal restoration.
[0055] On the other hand, the above flag being 0 indicates that temporal resampling was not performed.
[0056] The temporal resampling rate can be expressed as a power of 2. For example, a temporal resampling rate of 2^N indicates that one of 2^N images is selected as the image to be encoded / decoded. For instance, only images with a POC (Picture Order Count) that is a multiple of 2^N can be the targets for encoding / decoding. Information representing the temporal resampling rate can represent the power of the temporal resampling rate (i.e., N). For example, the above information can represent the value of the power of the temporal resampling rate or the value obtained by subtracting 1 from the power of the rate.
[0057] Spatial resampling information may include at least one of a flag indicating whether spatial resampling has been performed, or information indicating a scaling parameter for spatial resampling.
[0058] For example, the above flag being 1 indicates that spatial resampling has been performed. In this case, information representing scaling parameters may be additionally encoded. Specifically, information representing horizontal scaling parameters and information representing vertical scaling parameters may each be encoded and signaled. When spatial resampling is performed, the size and / or spatial resolution of the image may be reduced. In the image decoder, the size of the decoded image can be restored to the size of the original image or a preset size through spatial reconstruction. Meanwhile, information for specifying the preset size may be additionally encoded / decoded.
[0059] The above flag being 0 indicates that spatial resampling was not performed.
[0060] The region of interest processing information may include at least one of image size information or region of interest information.
[0061] The image size information may include information indicating whether retargeting has been performed. A retargeting flag of 1 indicates that the retargeted image is encoded / decoded instead of the original image. On the other hand, a retargeting flag of 0 indicates that the original image is encoded / decoded as is.
[0062] A retargeting image represents an image generated by performing at least one of resolution adjustment and position adjustment on at least one region of interest within an original image. Accordingly, the resolution or position of the region of interest within the retargeted image may differ from that of the original image. Additionally, the size of the retargeting image may be the same as or smaller than that of the original image.
[0063] If retargeting is allowed (i.e., when the retargeting flag is 1), size information of the retargeted image can be encoded / decoded. The size information of the retargeted image may include width information of the image and height information of the image.
[0064] Meanwhile, additional encoding / decoding information indicating the size difference between the original image and the retargeted image may be performed. For example, information indicating whether the difference in size between the retargeted image and the original image is encoded / decoded may be encoded / decoded.
[0065] For example, information indicating whether the magnitude difference is encoded / decoded is 0, which indicates that the magnitude difference between the retargeted image and the original image is not encoded / decoded.
[0066] On the other hand, if the information indicating whether the size difference is encoded / decoded is 1, it indicates that the size difference between the retargeted image and the original image is encoded / decoded. In this case, information indicating the difference in size between the retargeted image and the original image may be additionally encoded / decoded.
[0067] Information indicating the size difference represents the size difference between the original image and the retargeted image. Meanwhile, information indicating the size difference in the horizontal direction and information indicating the size difference in the horizontal direction can each be encoded and signaled.
[0068] The region of interest information may include at least one of a flag indicating whether a region of interest exists, number information of the region of interest, scaling parameters of the region of interest, or location information of the region of interest.
[0069] For example, the above flag being 1 indicates that information about the region of interest is encoded / decoded. In this case, at least one of the number of regions of interest, scaling parameter information of the region of interest, location information of the region of interest, or size information of the region of interest may be additionally encoded / decoded.
[0070] On the other hand, if the above flag is 0, it indicates that there is no region of interest.
[0071] Information on the number of regions of interest indicates the number of regions of interest. Meanwhile, the number of regions of interest can be calculated in units of image groups containing at least one image.
[0072] The scaling parameter of the region of interest represents the scaling parameter for the region of interest. Depending on the scaling parameter of the region of interest, the size of the region of interest can be adjusted.
[0073] The scaling parameter information of the region of interest may include information indicating whether the scaling parameter of the region of interest has been updated. If the information indicating whether to update indicates that the scaling parameter of the region of interest has not been updated, the scaling parameter of the region of interest may be set to a default value or the same value as in the previous frame. On the other hand, if the information indicating whether to update indicates that the scaling parameter of the region of interest should be updated, the information indicating the scaling parameter of the region of interest may be additionally encoded / decoded.
[0074] Meanwhile, scaling parameter information for the region of interest can be encoded / decoded individually for each region of interest.
[0075] The location information of the region of interest indicates the location of the region of interest in the original image. At this time, the horizontal position information (i.e., x-axis coordinate) and the vertical position information (i.e., y-axis coordinate) of the region of interest can be encoded / decoded, respectively.
[0076] The size information of the region of interest represents the size of the region of interest in the original image. In this case, the horizontal size (i.e., width) and vertical size (i.e., height) information of the region of interest can be encoded and decoded, respectively.
[0077] As described above, according to the present disclosure, through an image preprocessing process, the encoding / decoding efficiency of the image can be improved while maintaining machine mission performance.
[0078] RBSP of the present disclosure may be an abbreviation for Raw Byte Sequence Payload.
[0079] VCM of the present disclosure may be an abbreviation for Video Coding for Machines.
[0080]
[0081] Bit truncation (bit depth shift) is a technique that reduces the bit rate of a video sequence by lowering the bit depth of the luminance component (Y). This process can be achieved by performing a right shift operation at the encoder stage, followed by conditional left shift restoration based on the video resolution at the decoder stage. Bit truncation may operate differently depending on the resolution. For example, in the case of high resolution, restoration via left shift may generally not be performed.
[0082] The present disclosure proposes a technique for enhancing the Y component of a video frame after bit truncation. This enhancement technique aims to resolve the issue of image quality degradation that may occur due to a reduction in bit depth, particularly in high-resolution content where restoration is omitted.
[0083] The technique proposed in this disclosure may include a post-processing technique specifically designed to compensate for the loss of detail in luminance components caused by bit truncation. This allows for the improvement of the visual quality of a video sequence while maintaining the advantage of reduced bitrate.
[0084] The proposed technique can maintain the structure of existing bit truncation techniques. However, information indicating the type of enhancement technique applied at the encoder stage can be added.
[0085] For example, the above information may include a flag indicating whether an enhancement technique is applied and index information indicating the type of enhancement technique.
[0086] For example, the above information may include one index information indicating whether an enhancement technique is applied and the type of enhancement technique (e.g., 0: not applied, 1: first enhancement technique applied, 2: second enhancement technique applied, etc.).
[0087] For example, the information may include a flag indicating whether an enhancement technique is applied. In this case, a pre-defined technique may be used as the enhancement technique. The pre-defined technique may be at least one of CLAHE, an unsharp filter, or a reserved technique.
[0088] In this disclosure, CLAHE (Contrast Limited Adaptive Histogram Equalization) was primarily used to describe methods for improving frame contrast, but other enhancement techniques may also be utilized. For example, an unsharp filter may be used. For example, a Reserved technique may be used.
[0089] In one embodiment, the first enhancement technique represents CLAHE, the second enhancement technique represents an unsharp filter, and at least one of the first enhancement technique or the second enhancement technique can be adaptively performed.
[0090] In one embodiment, the first enhancement technique represents CLAHE, the second enhancement technique represents a reserved technique, and at least one of the first enhancement technique or the second enhancement technique can be adaptively performed.
[0091] In one embodiment, the first enhancement technique represents an Unsharp filter, the second enhancement technique represents a reserved technique, and at least one of the first enhancement technique or the second enhancement technique can be adaptively performed.
[0092] In one embodiment, the first enhancement technique represents CLAHE, the second enhancement technique represents an unsharp filter, and the third enhancement technique represents a reserved technique, and at least one of the first enhancement technique, the second enhancement technique, or the third enhancement technique can be adaptively performed.
[0093]
[0094] Figure 3 illustrates the syntax structure for bit cutting.
[0095] bit_depth_shift_flag can indicate whether bit cutting (bit depth shift process) is enabled. For example, if the bit value is 1, bit depth shift processing is enabled, and if it is 0, bit depth shift processing is disabled.
[0096] In one embodiment, bit_depth_shift_luma may indicate how much the input luminance samples are shifted during the bit depth shift process. For example, if the bit_depth_shift_luma value is 0, 1, 2, 3, 4, 5, 6, or 7, the input luminance samples may be shifted left by 0, 1, 2, 3, 4, 5, 6, or 7 bits, respectively, during the bit depth shift process. For example, if the value is not specified, the shift may be performed to a predefined value. For example, the predefined value may include 0.
[0097] In one embodiment, if the bit depth shift process includes a process of shifting to the right as well as to the left, bit_depth_right_shift_luma and bit_depth_left_shift_luma may be included instead of bit_depth_shift_luma. Here, bit_depth_right_shift_luma may indicate how much the input luminance samples are shifted to the right during the bit depth shift process. For example, if the bit_depth_right_shift_luma value is 0, 1, 2, 3, 4, 5, 6, or 7, the input luminance samples may be shifted to the right by 0, 1, 2, 3, 4, 5, 6, or 7 bits, respectively, during the bit depth shift process. Additionally, bit_depth_left_shift_luma may indicate how much the input luminance samples are shifted to the left during the bit depth shift process. For example, when the bit_depth_left_shift_luma value is 0, 1, 2, 3, 4, 5, 6, or 7, the input luma samples can be shifted left by 0, 1, 2, 3, 4, 5, 6, or 7 bits, respectively, during the bit depth shift process.
[0098] In one embodiment, bit_depth_shift_chroma may indicate how much the input chroma samples are shifted during the bit depth shift process. For example, if the bit_depth_shift_chroma value is 0, 1, 2, 3, 4, 5, 6, or 7, the input chroma samples may be shifted left by 0, 1, 2, 3, 4, 5, 6, or 7 bits, respectively, during the bit depth shift process. For example, if the value is not specified, the shift may be performed with a predefined value. For example, the predefined value may include 0.
[0099] In one embodiment, if the bit depth shift process includes a process of shifting to the right as well as to the left, bit_depth_right_shift_chroma and bit_depth_left_shift_chroma may be included instead of bit_depth_shift_chroma. Here, bit_depth_right_shift_chroma may indicate how much to the right the input chroma sample is shifted during the bit depth shift process. For example, if the bit_depth_right_shift_chroma value is 0, 1, 2, 3, 4, 5, 6, or 7, the input chroma sample may be shifted to the right by 0, 1, 2, 3, 4, 5, 6, or 7 bits, respectively, during the bit depth shift process. Additionally, bit_depth_left_shift_chroma may indicate how much to the left the input chroma sample is shifted during the bit depth shift process. For example, when the bit_depth_left_shift_chroma value is 0, 1, 2, 3, 4, 5, 6, or 7, the input chroma samples can be shifted left by 0, 1, 2, 3, 4, 5, 6, or 7 bits, respectively, during the bit depth shift process.
[0100] In one embodiment, bit_depth_luma_enhance may indicate whether to use an enhancement technique for the lumina channel (Y) and the type of enhancement technique. If such information is not signaled, it can be inferred that such information has a value of 0.
[0101] For example, a bit_depth_luma_enhance value of 0 may indicate that no enhancement technique is applied to the lumina channel.
[0102] For example, a bit_depth_luma_enhance value of 1 indicates that enhancement is applied using CLAHE (Contrast Limited Adaptive Histogram Equalization). CLAHE divides the frame into multiple regions and equalizes the histograms of each region to improve the contrast of the lumina channel while limiting contrast amplification to prevent excessive noise amplification. As a result, contrast can be enhanced and detail preserved, particularly in areas with uniform intensity.
[0103] For example, a bit_depth_luma_enhance value of 2 indicates that an enhancement feature is applied to sharpen the brightness channel using an Unsharp Filter. This feature can improve visual sharpness and clarity by emphasizing edges and fine details.
[0104] For example, when the bit_depth_luma_enhance value is 3, it may indicate that Reserved technology is applied to reinforce image contrast degradation caused by the right shift of the luma channel.
[0105] Since the values of the above examples are merely examples, other combinations of i) no application, ii) CLAHE application, iii) Unsharp filter application, and iv) Reserved application may also be possible.
[0106] For example, the values of bit_depth_luma_enhance can represent the following:
[0107] 0: Not applied, 1: CLAHE applied, 2: Unsharp filter applied, 3: Reserved applied
[0108] 0: Not applied, 1: CLAHE applied, 2: Unsharp filter applied (the order of 1 and 2 can be changed)
[0109] 0: Not applied, 1: CLAHE applied, 2: Reserved applied (the order of 1 and 2 may be swapped)
[0110] 0: Not applied, 1: Unsharp filter applied, 2: Reserved applied (the order of 1 and 2 can be changed)
[0111] 0: Not applied, 1: Apply CLAHE
[0112] 0: Not applied, 1: Unsharp filter applied
[0113] 0 : Not applied, 1 : Reserved applied
[0114]
[0115] In one embodiment, bit_depth_luma_enhance may indicate whether an enhancement technique is applied to the lumina channel (Y). In this case, the enhancement technique may be pre-defined as CLAHE or Unsharp Filter. If such information is not signaled, it can be inferred that such information has a value of 0.
[0116] For example, when bit_depth_luma_enhance is 0, it may indicate that the enhancement technique is not applied to the lumina channel (Y).
[0117] For example, when bit_depth_luma_enhance is 1, it indicates that the lumina channel (Y) enhancement technique is applied.
[0118] In one embodiment, the syntax structure for the bit truncation may further include enhance_thresh_ratio. enhance_thresh_ratio may specify a threshold value for enhancement processing when an enhancement technique is applied to the luminance channel (Y). For example, when an enhancement technique is applied to the luminance channel (Y), enhancement processing may be adaptively performed based on the threshold value for enhancement processing per application unit. For example, the threshold value may have a range of 0 to 3.
[0119] In one embodiment, the syntax structure for bit cutting can be called from the sequence restoration data syntax structure. Additionally, it can be additionally called from the picture restoration data syntax structure. For example, the picture restoration data syntax structure may call the syntax structure for bit cutting only when it is necessary to update the elements (parameters) of the syntax structure for bit cutting.
[0120] Figures 4 and 5 illustrate experimental results using CLAHE.
[0121] Figures 6 and 7 illustrate experimental results using an Unsharp Filter.
[0122] FIG. 8 illustrates a flowchart of an image decoding method for a machine of the present disclosure.
[0123] The image decoding method for a machine of the present disclosure may include: a step of decoding information regarding bit depth shift from a bitstream (S801); wherein the information may include at least one of information indicating whether a bit depth (left or right) shift is applied, information indicating the degree of bit depth shift, information indicating whether an enhancement technique is applied after bit depth shift, or information indicating the type of enhancement technique; a step of performing a bit depth shift based on the information regarding bit depth shift (S802); and a step of performing an enhancement technique based on the information regarding bit depth shift (S803).
[0124] FIG. 9 illustrates a flowchart of an image encoding method for a machine according to the present disclosure.
[0125] The image encoding method for a machine of the present disclosure may include the step of performing a bit depth shift (S901); the step of performing an enhancement technique (S902); and the step of encoding information regarding the bit depth shift into a bitstream based on the bit depth shift and the enhancement technique (S903).
[0126] FIG. 10 illustrates an example of the CLAHE technique.
[0127] The Contrast Limited Adaptive Histogram Equalization (CLAHE) technique may improve machine performance by dividing the frame into tiles, generating a Cumulative Distribution Function (CDF)-based LUT with clip-limit applied to the histogram of each tile, and then interpolating the LUT results of the four surrounding tiles with bilinear interpolation in the x and y directions to enhance image contrast.
[0128] FIG. 11 illustrates an example of the application of an unsharp filter.
[0129] By applying an unsharp filter, the edges and details of the luma channel can be emphasized, making it easier to distinguish between objects.
[0130] Figures 12 and 13 illustrate drawings related to the Reserved technique.
[0131] The Reserved technique can be a technology that can compensate for image contrast degradation caused by the right shift of the luma channel.
[0132] For example, referring to Fig. 12, the Reserved technique can suppress dark and bright areas and selectively enhance midtone contrast by analyzing alpha values determined through a non-linear tone mapping function and the current luma distribution. This can simultaneously improve the visual readability of the image and the feature discrimination ability in machine vision.
[0133] For example, referring to Fig. 13, the Reserved technique can improve machine vision performance by adjusting the gamma value based on the average brightness value of the decoder to correct the luminance value.
[0134] Figure 14 illustrates an example of parameters related to the external enhancement filter of the syntax structure for bit cutting.
[0135] In one embodiment, the syntax structure for bit cutting of the present disclosure may include use_external_bit_depth_luma_enhance_filter_flag. use_external_bit_depth_luma_enhance_filter_flag may indicate whether an external luminance enhancement filter is used.
[0136] For example, if the use_external_bit_depth_luma_enhance_filter_flag value is 1, it may indicate that an external luminance enhancement filter is used. If external luminance enhancement filter information is provided, bit_depth_luma_enhance is ignored and an externally loaded scaling table may be used.
[0137] For example, if the use_external_bit_depth_luma_enhance_filter_flag value is 0, the external luma enhancement filter may not be applied.
[0138] For example, if the use_external_bit_depth_luma_enhance_filter_flag value is 1, or if the external luma enhancement filter is not loaded due to reasons such as insufficient load information, an unacceptable filter size or format, or an empty table (file), the external luma enhancement filter may not be applied. In this case, bit_depth_luma_enhance may be used.
[0139] For example, whether to apply the aforementioned (internal) enhancement technique can be determined by considering bit_depth_luma_enhance, use_external_bit_depth_luma_enhance_filter_flag, and whether the external luma enhancement filter is loaded.
[0140] In one embodiment, whether an external luma enhancement filter is applied may be determined by the bit_depth_luma_enhance value. An embodiment thereof may include an embodiment in which, in the aforementioned embodiments, any one of the bit_depth_luma_enhance values is replaced with one that performs the same role as use_external_bit_depth_luma_enhance_filter_flag. For example, an embodiment may be included in which the bit_depth_luma_enhance value example described above, "0: not applied, 1: CLAHE applied, 2: Unsharp filter applied, 3: Reserved applied," is replaced with "0: not applied, 1: CLAHE applied, 2: Unsharp filter applied, 3: external filter applied."
[0141] In one embodiment, whether an external luma enhancement filter is applied may be determined by an additional bit_depth_luma_enhance value. An embodiment thereof may include an embodiment modified to include one additional value in bit_depth_luma_enhance in the aforementioned embodiments. For example, the embodiment of the bit_depth_luma_enhance value described above, "0: not applied, 1: CLAHE applied, 2: Unsharp filter applied, 3: Reserved applied," may be modified to "0: not applied, 1: CLAHE applied, 2: Unsharp filter applied, 3: Reserved applied, 4: external luma enhancement filter applied."
[0142] In one embodiment, whether an external luma enhancement filter is applied may be determined based on a bit_depth_luma_enhance value indicating a reserved application. For example, i) if the bit_depth_luma_enhance value indicates a reserved application and external enhancement filter information is loaded, an external enhancement filter may be applied. For example, ii) if the bit_depth_luma_enhance value indicates a reserved application and external enhancement filter information is not loaded, the reserved method described above may be applied.
[0143] In one embodiment, although the reserved method and the external enhancement filter were described separately in the above embodiments, conceptually, the reserved method may be considered to include a method for applying the external enhancement filter. In this case, the content regarding the external enhancement filter of the present disclosure may be considered to be included in the reserved method.
[0144] External luma enhancement filter information can be provided using supplementary enhancement information (SEI) messages or any method supported by the decoder.
[0145] Additionally, information on the external luminance enhancement filter can be acquired in at least one unit among sequences, pictures, tiles, blocks, samples, or filters. For example, information on the external luminance enhancement filter can be acquired in sequence units. For example, information on the external luminance enhancement filter can be acquired in picture units. For example, information on the external luminance enhancement filter can be acquired in sequence units and then updated in picture units.
[0146] Figure 15 illustrates an example in which several enhancement techniques (enhancement filters) of the syntax structure for bit cutting are applied.
[0147] use_multiple_luma_enhance_filter_flag can indicate whether lumina enhancement filters (luma enhancement techniques) are applied in series. For example, if the value of use_multiple_luma_enhance_filter_flag is 1, it indicates that multiple lumina enhancement filters are applied in series. That is, multiple filters can be applied sequentially. For example, if the value of use_multiple_luma_enhance_filter_flag is 0, it indicates that only one lumina enhancement filter is applied. Alternatively, if the value of use_multiple_luma_enhance_filter_flag is 0, it indicates that no lumina enhancement filter is applied or that only one lumina enhancement filter is applied.
[0148] num_luma_enhance_filter can specify the number of filters to apply sequentially.
[0149] luma_enhance_filter_idx can be an index specifying each filter applied sequentially.
[0150] Referring to Fig. 15, each enhancement filter (enhancement technique) applied sequentially can be determined by bit_depth_luma_enhance[ luma_enhance_filter_idx[i] ].
[0151] FIG. 15 illustrates an example in which multiple enhancement techniques are applied to an embedded filter, but the present disclosure is not limited thereto and may include and sequentially apply external filters.
[0152] For example, if an external filter is also included, this can be done by replacing or adding a value to bit_depth_luma_enhance that indicates the external filter is applied.
[0153] For example, if an external filter is also included, the value of bit_depth_luma_enhance indicates that the (internal) enhancement technique is not applied (e.g., 0), this can be done by additionally signaling use_multiple_luma_enhance_filter_flag.
[0154] In this case, similar to the content of the aforementioned use_multiple_luma_enhance_filter_flag, whether an external filter is applied to the corresponding filter (the filter specified by luma_enhance_filter_idx[i]) can be determined by additionally considering whether external filter information is loaded.
[0155]
[0156] The names of the syntax elements introduced in the embodiments described above are merely temporary for the purpose of describing the embodiments according to the present disclosure. Syntax elements may be named with names different from those proposed in the present disclosure.
[0157] The components described in the exemplary embodiments of the present disclosure may be implemented by hardware elements. For example, the hardware elements may include at least one of a digital signal processor (DSP), a processor, a controller, an application-specific integrated circuit (ASIC), a programmable logic element such as an FPGA, a GPU, other electronic devices, or a combination thereof. At least some of the functions or processes described in the exemplary embodiments of the present disclosure may be implemented in software, and the software may be recorded on a recording medium. The components, functions, and processes described in the exemplary embodiments may be implemented by a combination of hardware and software.
[0158] A method according to one embodiment of the present disclosure may be implemented as a program that can be executed by a computer, and said computer program may be recorded on various recording media such as magnetic storage media, optical reading media, digital storage media, etc.
[0159] The various technologies described in this disclosure may be implemented as digital electronic circuits or computer hardware, firmware, software, or a combination thereof. The technologies may be implemented as computer program products, namely, computer programs tangibly implemented on information media or computer programs (e.g., machine-readable storage devices (e.g., computer-readable media) or data processing devices), or as computer programs implemented as signals processed by or propagated to perform operations of data processing devices (e.g., programmable processors, computers, or a plurality of computers).
[0160] Computer program(s) may be written in any form of programming language, including compiled or interpreted languages, and may be distributed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. Computer programs may be executed on a single computer, or by multiple computers distributed across one site or multiple sites and interconnected by a communication network.
[0161] Examples of processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, and one or more processors of a digital computer. Generally, a processor receives instructions and data from read-only memory or random access memory, or both. Components of a computer may include at least one processor for executing instructions and one or more memory devices for storing instructions and data. Additionally, the computer may include one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or may be connected to said mass storage devices to receive and / or transmit data. Examples of information media suitable for implementing computer program instructions and data include semiconductor memory devices (magnetic media such as hard disks, floppy disks, and magnetic tapes), optical media such as compact disc read-only memory (CD-ROM) and digital video discs (DVD), magneto-optical media such as floptical disks, and Read Only Memory (ROM), Random Access Memory (RAM), flash memory, Erasable Programmable ROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), and other known computer-readable media. Processors and memory may be complemented or integrated by special-purpose logic circuits.
[0162] A processor may execute an operating system (OS) and one or more software applications running on the OS. A processor unit may also access, store, manipulate, process, and generate data in response to software execution. For simplification, a processor unit is described in the singular; however, those skilled in the art will understand that the processor unit may include multiple processing elements and / or various types of processing elements. For example, a processor unit may include multiple processors or a processor and a controller. It may also constitute different processing structures, such as parallel processors. Furthermore, a computer-readable medium means any medium accessible to a computer and may include both computer storage media and transmission media.
[0163] The present disclosure includes detailed descriptions of various detailed embodiments, but such details are not intended to limit the invention or claims proposed in the present disclosure and should be understood as describing the features of specific exemplary embodiments.
[0164] Features individually described in exemplary embodiments in this disclosure may be implemented by a single exemplary embodiment. Conversely, various features described with respect to a single exemplary embodiment in this disclosure may be implemented by a combination of multiple exemplary embodiments or a suitable sub-combination. Furthermore, in this disclosure, said features may operate by a specific combination and may be described as said combination first claimed, but in some cases, one or more features may be excluded from the claimed combination, or the claimed combination may be changed into a sub-combination or a modified form of a sub-combination.
[0165] Likewise, even if operations are described in a specific order in the drawings, it should not be understood that it is necessary to execute the operations in a specific sequence or order, or that all operations must be performed, in order to obtain the desired result. In certain cases, multitasking and parallel processing may be useful. Furthermore, it should not be understood that the various device components in the exemplary embodiments of all embodiments must be separated, and the aforementioned program components and devices may be packaged into a single software product or multiple software products.
[0166] The exemplary embodiments disclosed in this specification are merely illustrative and are not intended to limit the scope of this disclosure. Those skilled in the art will recognize that various modifications to the exemplary embodiments may be made without departing from the spirit and scope of the claims and their equivalents.
[0167] Accordingly, the present disclosure shall be deemed to include all other substitutions, modifications, and changes falling within the scope of the following claims.
[0168] The present disclosure may be used in industries related to image codecs for machines using methods, devices, and recording media for encoding / decoding images for machines.
Claims
1. A step of decoding bit depth shift related information from a bitstream; Based on the bit depth shift related information above, a step of performing a bit depth shift on a luminance sample; and An image decoding method characterized by including the step of performing an enhancement technique on the luminance sample based on the bit depth shift related information.
2. In Paragraph 1, An image decoding method characterized in that the bit depth shift related information includes a first flag indicating whether the bit depth shift is enabled.
3. In Paragraph 2, The bit depth shift related information includes shift degree information indicating the degree of shift of the input luminance sample in the bit depth shift, and An image decoding method characterized by decoding the shift degree information in response to the bit depth shift related information indicating that the bit depth shift is activated.
4. In Paragraph 1, An image decoding method characterized in that the bit depth shift related information includes at least one of a second flag indicating whether the enhancement technique is performed or a first index information indicating the type of the enhancement technique.
5. In Paragraph 1, An image decoding method characterized in that the bit depth shift related information includes a second index information indicating whether the enhancement technique is performed and the type of the enhancement technique.
6. In Paragraph 1, An image decoding method characterized in that the enhancement technique is at least one of the Contrast Limited Adaptive Histogram Equalization (CLAHE) technique, an Unsharp filter, or a Reserved technique.
7. A step of performing a bit depth shift on the luminance sample; A step of performing an enhancement technique on the above luminance sample; and An image encoding method characterized by including the step of encoding bit depth shift-related information into a bitstream based on the bit depth shift and enhancement technique.
8. In Paragraph 7, A video encoding method characterized in that the bit depth shift related information includes a first flag indicating whether the bit depth shift is enabled.
9. In Paragraph 8, The bit depth shift related information includes shift degree information indicating the degree of shift of the input luminance sample in the bit depth shift, and An image encoding method characterized by encoding the shift degree information in response to the bit depth shift-related information indicating that the bit depth shift is activated.
10. In Paragraph 7, An image encoding method characterized in that the bit depth shift related information includes at least one of a second flag indicating whether the enhancement technique is performed or a first index information indicating the type of the enhancement technique.
11. In Paragraph 7, An image encoding method characterized in that the bit depth shift related information includes a second index information indicating whether the enhancement technique is performed and the type of the enhancement technique.
12. In Paragraph 7, An image encoding method characterized in that the enhancement technique is at least one of the Contrast Limited Adaptive Histogram Equalization (CLAHE) technique, an Unsharp filter, or a Reserved technique.
13. A step of performing a bit depth shift on the luminance sample; A step of performing an enhancement technique on the above luminance sample; Based on the bit depth shift and the enhancement technique, a step of encoding bit depth shift-related information into a bitstream; and A bitstream transmission method characterized by including the step of transmitting the bitstream.
14. In Paragraph 13, A bitstream transmission method characterized in that the bit depth shift related information includes a first flag indicating whether the bit depth shift is enabled.
15. In Paragraph 14, The bit depth shift related information includes shift degree information indicating the degree of shift of the input luminance sample in the bit depth shift, and A bitstream transmission method characterized by encoding the shift degree information in response to the bit depth shift related information indicating that the bit depth shift is activated.
16. In Paragraph 13, A bitstream transmission method characterized in that the bit depth shift related information includes at least one of a second flag indicating whether the enhancement technique is performed or a first index information indicating the type of the enhancement technique.
17. In Paragraph 13, A bitstream transmission method characterized in that the bit depth shift-related information includes a second index information indicating whether the enhancement technique is performed and the type of the enhancement technique.
18. In Paragraph 13, A bitstream transmission method characterized in that the enhancement technique is at least one of the Contrast Limited Adaptive Histogram Equalization (CLAHE) technique, an Unsharp filter, or a Reserved technique.