Method of encoding and decoding by feature-based adaptive mapping of regions of interest in VCM
By using the region-of-interest-based adaptive mapping encoding and decoding method in VCM, the server load and power consumption problems in machine image analysis are solved, and efficient image processing is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANWHA VISION CO LTD
- Filing Date
- 2024-09-20
- Publication Date
- 2026-04-24
Smart Images

Figure CN121925850A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an encoding and decoding method in VCM using feature-based adaptive mapping of regions of interest. Background Technology
[0002] With the continuous development of the information and communication industry, broadcasting services with HD (high definition) resolution have spread globally.
[0003] This spread has made many users familiar with high-resolution and high-definition images and / or videos, and has increased the demand for high-resolution and high-quality images and videos, such as UHD (Ultra-High Definition) images / videos with higher definition, i.e., 4K or 8K or even higher definition, in various fields.
[0004] The technology used to encode this UHD image data was completed in 2013 using the standard technology HEVC (High Efficiency Video Coding).
[0005] HEVC is a next-generation image compression technology with higher compression ratio and lower complexity than the previous H.264 / AVC technology, and it is a core technology for effectively compressing large amounts of data in HD and UHD images.
[0006] Like previous compression standards, HEVC performs encoding on a block-by-block basis.
[0007] However, unlike H.264 / AVC, it has only one configuration file. The core coding techniques included in HEVC's single configuration file include layered coding structure techniques, transform techniques, quantization techniques, intra-frame predictive coding techniques, inter-frame motion prediction techniques, entropy coding techniques, loop filtering techniques, and other techniques, totaling eight areas.
[0008] Since the establishment of the HEVC video codec in 2013, with the expansion of services such as realistic images using 4K and 8K video images and virtual reality, a new standard, VVC (Variety Video Coding), has been developed as the next-generation video codec, aiming to improve performance by at least twice that of HEVC. VVC is also known as H.266.
[0009] H.266 (VVC) was developed with the goal of achieving twice the efficiency of the previous generation codec H.265 (HEVC). Initially developed with 4K resolution or higher in mind, VVC has been expanded to handle massive 16K ultra-high resolution images due to the growing VR market and its responsiveness to 360-degree imaging. Furthermore, as the HDR market has expanded along with display technology advancements, VVC supports not only 10-bit color depth but also 16-bit color depth to accommodate the HDR market, and supports brightness levels of 1000 nits, 4000 nits, and 10000 nits. Additionally, because VVC is being developed with the VR and 360-degree imaging markets in mind, it supports partial frame rates from 0 FPS to 120 FPS.
[0010] The Development of Artificial Intelligence
[0011] Artificial intelligence (AI) is also gradually developing. AI refers to the artificial imitation of human intelligence, that is, the intelligence that can recognize, classify, reason, predict, control / make decisions, etc.
[0012] With the development of artificial intelligence technology and the increase in Internet of Things (IoT) devices, machine-to-machine traffic is expected to increase dramatically, and machine-based image analytics is expected to be widely used.
[0013] However, server load and power consumption issues are expected as the number of images to be analyzed by machines is anticipated to increase exponentially. Summary of the Invention
[0014] Technical issues
[0015] Therefore, this disclosure provides an encoding and decoding method in VCM using feature-based adaptive mapping of regions of interest to efficiently perform machine-based image analysis.
[0016] Technical solutions
[0017] To achieve the objectives described above, based on a disclosure in this specification, a feature-based adaptive mapping method for encoding and decoding in VCM using regions of interest is proposed.
[0018] A VCM encoding apparatus may include: a region of interest extractor for extracting one or more regions of interest per frame of an input image; a region of interest feature-based processor for obtaining a mapping relationship of one or more regions of interest based on the features of one or more regions of interest and performing the mapping; and an internal encoder for generating a bitstream by encoding the input image using information about one or more regions of interest.
[0019] A VCM decoding apparatus according to one disclosure of this specification is provided. A VCM decoding apparatus may include: an internal decoder for performing decoding on a bitstream to generate a restored image and extracting information about regions of interest (ROIs) in the restored image; and a ROI restorer for performing restoration on each or more ROIs included in the restored image based on the information about the ROIs.
[0020] According to one disclosure of this specification, as a non-volatile computer-readable storage medium for recording commands, when executed by at least one processor, the at least one processor may include: extracting one or more regions of interest per frame of an input image; obtaining a mapping relationship of one or more regions of interest based on features of the one or more regions of interest and performing the mapping; and generating a bitstream by encoding the input image using information about the one or more regions of interest.
[0021] Technical effect
[0022] Based on this disclosure, machine-based image analysis can be performed efficiently. Attached Figure Description
[0023] Figure 1 An example illustrating a video / image coding system.
[0024] Figure 2 This is a diagram schematically illustrating the configuration of a video / image encoding device.
[0025] Figure 3 This is a diagram schematically illustrating the configuration of a video / image decoding device.
[0026] Figures 4a to 4d This is an exemplary diagram representing a VCM encoder and a VCM decoder.
[0027] Figure 5 This is a block diagram of an encoding apparatus according to an embodiment of the present disclosure.
[0028] Figure 6 The diagram illustrates the region of interest extraction and processing based on region of interest features of an encoding apparatus according to an embodiment of the present disclosure.
[0029] Figure 7 Examples of (a) object extraction, (b) region of interest selection, and (c) region of interest classification according to embodiments of this disclosure.
[0030] Figure 8 This is an example of a histogram of the chromaticity components of a region of interest according to an embodiment of this disclosure.
[0031] Figure 9 This is an example of mapping information according to an implementation of this disclosure.
[0032] Figure 10 Examples of histograms (a) before mapping and (b), (c), and (d) after mapping according to embodiments of this disclosure.
[0033] Figure 11 This is a block diagram of a decoding apparatus according to an embodiment of the present disclosure.
[0034] Figure 12 The process of recovering the region of interest in a decoding apparatus according to an embodiment of this disclosure is illustrated. Detailed Implementation
[0035] The specific structural or phased descriptions of embodiments of the concept of this disclosure disclosed in this specification or application are shown only for the purpose of describing embodiments of the concept of this disclosure, and embodiments of the concept of this disclosure may be implemented in various forms and should not be construed as limited to the embodiments described in this specification or application.
[0036] Because various modifications and forms can be made to the embodiments based on the concept of this disclosure, specific embodiments are shown in the accompanying drawings and described in detail in this specification or application. However, it is not intended to limit the embodiments based on the concept of this disclosure to the specific forms disclosed, and it should be understood that they include all changes, equivalents, or substitutions included within the spirit and scope of this disclosure.
[0037] Various components may be described using terms such as first, second, etc., but components should not be limited by these terms. These terms are used only for the purpose of distinguishing one component from other components, and, for example, without departing from the scope of the claims based on the concept of this disclosure, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component.
[0038] When a component is referred to as "linked" or "connected" to other components, it should be understood that the component can be directly linked or connected to other components, but other components may exist in between. On the other hand, when a component is referred to as "directly linked" or "directly connected" to other components, it should be understood that no other components exist in between. Other expressions describing the relationship between components, namely "between" and "directly between," or "adjacent to" and "directly adjacent to," should also be interpreted similarly.
[0039] Since the terminology used in this specification is for describing particular embodiments only, it is not intended to limit the scope of this disclosure. Singular expressions include plural expressions unless the singular expression clearly has a different meaning depending on the context. In this specification, it should be understood that terms such as "comprising" or "having" refer to the presence of the described features, numbers, steps, movements, components, portions, or combinations thereof, but do not preclude the presence or prior addition of one or more other features, numbers, steps, movements, components, portions, or combinations thereof.
[0040] Unless otherwise defined, all terms used herein, including technical or scientific terms, shall have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0041] Terms defined in commonly used dictionaries should be interpreted as having the same meaning as in the context of the relevant technology, and should not be interpreted in an ideal or overly formalized sense unless explicitly defined in this specification.
[0042] In describing the implementation methods, technical details that are well-known in the technical field to which this disclosure pertains and are not directly related to this disclosure have been omitted.
[0043] This is to convey the key points of this disclosure more clearly, without obscuring them by omitting unnecessary descriptions.
[0044] This document relates to video / image coding. For example, the methods / implementations disclosed in this document may relate to the Multifunctional Video Coding (VVC) standard (ITU-T Rec.H.266), next-generation video / image coding standards that comply with VVC, or other video coding-related standards (e.g., High Efficiency Video Coding (HEVC) standard (ITU-T Rec.H.265), Basic Video Coding (EVC) standard, AVS2 standard, etc.).
[0045] This document presents various implementations related to video / image coding, and unless otherwise stated, these implementations can be combined with each other.
[0046] In this document, video can refer to a set of images over time. An image is generally a unit representing an image at a specific time interval, and a slice / tile is a unit that configures a portion of an image in the encoding.
[0047] A slice / tile may include one or more coding tree units (CTUs). An image may consist of one or more slices / tiles. An image may consist of one or more tile groups. A tile group may include one or more tiles.
[0048] A pixel, or pel, can refer to the smallest unit that configures a picture (or image). Alternatively, a "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or pixel value, or it can represent only the pixel / pixel value of the luminance component, or it can represent only the pixel / pixel value of the chrominance component. Alternatively, a sample can refer to a pixel value in the spatial domain, or, when such a pixel value is transformed to the frequency domain, it can refer to the transform coefficients in the frequency domain.
[0049] A unit can represent a basic unit used for image processing. A unit may include a specific region of an image and at least one of the information associated with that region.
[0050] A unit may include one luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the term "unit" can be used interchangeably with terms such as "block" and "region." Typically, A block may include a sample (or sample array) consisting of M columns and N rows or a set (or array) of transform coefficients.
[0051] Figure 1 An example illustrating a video / image coding system.
[0052] Reference Figure 1 A video / image encoding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.
[0053] Source devices may include video sources, encoding devices, and transmitters. Receiving devices may include receivers, decoding devices, and renderers.
[0054] The encoding device can be referred to as a video / image encoding device, and the decoding device can be referred to as a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can include a display unit, and the display unit can include a separate device or an external component.
[0055] Video sources can acquire video / images through processes such as video / image capture, compositing, and generation. Video sources may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more camera devices, video / image files including previously captured video / images, etc. Video / image generation devices may include, for example, computers, tablets, smartphones, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated by computers, etc., and in this case, video / image capture processing can be replaced by processing that generates related data.
[0056] An encoding device can encode input video / images. The encoding device can perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0057] A transmitter can transmit encoded video / image information or data, output in bitstream form, to a receiver of a receiving device via a digital storage medium or network, either as a file or a stream. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network.
[0058] The receiver can receive / extract bit streams and transmit them to the decoding device.
[0059] Decoding devices can decode video / images by performing a series of processes corresponding to the operations of encoding devices, such as dequantization, inverse transform, and prediction.
[0060] The renderer can render decoded video / images. The rendered video / images can then be displayed using a display unit.
[0061] Figure 2 This is a diagram schematically illustrating the configuration of a video / image encoding device.
[0062] In the following text, a video encoding device may include an image encoding device.
[0063] Reference Figure 2The encoding device 10a can be configured to include an image segmenter 10a-10, a predictor 10a-20, a residual processor 10a-30, an entropy encoder 10a-40, an adder 10a-50, a filter 10a-60, and a memory 10a-70. The predictor 10a-20 may include an inter-frame predictor 10a-21 and an intra-frame predictor 10a-22. The residual processor 10a-30 may include a transformer 10a-32, a quantizer 10a-33, a dequantizer 10a-34, and an inverse transformer 10a-35. The residual processor 10a-30 may also include a subtractor 10a-31. The adder 10a-50 may be referred to as a reconstructor or a reconstructed block generator. According to the implementation, the image segmenter 10a-10, predictor 10a-20, residual processor 10a-30, entropy encoder 10a-40, adder 10a-50, and filter 10a-60 described above can be configured by one or more hardware components (e.g., encoder chipsets or processors). Furthermore, the memory 10a-70 may include a decoded image buffer (DPB) and can be configured by a digital storage medium. The hardware components may also include the memory 10a-70 as an internal / external component.
[0064] Image segmenters 10a-10 can split an input image (or picture, frame) input to encoding device 10a into one or more processing units.
[0065] As an example, a processing unit may be referred to as a coding unit (CU). In this case, the coding unit can be recursively split from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree-binary-tritree (QTBTTT) structure. For example, a coding unit can be split into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, a quadtree structure can be applied first, and a binary tree structure and / or a ternary tree structure can be applied later. Alternatively, a binary tree structure can be applied first. The coding process according to this document can be performed based on the final coding unit that is no longer split. In this case, based on coding efficiency and other factors, and according to image characteristics, the maximum coding unit can be directly used as the final coding unit, or, if necessary, the coding unit can be recursively split into deeper coding units, and the optimally sized coding unit can be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, reconstruction, etc., as described later. As another example, the processing unit may also include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transform unit can be split or divided from the final coding unit described above. The prediction unit can be a unit for sample prediction, and the transform unit can be a unit for obtaining transform coefficients and / or a unit for obtaining residual signals from the transform coefficients.
[0066] In some cases, the term "unit" can be used interchangeably with terms such as "block" or "region." Generally, A block can represent a set of transform coefficients or samples consisting of M columns and N rows. Samples can typically represent pixels or pixel values, or they can represent only pixel / pixel values of the luminance component, or only pixel / pixel values of the chrominance component. Samples can be used as a term corresponding to pixels or pels in a picture (or image).
[0067] Subtractor 10a-31 generates a residual signal (residual block, residual sample, or residual sample array) by subtracting the prediction signal (prediction block, prediction sample, or prediction sample array) output from predictor 10a-20 from the input image signal (original block, original sample, or original sample array), and the generated residual signal is transmitted to transformer 10a-32. Predictor 10a-20 can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block including the prediction sample for the current block.
[0068] Predictors 10a-20 can determine whether to apply intra-frame prediction or inter-frame prediction on a block or CU basis. As described below in the description of each prediction mode, the predictor can generate various prediction-related information, such as prediction mode information, and pass it to entropy encoder 10a-40. The prediction-related information can be encoded by entropy encoder 10a-40 and output as a bitstream.
[0069] The intra-frame predictor 10a-22 can predict the current block by referring to samples within the current image. Depending on the prediction mode, the reference sample can be located near or far from the current block.
[0070] In intra-frame prediction, the prediction modes can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC modes and planar modes. Depending on the granularity of the prediction direction, the directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes.
[0071] However, this is just an example, and more or fewer directional prediction modes can be used depending on the settings. The intra-frame predictor 10a-22 can also determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.
[0072] Inter-frame predictors 10a-21 can obtain the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially adjacent blocks existing in the current image and temporally adjacent blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally adjacent block may be the same or different. The temporally adjacent block may be referred to as a juxtaposed reference block, juxtaposed CU (colCU), etc., and the reference image including the temporally adjacent block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictors 10a-21 can construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to obtain the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes, and for example, for skip and merge modes, the inter-frame predictor 10a-21 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be transmitted. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and the motion vector difference can be signaled to indicate the motion vector of the current block.
[0073] Predictors 10a-20 can generate prediction signals based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Additionally, the predictor can perform intra-frame block copying (IBC) to predict blocks. Intra-frame block copying can be used, for example, for content image / video coding including games, such as Screen Content Coding (SCC). IBC essentially performs prediction within the current frame, but can be performed in a similar manner to inter-frame prediction because it obtains a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described in this document.
[0074] The predicted signals generated by the inter-frame predictors 10a-21 and / or the intra-frame predictors 10a-22 can be used to generate the reconstructed signal or the residual signal. The transformer 10a-32 can apply transformation techniques to the residual signal to generate transform coefficients. For example, transform techniques may include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Graph-Based Transform (GBT), Conditional Nonlinear Transform (CNT), etc. Here, GBT refers to the transform obtained from a graph when the relationship information between pixels is represented in the graph. CNT refers to the transform obtained by generating the predicted signal using all previously reconstructed pixels and based on it. Furthermore, the transform process can be applied to square pixel blocks of the same size or to non-square blocks of variable size.
[0075] The quantizer 10a-33 quantizes the transform coefficients and transmits them to the entropy encoder 10a-40, which encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. This information about the quantized transform coefficients can be referred to as residual information.
[0076] The quantizer 10a-33 can reorder the block-shaped quantized transform coefficients into a one-dimensional vector shape based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the one-dimensional vector shape of the quantized transform coefficients. The entropy encoder 10a-40 can perform various encoding methods, such as exponential Columbus, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC).
[0077] The entropy encoder 10a-40 can also encode information required for video / image reconstruction, excluding quantized transform coefficients (e.g., values of syntax elements, etc.), either together or separately. The encoded information (e.g., encoded video / image information) can be transmitted as a bitstream or stored in units of the network abstraction layer. The video / image information can also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc. Additionally, the video / image information can also include general constraint information. Information and / or syntax elements signaled / transmitted later in this document can be encoded and included in the bitstream using the encoding process described above. The bitstream can be transmitted over a network or stored on a digital storage medium. Here, the network can include broadcast networks and / or communication networks, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting signals output from the entropy encoder 10a-40 and / or a storage device (not shown) for storing signals output from the entropy encoder 10a-40 may be configured as an internal / external element of the encoding device 10a, or the transmitter may be included in the entropy encoder 10a-40.
[0078] The quantized transform coefficients output from quantizer 10a-33 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients via dequantizer 10a-34 and inverse transform 10a-35. Adder 10a-50 can generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample, or reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from predictor 10a-20. When there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as the reconstructed block. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current image, and can also be used for inter-frame prediction of the next image by filtering, as described below.
[0079] Simultaneously, during image encoding and / or reconstruction processing, luminance mapping with chroma scaling (LMCS) can be applied.
[0080] Filters 10a-60 can apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, filters 10a-60 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and the modified reconstructed image can be stored in memory 10a-70, specifically in the DPB of memory 10a-70. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filtering, bilateral filtering, etc. Filters 10a-60 can generate various filter-related information and transmit it to entropy encoder 10a-90, as described later in the description of each filtering method. The filter-related information can be encoded by entropy encoder 10a-90 and output as a bitstream.
[0081] The modified reconstructed images transferred to the memories 10a-70 can be used as reference images in the inter-frame predictors 10a-80. When inter-frame prediction is applied in this way, the encoding device can avoid prediction mismatch between the encoding device 10a and the decoding device, and can also improve encoding efficiency.
[0082] The DPB of memory 10a-70 can store a modified reconstructed image that will be used as a reference image in inter-frame predictor 10a-21. Memory 10a-70 can store motion information of blocks from which motion information is derived (or encoded) within the current image, and / or motion information of blocks within previously reconstructed images. The stored motion information can be transmitted to inter-frame predictor 10a-21 as motion information for spatially adjacent blocks or temporally adjacent blocks. Memory 10a-70 can store reconstructed samples of reconstructed blocks within the current image and transmit them to intra-frame predictor 10a-22.
[0083] Figure 3 This is a diagram schematically illustrating the configuration of a video / image decoding device.
[0084] Reference Figure 3The decoding device 10b can be configured including an entropy decoder 10b-10, a residual processor 10b-20, a predictor 10b-30, an adder 10b-40, a filter 10b-50, and a memory 10b-60. The predictor 10b-30 may include an inter-frame predictor 10b-31 and an intra-frame predictor 10b-32. The residual processor 10b-20 may include a dequantizer 10b-21 and an inverse transformer 10b-21. According to an embodiment, the entropy decoder 10b-10, residual processor 10b-20, predictor 10b-30, adder 10b-40, and filter 10b-50 described above can be configured by a single hardware component (e.g., a decoder chipset or processor). Furthermore, the memory 10b-60 may include a decoded picture buffer (DPB) and can be configured by a digital storage medium. The hardware component may also include the memory 10b-60 as an internal / external component.
[0085] When a bitstream including video / image information is input, the decoding device 10b can respond to the video / image information in it. Figure 2 The image is reconstructed through processing within the encoding apparatus. For example, the decoding apparatus 10b can obtain units / blocks based on block splitting information obtained from the bitstream. The decoding apparatus 10b can perform decoding using processing units applied in the encoding apparatus. Therefore, the processing unit used for decoding can be, for example, an encoding unit, and the encoding unit can be split from encoding tree units or maximally encoded units according to a quadtree structure, binary tree structure, and / or ternary tree structure. One or more transform units can be obtained from the encoding unit. Furthermore, the reconstructed image signal decoded and output by the decoding apparatus 10b can be played back by a playback device.
[0086] Decoding device 10b can receive data in bitstream form from... Figure 2 The signal output by the encoding device, and the received signal, can be decoded by the entropy decoder 10b-10. For example, the entropy decoder 10b-10 can parse the bitstream to obtain the information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), video parameter sets (VPS), etc. In addition, the video / image information may also include general constraint information.
[0087] The decoding device can additionally decode the image based on information about the parameter set and / or general constraint information. Information and / or syntax elements notified / received by signals, as described later in this document, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 10b-10 can decode the information within the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and output the values of the syntax elements required for image reconstruction, as well as the values of the quantized transform coefficients with respect to the residuals.
[0088] More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element from the bitstream, determine a context model by using information about the decoded target syntax element and decoding information of adjacent target blocks and decoded target blocks, or information about symbols / bins decoded in previous steps, and generate symbols corresponding to the values of each syntax element by performing arithmetic decoding of bins to predict the occurrence probability of bins based on the determined context model. In this case, the CABAC entropy decoding method can update the context model after determining the context model by using information about the decoded symbols / bins for the context model of the next symbol / bin. The prediction information in the information decoded by the entropy decoder 10b-10 can be provided to the predictor 10b-30, and the residual information, i.e., the quantized transform coefficients and related parameter information, obtained by the entropy decoder 10b-10, can be input to the dequantizer 10b-21.
[0089] Additionally, filtering information from the information decoded by the entropy decoder 10b-10 can be provided to the filter 10b-50. Meanwhile, the receiver (not shown) receiving the signal output from the encoding device can be configured as an internal / external element of the decoding device 10b, or the receiver can be a component of the entropy decoder 10b-10. Furthermore, the decoding device according to this document can be referred to as a video / video / image decoding device, and this decoding device can be divided into an information decoder (video / video / image information decoder) and a sample decoder (video / video / image sample decoder). The information decoder can include the entropy decoder 10b-10, and the sample decoder can include at least one of a dequantizer 10b-21, an inverse transformer 10b-22, a predictor 10b-30, an adder 10b-40, a filter 10b-50, and a memory 10b-60.
[0090] Dequantizer 10b-21 can dequantize quantized transform coefficients to output transform coefficients. Dequantizer 10b-21 can reorder the quantized transform coefficients in the form of two-dimensional blocks. In this case, the reordering can be performed based on the coefficient scan order executed by the encoding device. Dequantizer 10b-21 can dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0091] The inverse transformer 10b-22 performs an inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0092] The predictor can perform predictions on the current block and generate a prediction block for the current block, including the predicted samples.
[0093] The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on prediction-related information output from the entropy decoder 10b-10, and can determine a specific intra-frame prediction mode / inter-frame prediction mode.
[0094] The predictor can generate a predicted signal based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Additionally, the predictor can perform intra-frame block copying (IBC) to predict blocks. Intra-frame block copying can be used, for example, for content image / video coding including games, such as Screen Content Coding (SCC). IBC essentially performs prediction within the current frame, but can be performed in a similar manner to inter-frame prediction because it obtains a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described in this document.
[0095] The intra-frame predictor 10b-32 can predict the current block by referring to samples within the current image. Depending on the prediction mode, the reference sample can be located near or far from the current block.
[0096] In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-frame predictor 10b-32 can also determine the prediction mode applied to the current block by using the prediction modes applied to neighboring blocks.
[0097] The inter-frame predictor 10b-31 can obtain the predicted block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.).
[0098] For inter-frame prediction, neighboring blocks can include spatially adjacent blocks existing in the current image and temporally adjacent blocks existing in the reference image. For example, inter-frame predictor 10b-31 can construct a motion information candidate list based on neighboring blocks and obtain the motion vector of the current block and / or the reference image index based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and prediction-related information can include information indicating the inter-frame prediction mode of the current block.
[0099] Adder 10b-40 generates the reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (predicted block, predicted sample array) output from predictor 10b-30. When there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block.
[0100] Adder 10b-40 can be referred to as a rebuilder or rebuild block generator.
[0101] The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current image, and as described below, it can also be output by filtering or used for inter-frame prediction of the next image.
[0102] Meanwhile, luminance mapping with chroma scaling (LMCS) can be applied during image decoding processing.
[0103] Filters 10b-50 can apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, filters 10b-50 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and the modified reconstructed image can be transferred to memory 60, specifically to the DPB of memory 10b-60. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.
[0104] The (modified) reconstructed image stored in the DPB of memory 10b-60 can be used as a reference image in inter-frame predictor 10b-31. Memory 10b-60 can store motion information of blocks from which motion information is derived (or decoded) within the current image, and / or motion information of blocks within previously reconstructed images. The stored motion information can be transmitted to inter-frame predictor 10b-31 as motion information for spatially adjacent blocks or temporally adjacent blocks. Memory 10b-60 can store reconstructed samples of reconstructed blocks within the current image and transmit them to intra-frame predictor 10b-32.
[0105] In this specification, the embodiments described in the predictor 10b-30, dequantizer 10b-21, inverse transformer 10b-22 and filter 10b-50 of the decoding device 10b can also be applied in the same or corresponding manner to the predictor 10a-20, dequantizer 10a-34, inverse transformer 10a-35 and filter 10a-60 of the encoding device 10a.
[0106] As described above, prediction is performed during video encoding to improve compression efficiency. This allows the generation of a prediction block, including predicted samples, for the current block, which is the target block for encoding. Here, the prediction block includes predicted samples in the spatial domain (or pixel domain). The prediction block is obtained by both the encoding and decoding units in the same manner, and the encoding unit can improve image encoding efficiency by signaling information about the residual between the original block and the prediction block (residual information) to the decoding unit, rather than the original sample values of the original block themselves. The decoding unit can obtain a residual block including residual samples based on the residual information, generate a reconstruction block including reconstructed samples by combining the residual block and the prediction block, and generate a reconstructed image including the reconstruction block.
[0107] Residual information can be generated through transformation and quantization processes.
[0108] For example, the encoding device can obtain a residual block between the original block and the prediction block, perform a transform process on the residual samples (residual sample array) included in the residual block to obtain transform coefficients, and signal the relevant residual information to the decoding device (via a bitstream) by obtaining the transform coefficients quantized through a quantization process on the transform coefficients. Here, the residual information may include information such as the value information of the quantized transform coefficients, position information, transform technique, transform kernel, quantization parameters, etc. The decoding device can perform a dequantization / inverse transform process based on the residual information and obtain residual samples (or residual blocks). The decoding device can generate a reconstructed image based on the prediction block and the residual block. The encoding device can also obtain residual blocks by performing dequantization / inverse transform on the quantized transform coefficients for inter-frame prediction reference of subsequent images, and generate a reconstructed image based on this.
[0109] <Video Coding for Machines (VCM)>
[0110] With recent developments in various industrial fields such as surveillance, intelligent transportation, smart cities, smart industries, and smart content, the amount of image or feature map data used by machines is increasing. In contrast, since the conventional image compression methods currently in use are technologies developed by considering the human visual characteristics recognized by viewers, which include unnecessary information, it is inefficient to perform machine tasks. For example, an image based on the vision of a viewer may have a higher resolution than an image based on the data-usage vision of a machine (e.g., a feature map). Therefore, research on video codec technologies for efficiently compressing feature maps is needed to perform machine tasks.
[0111] Video Coding for Machines (VCM) technology is being discussed by the Moving Picture Experts Group (MPEG), an international standardization organization for multimedia coding. VCM is an image or feature map coding technology based on the data-usage vision of machines (machine vision), rather than the vision of human viewers. In this document, a feature map may be referred to as a feature map, and a feature may be referred to as a feature.
[0112] Figures 4a to 4d are exemplary diagrams showing a VCM encoder and a VCM decoder.
[0113] Referring to Figure 4a , a VCM encoder 100a and a VCM decoder 100b are shown.
[0114] When the VCM encoder 100a encodes a video and / or a feature map and transmits it as a bitstream, the VCM decoder 100b can decode the bitstream and output it. In this case, the VCM decoder 100b can output one or more videos and / or feature maps. For example, the VCM decoder 100b can output a first feature map for analysis using a machine, or output a first image for a user to view. The first image may have a higher resolution than the first feature map.
[0115] Referring to Figure 4b , a feature extractor for extracting a feature map may be connected to the front end of the VCM encoder 100a.
[0116] The VCM encoder 100a may include a feature encoder.
[0117] The VCM decoder 100b may include a feature decoder and a video reconstructor. The feature decoder can decode a feature map from the bitstream and output a first feature map for analysis using a machine. The video reconstructor can reconstruct a first image from the bitstream and output the first image for a user to view.
[0118] Reference Figure 4c A feature extractor for extracting feature maps can be connected to the front end of the VCM encoder 100a. The VCM encoder 100a may include a feature encoder.
[0119] VCM decoder 100b may include a feature decoder. The feature decoder can decode feature maps from a bitstream and output a first feature map for machine analysis. In other words, the bitstream may be encoded only as feature maps, rather than images. To explain further, a feature map may be data that includes information about features used for a specific task of image-based machine processing.
[0120] Reference Figure 4d The feature extractor can be connected to the front end of the VCM encoder 100a.
[0121] The VCM encoder 100a may include a feature converter and a video encoder. The video encoder may be... Figure 2 The encoding device 10a shown is shown.
[0122] Figure 4d The VCM decoder 100b shown may include a video decoder and an inverse converter. The video decoder may be... Figure 3 The decoding device 10b shown is illustrated.
[0123] Figure 5 This is a block diagram of an encoding apparatus according to an embodiment of the present disclosure.
[0124] The encoding apparatus 10a according to embodiments of this disclosure can encode the input signal (input image) using adaptive mapping based on region of interest (ROI) features for VCM. Machine vision applications primarily perform tracking, recognition, classification, segmentation, etc., centered on objects present within an image, and the background is of relatively low importance. In human vision, the size, color, position, etc., of actual objects within an image correspond to important information along with the background; however, in machine vision, the accuracy of machine vision applications (e.g., object recognition, tracking, classification, segmentation, etc.) is important, regardless of information such as the original color or size. The encoding / decoding apparatus according to various embodiments of this disclosure can improve encoding and decoding efficiency through object-centered ROI-based encoding and decoding for the sake of encoding and decoding efficiency without degrading the machine vision application. The ROI can be mapped based on features within the input image, and encoding and decoding efficiency can be improved by performing encoding and decoding on the mapped signal. If necessary, the encoding apparatus 10a can transmit information about the ROI to the decoding apparatus 10b, enabling the decoding apparatus 10b to recover the original image and the ROI.
[0125] The encoding apparatus 10a, according to various embodiments, can receive an image (video) as input and perform encoding to generate a bitstream by adaptive mapping based on regions of interest (ROIs). The encoding apparatus 10a may include a temporal resampler 510, a spatial resampler 520, a ROI extractor 530, a processor 540 based on ROI features, and an internal encoder 550. In various embodiments, Figure 5 The order of the components included in the encoding device 10a shown can be changed, and some components can be omitted. (See reference...) Figure 5 The encoding device 10a can first perform temporal resampling / spatial resampling, and then extract and process the region of interest. Figure 5 Unlike other methods, the encoding apparatus 10a according to the embodiment can first extract the region of interest (ROI), and then perform encoding in the order of ROI extractor 530, temporal resampler 510, spatial resampler 520, ROI-based processor 540, and internal encoder 550 to perform temporal / spatial resampling based on ROI information. Alternatively, in the embodiment, the encoding apparatus 10a can perform encoding in the order of ROI extractor 530, ROI-based processor 540, temporal resampler 510, spatial resampler 520, and internal encoder 550.
[0126] In this disclosure, for ease of explanation, encoding device 10a may be referred to as encoder or encoder, and decoding device 10b may be referred to as decoder or decoder.
[0127] The input image (video) is the input signal to the encoding device 10a for machine vision, and can be an image acquired by various sensors with video characteristics, such as images from general optical cameras, thermal cameras, LiDAR sensors, etc. Depending on the characteristics of the sensor, the configuration of the input image can include a single component or multiple components. For example, an image acquired by an optical camera may include multiple chromaticity components in the RGB and YUV domains.
[0128] The temporal resampler 510 is a step for changing the frame rate of the input image and can perform upsampling or downsampling on a frame-by-frame basis. The frame rate can be determined according to the application purpose of machine vision, and in an embodiment, temporal resampling can be performed according to a predetermined frame rate. Alternatively, the frame rate can be variable in units of a series of frame groups or sequences. When the frame rate is applied variably, the encoding device 10a can transmit information related to temporal resampling (e.g., frame rate, resampling unit, etc.) to the decoding device 10b.
[0129] Spatial resampler 520 is a step for changing the resolution of each frame of the input image, and can perform upsampling or downsampling on a frame-by-frame basis. According to one embodiment, spatial resampler 520 can perform upsampling or downsampling at the same resolution for all frames of the input image. According to another embodiment, spatial resampler 520 can perform upsampling or downsampling by applying different resolutions on a frame-by-frame basis to the input image. Encoding device 10a can transmit information related to spatial resampling (e.g., target resolution, resampling unit, etc.) to decoding device 10b.
[0130] In various embodiments, the encoding device 10a may extract objects from each frame of the input image for a specific purpose required for a machine vision application of the input image or for general machine vision purposes. For example, machine vision applications may include object tracking, recognition, classification, segmentation, etc. In various embodiments, objects may be, for example, people, objects, animals, robots, drones, mobile electronic devices, or vehicles, and may be defined in various forms for various applications.
[0131] The region of interest extractor 530 is a step for distinguishing regions in the input image that are expected to be objects from the background corresponding to parts other than the objects, and can extract the position and size information of the objects by using in-image learning or image processing techniques. According to an embodiment, the region of interest extractor 530 can extract multiple regions of interest from the input image and transmit information about the size and position of the regions of interest to the internal encoder 550 for efficient encoding. The internal encoder 550 can perform encoding by changing the image quality by applying different quantization methods to the objects and the background.
[0132] The feature-based processor 540 is used to analyze the features of the region of interest extracted by the previous region of interest extractor 530 and to perform specific processing on the values inside and outside the region of interest using these features. Detailed steps of the specific processing can be discussed with... Figure 6 The steps are the same.
[0133] According to the implementation, the processor 540 based on region of interest (ROI) features can analyze and process the features by distinguishing the entire ROI extracted from the image from the entire non-ROI region, or it can analyze and process the features of each of the multiple ROI regions obtained from the ROI extractor 530. Alternatively, there may be a method for analyzing the features of the entire ROI and applying the processing for the corresponding features to both the ROI and non-ROI regions equally.
[0134] According to the implementation, the processor 540 based on region of interest features can analyze and process the features of each component of the entire region of interest, or it can analyze and process the features of each component of each of the multiple regions of interest obtained from the region of interest extractor 530.
[0135] According to the implementation, the regions of interest (ROIs) whose features are analyzed and processed can correspond to some of the multiple ROIs extracted by the ROI extractor 530. For example, when the number of ROIs extracted by the ROI extractor 530 is K, only some of the K ROIs can undergo feature analysis and processing, which can be determined by the size, position, shape, etc. of the ROIs. For example, there may be methods for determining the features of ROIs to be processed by the ROI-based processor 540, such as ROIs having a size greater than or equal to / less than or equal to a specific size, ROIs confirmed to have occurred consecutively in multiple frames, ROIs appearing in a specific location in the image, ROIs extracted as rectangular / square regions, etc., and for limiting only ROIs that satisfy the corresponding features to the target of the ROI-based processor 540. Alternatively, the following methods may exist: restricting regions of interest (ROIs) that satisfy only a specific ranking by sorting them by size to the target of processor 540 based on ROI features; restricting the area ratio of the entire ROI to below a certain ratio in the image; and restricting only ROIs within a specific ranking to the target of processor 540 based on ROI features to ensure that the area of the ROIs is restricted to below a finite ratio. When the target of processor 540 based on ROI features is restricted, and according to the implementation, when there are no ROIs satisfying the corresponding features, the steps of the processor after region of interest extractor 530 can be omitted.
[0136] According to the implementation, the components of the region of interest where features are analyzed and processed can be some components, rather than all components. For example, when the input image includes three components in YUV, feature analysis and processing can be performed using one or more components of YUV. In this case, the number of components for feature analysis may not be the same as the number of components for which processing is performed through feature analysis. For example, feature analysis can be performed on the Y component, and all processing based on the corresponding features can be applied to the YUV. Alternatively, processing can be performed only on the components in which feature analysis is performed, and processing can be omitted on the components in which feature analysis is not performed.
[0137] According to embodiments, when the present invention includes as follows Figure 4b When using the feature map extractor shown, the components can be feature maps.
[0138] The internal encoder 550 can perform encoding on the input image. The internal encoder 550 can utilize temporal resampling information in a pre-execution step targeting a reference structure. The internal encoder 550 can determine the reference method for the object and background based on information about the region of interest, or it can change the quantization parameters to perform encoding more efficiently.
[0139] According to the implementation, the execution order of the region of interest extractor 530, the processor 540 based on region of interest features, the temporal resampler 510, and the spatial resampler 520 can be changed. However, when the information of the region of interest extracted from the region of interest extractor 530 of the encoding device 10a is used in the temporal resampler 510 and / or the spatial resampler 520, the region of interest extractor 530 can be executed before the temporal resampler 510 and the spatial resampler 520.
[0140] Figure 6 The diagram illustrates the region of interest extraction and processing based on region of interest features of an encoding apparatus according to an embodiment of the present disclosure.
[0141] The encoding apparatus 10a according to various embodiments can extract regions of interest (ROIs) of an input image and perform one or more processes based on the features of the ROIs. (See also...) Figure 6 and Figure 7 The detailed operation of the region of interest extractor 530 and the processor 540 based on region of interest features of the encoding apparatus 10a is described. Figure 7 Examples of (a) object extraction, (b) region of interest selection, and (c) region of interest classification according to the implementation method.
[0142] In object region extraction step 610, region of interest extractor 530 can extract one or more object regions from the input image. Region of interest extractor 530 can extract six object regions from a single frame. ),like Figure 7 As shown in (a), the region of interest extractor 530 can extract objects using a learning-based neural network model. The region of interest extractor 530 can extract objects frame-by-frame from all frames of the input image, and the number of objects extracted per frame can vary. In some cases, there may be frames from which no objects are extracted. Multiple objects can be extracted within a single frame, and the same object can be extracted from multiple frames. Multiple frames containing the same object can be consecutive or discontinuous. The object extraction results can vary depending on the input image. Frames from which no objects are extracted can be encoded immediately.
[0143] In step 620, the region of interest extractor 530 can select one or more regions of interest within the extracted object region. The region of interest extractor 530 can then select from... Figure 7 Select three object regions from the six object regions in (a) ) as region of interest ( ),like Figure 7 As shown in (b), the region of interest (ROI) can be the target of encoding. According to various embodiments, the ROI extractor 530 can select ROIs within the object region in different ways. In one embodiment, the ROI extractor 530 can select ROIs based on the size information of the object region (e.g., the number of horizontal pixels, the number of vertical pixels, the area of the region, etc.). In another embodiment, the ROI extractor 530 can select an object region larger than a predetermined specific size as the ROI. In another embodiment, the ROI extractor 530 can select an object region smaller than a predetermined specific size as the ROI. In another embodiment, the ROI extractor 530 can select ROIs based on the accuracy of the extracted object using a learning-based neural network model. According to embodiments, the ROI extractor 530 can apply various methods to select ROIs within the object region. The number of ROIs can be greater than or equal to 0 and less than or equal to n, and the number of ROIs can be the same as the number of object regions. When there are 0 ROIs, the corresponding frame can be encoded immediately.
[0144] In the region of interest (ROI) feature analysis step 630, the processor 540, based on ROI features, analyzes the features of each of one or more selected ROIs. According to the implementation, the features of the ROI may be histograms, averages, variances, object edges, object classifications, etc., of the chromaticity components of the ROI, and one or more features may be used. Figure 8 This is an example of a histogram of the chromaticity components of a region of interest according to an embodiment of this disclosure. According to the embodiment, when analysis is performed using the histogram of the chromaticity components of the region of interest as a feature according to the embodiment, the features of each RGB component can be as follows: Figure 8 It is shown in the middle.
[0145] In the region of interest (ROI) classification and chroma sampling step 640, the processor 540 based on ROI features can classify one or more ROIs into M categories based on the features of the analyzed ROIs. For example, when performing the analysis using histograms of each chroma component as features of the ROIs, they can be classified into M categories based on the similarity of the histograms. For example, when the features of the ROIs are as follows: Figure 8When shown in the diagram, (a) and (b) strongly indicate the characteristics of the chromaticity components of R, and (c) strongly indicates the characteristics of the chromaticity components of G. Therefore, (a) and (b) can be classified into one class, and (c) into another. Thus, as... Figure 7 As shown in (c), and It can be classified as ,and It can be classified as Classifying multiple regions of interest (ROIs) based on their chromaticity components is an example, and depending on the implementation, they can be classified in various ways based on various criteria. The classification criteria can be subdivided or simplified depending on the purpose of the machine vision application. Depending on the implementation, object sampling can be performed during the process of classifying ROIs. Chromaticity sampling refers to removing some chromaticity signals from an image signal (input image). For example, it refers to removing one or more of the R, G, and B components from an image that includes RGB. Typically, one component can be selected and removed.
[0146] In the mapping relationship extraction and mapping step 650, the processor 540 based on the region of interest features can obtain the mapping relationship for each chroma component, and perform mapping for the signals in which chroma sampling is performed in the regions of interest divided into M classes. Figure 9 This is an example of mapping information according to an implementation of this disclosure.
[0147] If a component is preserved through chroma sampling, or if the input image is an image signal composed of a single component, the processor 540 based on region of interest features can extract a mapping relationship for that component and perform the mapping. According to the implementation, since chroma sampling is not performed, if the image includes multiple chroma signals, the processor 540 based on region of interest features can perform corresponding processing based on the number of chroma signals. The mapping relationship can be set to various relationships for encoding efficiency.
[0148] Figure 10Examples of histograms before mapping (a) and after mapping (b), (c), and (d) according to embodiments of this disclosure are provided. (b) refers to the changes in the histogram that occur in the mapping relationship, where the min_value corresponding to the minimum value of the determined chromaticity component is mapped to 0, and values greater than this value are mapped by subtracting the min_value, such that these values are generally mapped to values less than the min_value. (c) and (d) refer to the changes in the histogram when the minimum value is mapped to 0 and when the segment size is adaptively changed before and after mapping values greater than or equal to the minimum value. When the gradient between the segment before mapping (delta_value) and the segment after mapping (delta_mapping_value) is less than or equal to 1 / 2, signal loss may occur when the segment before mapping is wider and the segment after mapping is narrower. However, when the mapping segment is appropriately adjusted for cases where there is no signal on the histogram or for segments where the signal size is very small, effective representation can be achieved through mapping without significant signal loss. In various implementations, the processor 540 based on region of interest features can adaptively set the mapping segments and effectively transmit the corresponding information using syntaxes such as VCM_Colour mapping_SP(), mapping_fun_list_data(), VCM_Colour mapping_FP(), ROI_Ioc(), etc.
[0149] The semantics and syntax used in the implementation of this disclosure are as follows.
[0150] In various embodiments, the encoding device 10a can encode information about the mapping relationship of the region of interest, as shown in Tables 1 to 6.
[0151] [Table 1]
[0152]
[0153] [Table 2]
[0154]
[0155] [Table 3]
[0156]
[0157] [Table 4]
[0158]
[0159] When the number and increment of segments and the size of segments are fixed by the agreement between the encoder and the decoder, the processor 540 based on region of interest features according to the implementation can omit the corresponding information and transmit the mapping information as shown in Table 5 below.
[0160] [Table 5]
[0161]
[0162] According to Table 5, the mapping ratio between delta_value[k] and mapping_delta_value[k] can be a fixed value, and for mapping efficiency, it can be... The mapping is performed in the form of . n is fixed to a specific value by the agreement between the encoder and decoder, and may not be transmitted. Here, n can be a positive number, 0, or a negative number. Depending on the implementation, the mapping ratio n between delta_value[k] and mapping_delta_value[k] can be transmitted from the encoder to the decoder. The syntax for the case where the mapping ratio n is transmitted to the decoder can be defined as shown in Table 6 below.
[0163] [Table 6]
[0164]
[0165] In Tables 2, 5, and 6, min_value[i] refers to the minimum value of the input signal (input image), and for transmission efficiency, it can be obtained by taking... The obtained value is transmitted. According to the implementation, when the value of min_value[i] is 0 and the minimum value before and after mapping is the same as 0, the corresponding value can be omitted. min_value[i] can be transmitted in the form of the difference between min_value[i-1] and the previously transmitted min_value[i-1] for transmission efficiency.
[0166] Basically, mapping relationships can be extracted for each region of interest (ROI) or for each ROI category, and the corresponding information can be transmitted to the decoder. According to the implementation, the coding efficiency of transmitting mapping relationships can be improved by using the following method: delivering a list of mapping relationships in units of sequences or GOPs for transmission efficiency, and transmitting only the index of the corresponding list within each ROI per frame.
[0167] In the non-region of interest (ROI) processing step 660, the processor 540 based on ROI features processes regions other than the ROI. The processor 540 based on ROI features performs the same processing for internal encoding by using information corresponding to the widest region in the ROI classification. According to the implementation, the method used in a specific classification within the ROI classification can be used for encoding efficiency, and the corresponding information can be delivered, which can be transmitted in frames.
[0168] The internal encoder 550 performs encoding on the mapped signal. The signal input to the internal encoder 550 is the chroma signal selected in the mapping step for encoding. Typically, the internal encoder 550 supports various formats of the input chroma signal; however, depending on the implementation, the number of chroma components encoded can be fixed to a specific number. For example, in chroma sampling processing, when only one chroma signal is retained and sampled, processing based on the region of interest features is performed, and the corresponding execution result is delivered to the internal encoder 550; the encoder may necessarily require three components. In this case, the signal sampled before the encoding step can be delivered as input to the internal encoder 550 by copying the same signal into three components to ensure that the required number of signals is met.
[0169] Figure 11 This is a block diagram of a decoding apparatus according to an embodiment of the present disclosure.
[0170] The decoding apparatus 10b according to embodiments of this disclosure can receive a bitstream as input, and generate and output a restored image through decoding. The decoding apparatus 10b may include an internal decoder 1110, a region of interest restorer 1120, a spatial restorer 1130, and a temporal restorer 1140. Figure 11 In this document, decoding device 10b represents the structure of a decoder for machine vision. Depending on the implementation, detailed configuration steps may be omitted and some steps may be modified. The decoder order is typically the reverse of the encoder order; however, when there is no correlation between the region of interest restorer 1120, the temporal restorer 1140, and the spatial restorer 1130, the decoder order may not correspond to the encoder order. However, when the information about the region of interest from the encoder is used in the spatial processor (spatial resampler 520 or region of interest feature-based processor 540) and the temporal processor (temporal resampler 510 or region of interest feature-based processor 540), and the corresponding information is used in the spatial restorer 1130 and the temporal restorer 1140, the decoder order may be the same as the encoder order. Figure 11 Same as above.
[0171] The internal decoder 1110 receives the bitstream as input and performs decoding. Although various decoders are used depending on the implementation, a decoder corresponding to the internal encoder used in the encoder is used.
[0172] The region of interest (ROI) restorer 1120 performs restoration by region of interest classification in the image signal decoded by the internal decoder 1110. Resolution restoration can be performed on the bitstream when the encoding device 10a, which encodes the input image, requests it. For example, restoration processing can be performed only when the color_restoration_flag in the transmitted bitstream is 1. The ROI restorer 1120 performs restoration by inverse mapping based on the position information of each ROI transmitted from the encoder, and performs restoration of sampled components by using a transmission method when the chroma components are additionally sampled.
[0173] Figure 12 The process of recovering the region of interest in a decoding apparatus according to an embodiment of this disclosure is illustrated.
[0174] The internal decoder 1110 can decode the bitstream and transmit it to the region of interest restorer 1120.
[0175] In the inverse mapping information extraction step 1201, the region of interest restorer 1120 can decode the information transmitted through the syntax defined in Tables 1 to 6 and obtain the inverse mapping relationship. The inverse mapping relationship is similar to... Figure 9 The mapping function in the encoder is related, and the inverse mapping relationship can be obtained by calculating the value corresponding to each segment using values such as delta_mapping_value and delta_value transmitted as differences. The region of interest restorer 1120 can decode the corresponding inverse mapping information of each region of interest transmitted from the encoder and obtain the inverse mapping relationship.
[0176] In step 1202 of extracting region of interest location information, the region of interest restorer 1120 can recover the region of interest location information through... Figures 1 to 6 The region of interest (ROI) location information in the transmitted syntax is decoded to obtain the position and size information of each ROI in the actual frame. The ROI restorer 1120 can apply the corresponding information to the image decoded by the internal decoder 1110 to specify the regions in which the restoration of the internally decoded image will be performed through inverse mapping.
[0177] In the region of interest inverse mapping step 1203, the region of interest restorer 1120 can perform inverse mapping by using inverse mapping information for the location of the region of interest in a specific image.
[0178] In component recovery step 1204, for an image decoded internally, when num_of_component is 1 and sensor_info is 0, i.e., when it is an image from an optical imaging device, the region of interest restorer 1120 can perform recovery on the chroma components. Recovery of the chroma components can be performed based on information transmitted via Tables 1 to 6. The image signal in which the chroma components are recovered in this manner can be delivered to the next step. For example, the image signal in which the chroma components are reconstructed can be delivered to the spatial restorer 1130 or the temporal restorer 1140.
[0179] Spatial restorer 1130 can perform spatial restoration using spatial restoration data transmitted from the encoder. Spatial restoration refers to the restoration of resolution per frame, and can have a modified resolution signal output to the signal input to the spatial restorer based on information transmitted from the encoder. Spatial restoration performed by spatial restorer 1130 refers to restoration using spatial restoration-related information transmitted from the encoder, and according to embodiments, the resolution restored by spatial restorer 1130 may be different from the resolution input to the spatial resampler 520 in the encoder.
[0180] The time restorer 1140 can perform time restoration using time restoration data transmitted from the encoder. Time restoration refers to the restoration of frames per second, and the output signal of the signal input to the time restorer 1140 can have a modified frame rate based on the information transmitted from the encoder. The time restoration performed by the time restorer 1140 refers to the restoration performed using time restoration-related information transmitted from the encoder, and according to the embodiment, the frame rate restored by the time restorer 1140 may be different from the frame rate input to the time resampler 510 in the encoder.
[0181] The output signal is the image signal from which the bitstream has been decoded to perform steps up to the time recovery stage, and filtering can be performed on the signal recovered during the spatial / temporal recovery process for reasons such as image quality degradation. Depending on the application of machine vision, the output signal can become the input to deep neural networks, etc.
[0182] The examples and figures presented in this specification are merely specific examples used to readily explain the technical content of this disclosure and to aid in understanding it, and are not intended to limit the scope of this specification. It will be apparent to those skilled in the art that other variations besides the examples described above may be feasible.
[0183] The claims set forth in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented in an apparatus, or the technical features of the apparatus claims in this specification can be combined and implemented in a method. Furthermore, the technical features of the method claims and the apparatus claims in this specification can be combined and implemented in an apparatus, or the technical features of the method claims and the apparatus claims in this specification can be combined and implemented in a method.
Claims
1. A machine-oriented video coding (VCM) encoding apparatus, the encoding apparatus comprising: A region of interest extractor, which is used to extract one or more regions of interest per frame of an input image; A processor based on region of interest features, wherein the processor is used to obtain the mapping relationship of the one or more regions of interest based on the features of the one or more regions of interest and to perform the mapping; as well as An internal encoder is used to generate a bitstream by encoding the input image using information about one or more regions of interest.
2. The apparatus according to claim 1, wherein: The region of interest extractor extracts one or more object regions from the input image and selects one or more regions of interest from the one or more object regions according to preset criteria.
3. The apparatus according to claim 2, wherein: The preset standard is set based on the size information of the object region, and the size information of the object region includes at least one of the width, height and area of the object region.
4. The apparatus according to claim 1, wherein: The region of interest extractor extracts one or more objects from the input image using a learning-based neural network model and selects regions of interest based on the detection accuracy of the extracted objects.
5. The apparatus according to claim 1, wherein: The features of the region of interest include at least one of the following: histogram of chromaticity components, mean, variance, object edges, and object classification.
6. The apparatus according to claim 5, wherein: The processor based on region of interest features classifies one or more regions of interest according to the similarity of their chromaticity components.
7. The apparatus according to claim 6, wherein: The processor based on region of interest features performs chroma sampling during the process of classifying the one or more regions of interest to remove some chroma signals from each region of interest.
8. The apparatus according to claim 6, wherein: The processor based on region of interest features determines the classification degree of one or more regions of interest according to the machine vision application purpose of the input image.
9. The apparatus according to claim 1, wherein: The processor based on region of interest features performs per-chroma component mapping for each region of interest.
10. The apparatus according to claim 1, wherein: The processor based on region of interest features adaptively sets the mapping segments by taking into account signal loss, according to the gradient of the segments before and after the mapping between each region of interest.
11. The apparatus according to claim 1, wherein: Extract the mapping relationship between one or more regions of interest according to the region of interest or the classification of the regions of interest.
12. The apparatus according to claim 11, wherein: The list of mappings between one or more regions of interest is included in the bitstream in units of sequences of the input images or groups of pictures (GOPs), and information about the regions of interest for each frame includes the index of the list of mappings.
13. The apparatus according to claim 1, wherein: The internal encoder encodes information about the one or more regions of interest and information about the mapping relationships between the one or more regions of interest.
14. The apparatus according to claim 1, wherein: The apparatus further includes a temporal resampler for changing the frame rate of the input image, and The temporal resampler performs upsampling or downsampling on a frame-by-frame basis based on the information about the one or more regions of interest.
15. The apparatus according to claim 1, wherein: The apparatus also includes a spatial resampler for changing the resolution of the input image, and The spatial resampler performs upsampling or downsampling on a frame-by-frame basis based on the information about the one or more regions of interest.
16. A machine-oriented video coding (VCM) decoding apparatus, the decoding apparatus comprising: An internal decoder is used to perform decoding on the bitstream to generate a restored image and extract information about the region of interest in the restored image; as well as A region of interest restorer, which is used to perform restoration on each or more regions of interest included in the restored image based on the information about the regions of interest.
17. The apparatus according to claim 16, wherein: The region of interest restorer performs inverse mapping on the one or more regions of interest based on the location information of each region of interest.
18. The apparatus according to claim 17, wherein: The region of interest restorer recovers the chromaticity components of the region of interest based on the information about the region of interest.
19. The apparatus according to claim 16, wherein: The device further includes: A spatial restorer, configured to restore the resolution of the restored image using spatial restoration data extracted from the bitstream; and A time restorer, used to restore the frame rate of the restored image by using time-restored data extracted from the bitstream. The spatial restorer performs spatial restoration based on the information about the region of interest, and The time restorer performs time restoration based on the information about the region of interest.
20. A non-volatile computer-readable storage medium for recording commands, wherein, When the command is executed by at least one processor, the at least one processor includes: Extract one or more regions of interest from frames of the input image; Based on the features of the one or more regions of interest, obtain the mapping relationship of the one or more regions of interest and perform the mapping; and A bitstream is generated by encoding the input image using information about one or more regions of interest.