Image data processing method and apparatus, recording medium storing bit stream, and bit stream transmission method
By decoding and post-processing image data, noise is removed and information is improved using predetermined optimization methods. This solves the problem that image compression is not suitable for artificial intelligence services in existing technologies, and achieves efficient image data processing and bitstream generation and storage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2023-10-06
- Publication Date
- 2026-05-01
Smart Images

Figure CN121970353A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an image data processing method and apparatus, and more specifically, to an image data processing method and apparatus optimized based on image data, a recording medium for storing a bitstream generated by the image data processing method / apparatus of this disclosure, and a bitstream transmission method. Background Technology
[0002] With the development of machine learning technology, the demand for image processing-based artificial intelligence services is constantly increasing. To effectively handle the large amounts of image data required for AI services within limited resources, optimizing image compression techniques for performing machine tasks is essential. However, existing image compression techniques, developed with the goal of high-resolution and high-quality image processing for human vision, are not suitable for AI services. Therefore, research and development of new machine-oriented image compression techniques suitable for AI services have been actively undertaken. Summary of the Invention
[0003] Technical issues
[0004] This disclosure aims to provide an image data processing method and apparatus with improved processing efficiency.
[0005] In addition, this disclosure aims to provide an image data processing method and apparatus for performing preprocessing / postprocessing on an input source based on a predetermined optimization method.
[0006] In addition, this disclosure aims to provide an image data processing method and apparatus based on a pre-determined interface with defined optimization information.
[0007] In addition, this disclosure aims to provide an image processing method and apparatus that can remove noise from the input source and improve information based on a predetermined optimization method.
[0008] In addition, this disclosure aims to provide a method for transmitting a bit stream generated by an image processing method or apparatus according to this disclosure.
[0009] In addition, this disclosure aims to provide a recording medium for storing a bitstream generated by an image processing method or apparatus according to this disclosure.
[0010] In addition, this disclosure aims to provide a recording medium for storing bitstreams that are received and decoded by the image processing method according to this disclosure and used for image reconstruction.
[0011] The technical problems to be solved by this disclosure are not limited to those described above. Other technical problems not described can be clearly understood by those skilled in the art through the following description.
[0012] Technical solution
[0013] An image data processing method according to one aspect of the present disclosure includes: decoding image data obtained from a bitstream; obtaining information related to an optimization method for the image data; and post-processing the decoded image data based on the information related to the optimization method, wherein the information related to the optimization method may include optimization attribute information indicating the type of at least one optimization method applied to the image data.
[0014] The receiving apparatus according to another aspect of this disclosure includes a memory and at least one processor, wherein the at least one processor decodes image data obtained from a bitstream, obtains information related to an optimization method for the image data, and performs post-processing on the decoded image data based on the information related to the optimization method, wherein the information related to the optimization method may include optimization attribute information indicating the type of at least one optimization method applied to the image data.
[0015] An image data processing method according to another aspect of this disclosure includes: determining an optimization method for image data; preprocessing the image data based on the optimization method; and encoding the preprocessed image data and information related to the optimization method, wherein the information related to the optimization method may include optimization attribute information indicating the type of at least one optimization method applied to the image data.
[0016] According to another aspect of this disclosure, the transmitting apparatus includes a memory and at least one processor, wherein the at least one processor determines an optimization method for image data, preprocesses the image data based on the optimization method, and encodes the preprocessed image data and information related to the optimization method, wherein the information related to the optimization method may include optimization attribute information indicating the type of at least one optimization method applied to the image data.
[0017] According to another aspect of this disclosure, the recording medium can store bit streams generated by the image processing method or transmitting apparatus of this disclosure.
[0018] According to another aspect of this disclosure, a bitstream transmission method can transmit a bitstream generated by the image processing method or transmitting apparatus of this disclosure to a receiving apparatus.
[0019] The features briefly outlined above for this disclosure are merely exemplary aspects of the detailed description of this disclosure that follows, and do not limit the scope of this disclosure.
[0020] Beneficial effects
[0021] According to this disclosure, an image data processing method and apparatus with improved processing efficiency can be provided.
[0022] In addition, according to this disclosure, an image data processing method and apparatus for performing preprocessing / postprocessing on an input source based on a predetermined optimization method can be provided.
[0023] In addition, according to this disclosure, an image data processing method and apparatus based on a pre-determined interface with defined optimization information can be provided.
[0024] In addition, according to this disclosure, an image processing method and apparatus capable of removing noise from the input source and improving information based on a predetermined optimization method can be provided.
[0025] Additionally, according to this disclosure, a method for transmitting a bit stream generated by an image processing method or apparatus according to this disclosure can be provided.
[0026] Additionally, according to this disclosure, a recording medium for storing bitstreams generated by an image processing method or apparatus according to this disclosure can be provided.
[0027] Additionally, according to this disclosure, a recording medium for storing bitstreams received and decoded by a receiving device according to this disclosure and used for image reconstruction can be provided.
[0028] The effects achievable by this disclosure are not limited to those described above, and other effects not described herein can be clearly understood by those skilled in the art from the following description. Attached Figure Description
[0029] Figure 1 This is a schematic diagram illustrating a VCM system to which embodiments of the present disclosure can be applied.
[0030] Figure 2 This is a schematic diagram illustrating a VCM pipeline structure to which embodiments of the present disclosure can be applied.
[0031] Figure 3 This is a schematic diagram illustrating an image / video encoder to which embodiments of the present disclosure may be applied.
[0032] Figure 4 This diagram schematically illustrates an image / video decoder to which embodiments of the present disclosure may be applied.
[0033] Figure 5 This is a flowchart illustrating a feature / feature map encoding process to which embodiments of the present disclosure can be applied.
[0034] Figure 6 This is a flowchart illustrating a feature / feature map decoding process to which embodiments of the present disclosure can be applied.
[0035] Figure 7This is a diagram illustrating an example of a VCM layer structure.
[0036] Figure 8 This is a diagram representing an example of a VCM bitstream consisting of encoded abstract features and NNAL information.
[0037] Figure 9 It is a diagram used to describe various problems in a VCM system.
[0038] Figure 10 and Figure 11 This is a diagram illustrating a VCM system according to an embodiment of the present disclosure.
[0039] Figure 12 and Figure 13 It is a diagram representing the hierarchical interface structure between the optimizer and the encoder.
[0040] Figure 14 This is an example of expressing the position and size of the ROI region within the input source in a pixel unit.
[0041] Figure 15 This is an example of expressing the location and size of the ROI region within the input source in a predefined cell.
[0042] Figure 16 This is a diagram illustrating the process of optimizing the input source and sending it to the encoder.
[0043] Figure 17 It means to Figure 16 A diagram illustrating the process of encoding optimized input sources.
[0044] Figure 18 This is a diagram illustrating the order of application of the optimization method according to embodiments of the present disclosure.
[0045] Figure 19 This is a diagram illustrating a VCM system according to an embodiment of the present disclosure.
[0046] Figure 20 This is a diagram used to describe an optimization setup process according to an embodiment of the present disclosure.
[0047] Figure 21 This is a flowchart illustrating an image data processing method according to an embodiment of the present disclosure.
[0048] Figure 22 This is a flowchart illustrating an image data processing method according to an embodiment of the present disclosure.
[0049] Figure 23 This is a diagram illustrating an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0050] Figure 24This is another example of a content streaming system to which embodiments of the present disclosure can be applied. Detailed Implementation
[0051] In the following description, embodiments of the present disclosure will be detailed with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, the present disclosure can be implemented in various different forms and is not limited to the embodiments described herein.
[0052] In describing embodiments of this disclosure, detailed explanations of well-known configurations or functions have been omitted where they are deemed to obscure the essential points of this disclosure. Additionally, portions irrelevant to the description of this disclosure have been omitted from the accompanying drawings, and similar reference numerals have been assigned to similar portions.
[0053] In this disclosure, when a particular component is described as “connected,” “coupled,” or “linked” to another component, this may include not only direct connections but also indirect connections in which another component may be present in between. Furthermore, when a component is described as “including” or “having” another component, this means that, unless otherwise expressly stated, it does not exclude other components but may further include additional components.
[0054] In this disclosure, unless otherwise expressly stated, the terms first, second, etc., are used only to distinguish one component from another and do not limit the order or importance of the components. Therefore, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0055] In this disclosure, distinguishable components are described to clearly explain their respective characteristics, and this does not necessarily mean that the components are separate. In other words, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed across multiple hardware or software units. Therefore, such integrated or distributed embodiments are also included within the scope of this disclosure unless they are explicitly described.
[0056] In this disclosure, the components described in the various embodiments are not necessarily essential components, and some components may be optional components. Therefore, embodiments comprising a subset of the components described in one embodiment are also included within the scope of this disclosure. Furthermore, embodiments that include additional components in addition to those described in the various embodiments are also included within the scope of this disclosure.
[0057] This disclosure relates to the encoding and decoding of images, and the terms used herein may have their common meanings as commonly used in the art to which this disclosure pertains, unless these terms are redefined in this disclosure.
[0058] This disclosure can be applied to methods disclosed in the Versatile Video Coding (VVC) standard and / or the Video Coding for Machines (VCM) standard. Additionally, this disclosure can be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second-generation Audio Video Coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268, etc.).
[0059] This disclosure presents various embodiments related to video / image coding, and unless otherwise stated, embodiments may be performed in combination with each other. In this disclosure, "video" can refer to a collection of images in chronological order. "Image" can be information generated by artificial intelligence (AI). Input information used by AI in performing a series of tasks, information generated during information processing, and output information can be used as images. In this disclosure, "picture" generally refers to a unit indicating a single image at a specific point in time, and a slice / tile is a coding unit that constitutes a part of a picture. A picture can consist of at least one slice / tile. Additionally, a slice / tile can include at least one coding tree unit (CTU). A CTU can be partitioned into at least one CU. A tile is a rectangular region existing within a specific tile row and a specific tile column within a picture, and can consist of multiple CTUs. A tile column can be defined as a rectangular region of a CTU and can have the same height as the picture and a width specified by syntax elements signaled from a bitstream portion (such as a picture parameter set). A tile row can be defined as a rectangular region of a CTU and can have the same width as the image and a height specified by a syntax element signaled from a bitstream portion (such as an image parameter set). Tile scanning is a method of predefined ordering of CTUs within a tile partition. Here, CTUs can be ordered according to the raster scan order of CTUs within a tile, and tiles within an image can be ordered sequentially according to the raster scan order of tiles in the image. A slice can include an integer number of complete tiles or an integer number of sequential complete CTU rows within a tile of an image. Slices can be specifically included in a single NAL unit. An image can consist of at least one tile group. A tile group can include at least one tile. A brick can indicate a rectangular region of a CTU row within a tile of an image. A tile can include at least one brick. A brick can indicate a rectangular region of a CTU row within a tile. A tile can be partitioned into multiple bricks, and each brick can include at least one CTU row belonging to the tile. A tile that is not partitioned into multiple bricks can also be considered a brick.
[0060] In this disclosure, a "pixel" or "cell" can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as the corresponding term for a pixel. A sample can generally indicate a pixel or a pixel value, and can indicate only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0061] In one embodiment, particularly when applied to VCM, when an image exists consisting of a set of components with different characteristics and meanings, the pixel / pixel value can indicate the pixel / pixel value of the component generated by combining, synthesizing, and analyzing the independent information of each component. For example, in RGB input, it can indicate only the pixel / pixel value of R, only the pixel / pixel value of G, or only the pixel / pixel value of B. For example, it can indicate only the pixel / pixel value of the luminance component synthesized using the R, G, and B components. For example, it can indicate only the pixel / pixel value of the information or image extracted by analyzing the R, G, and B components.
[0062] In this disclosure, "unit" can refer to a basic unit of image processing. A unit may include at least one of a specific region of an image or information associated with that region. A unit may include a luminance block and two chrominance (e.g., Cb, Cr) blocks. Depending on the context, the term "unit" may be used interchangeably with "sample array," "block," "region," etc. Typically, an M×N block may include a set (or array) of samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows. In one embodiment, particularly when applied to VCM, a unit may refer to a basic unit that includes information for performing a specific task.
[0063] In this disclosure, the term "current block" may refer to one of the following: "current coding block," "current coding unit," "encoding target block," "decoding target block," or "processing target block." When performing prediction, "current block" may refer to either the "current prediction block" or the "prediction target block." When performing transform (inverse transform) / quantization (dequantization), "current block" may refer to either the "current transform block" or the "transform target block." When performing filtering, "current block" may refer to the "filter target block."
[0064] Additionally, in this disclosure, unless explicitly stated as a chroma block, "current block" may refer to "the luminance block of the current block". "The chroma block of the current block" can be expressed by explicitly including an explicit description of a chroma block such as "chroma block" or "current chroma block".
[0065] In this disclosure, " / " and "," can mean "and / or". For example, "A / B" and "A, B" can mean "A and / or B". In addition, "A / B / C" and "A, B, C" can mean "at least one of A, B and / or C".
[0066] In this disclosure, "or" can mean "and / or". For example, "A or B" can mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in this disclosure, "or" can also mean "additionally or alternatively".
[0067] This disclosure relates to video / image coding (VCM) for machines.
[0068] VCM refers to compression techniques that encode / decode a portion of a source image / video or information for use in machine vision. In VCM, the encoding / decoding target can be called a feature. Features can refer to information extracted from the source image / video based on task objectives, requirements, and the surrounding environment. Features can have different information formats than the source image / video, and therefore, feature compression methods and representation formats can also differ from the video source.
[0069] VCMs can be applied to a wide range of applications. For example, in surveillance systems that identify and track objects or people, VCMs can be used to store or transmit object identification information. Additionally, in intelligent transportation or smart mobility systems, VCMs can be used to transmit vehicle location information collected from GPS, sensor information collected from LiDAR, radar, etc., and various vehicle control information to other vehicles or infrastructure. Furthermore, in the field of smart cities, VCMs can be used to perform individual tasks for interconnected sensor nodes or devices.
[0070] This disclosure provides various embodiments regarding feature / feature map encoding. Unless otherwise specifically stated, embodiments of this disclosure can be implemented individually or in combination of at least two.
[0071] Overview of VCM Systems
[0072] Figure 1 This is a schematic diagram illustrating a VCM system to which embodiments of the present disclosure can be applied.
[0073] refer to Figure 1 The VCM system may include an encoding device 10 and a decoding device 20.
[0074] Encoding device 10 can compress / encode features / feature maps extracted from source images / videos to generate a bitstream, and send the generated bitstream to decoding device 20 via a storage medium or network. Encoding device 10 can also be referred to as a feature encoding device. In a VCM system, features / feature maps can be generated in each hidden layer of a neural network. The size and number of channels of the generated feature map can vary depending on the type of neural network or the location of the hidden layers. In this disclosure, a feature map can be referred to as a feature set, and a feature or feature map can be referred to as "feature information".
[0075] The encoding device 10 may include a feature acquisition unit 9, an encoder 10, and a transmitter 13.
[0076] Feature acquirer 9 can obtain features / feature maps for the source image / video. According to an embodiment, feature acquirer 9 can obtain features / feature maps from an external device (e.g., a feature extraction network). In this case, feature acquirer 9 performs a feature receiving interface function. Alternatively, feature acquirer 9 can obtain features / feature maps by executing a neural network (e.g., CNN, DNN, etc.) using the source image / video as input. In this case, feature acquirer 9 performs a feature extraction network function.
[0077] According to an embodiment, the encoding device 10 may further include a source image generator (not shown) for obtaining source images / videos. The source image generator can be implemented using an image sensor, camera module, etc., and the source images / videos can be obtained through a process of capturing, synthesizing, or generating images / videos. In this case, the generated source images / videos can be sent to a feature extraction network and used as input data for extracting features / feature maps.
[0078] Encoder 10 can encode the features / feature maps obtained by feature acquirer 9. Encoder 10 can perform a series of processes, such as prediction, transformation, and quantization, to increase encoding efficiency. The encoded data (encoded feature / feature map information) can be output in the form of a bitstream. The bitstream including the encoded feature / feature map information can be referred to as the VCM bitstream.
[0079] Transmitter 13 can send feature / feature map information or data, output in bitstream form, to decoding device 20 via digital storage medium or network in the form of file or streaming transmission. Here, digital storage medium can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmitter 13 can include elements for generating media files with a predetermined file format or elements for transmitting data via broadcast / communication network.
[0080] The decoding device 20 can obtain feature / feature map information from the encoding device 10 and reconstruct the feature / feature map based on the obtained information.
[0081] The decoding device 20 may include a receiver 21 and a decoder 22.
[0082] Receiver 21 can receive bit streams from encoding device 10 and obtain feature / feature map information from the received bit streams to send to decoder 22.
[0083] Decoder 22 can decode features / feature maps based on the obtained feature / feature map information. Decoder 22 can perform a series of processes corresponding to the operations of encoder 14, such as dequantization, inverse transform, prediction, etc., to increase decoding efficiency.
[0084] According to an embodiment, the decoding device 20 may further include a task analysis / rendering unit 23.
[0085] The task analysis / rendering unit 23 can perform task analysis based on decoded features / feature maps. Furthermore, the task analysis / rendering unit 23 can render decoded features / feature maps into a form suitable for task execution. Based on the task analysis results and the rendered features / feature maps, various (machine-oriented) tasks can be executed.
[0086] Therefore, a VCM system can encode / decode features extracted from source images / videos based on user and / or machine requests, task objectives, and the surrounding environment, and perform various (machine-oriented) tasks based on the decoded features. A VCM system can also be implemented by extending / redesigning the video / image coding system and can execute various encoding / decoding methods defined in the VCM standard.
[0087] VCM production line
[0088] Figure 2 This is a schematic diagram illustrating a VCM pipeline structure to which embodiments of the present disclosure can be applied.
[0089] refer to Figure 2 The VCM pipeline 200 may include a first pipeline 210 for encoding / decoding images / videos and a second pipeline 220 for encoding / decoding features / feature maps. In this disclosure, the first pipeline 210 may be referred to as a video codec pipeline, and the second pipeline 220 may be referred to as a feature codec pipeline.
[0090] The first pipeline 210 may include a first stage 211 for encoding the input image / video and a second stage 212 for decoding the encoded image / video to generate a reconstructed image / video. The reconstructed image / video can be used for human viewing, i.e., human vision.
[0091] The second pipeline 220 may include a third stage 221 for extracting features / feature maps from the input image / video, a fourth stage 222 for encoding the extracted features / feature maps, and a fifth stage 223 for decoding the encoded features / feature maps to generate reconstructed features / feature maps. The reconstructed features / feature maps can be used for machine (vision) tasks. Here, a machine (vision) task can refer to a task where a machine consumes images / videos. Machine vision tasks can be applied to service scenarios such as, for example, surveillance, intelligent transportation, smart cities, smart industries, and smart content. According to one embodiment, the reconstructed features / feature maps can also be used for human vision.
[0092] According to an embodiment, the features / feature maps encoded in the fourth stage 222 can be sent to the first stage 221 and used to encode the image / video. In this case, an additional bitstream can be generated based on the encoded features / feature maps, and the generated additional bitstream can be sent to the second stage 222 and used to decode the image / video.
[0093] According to an embodiment, the features / feature maps decoded in the fifth stage 223 can be sent to the second stage 222 and used to decode the image / video.
[0094] although Figure 2 The illustration shows a VCM pipeline 200 including a first pipeline 210 and a second pipeline 220, but this is merely exemplary, and embodiments of this disclosure are not limited thereto. For example, the VCM pipeline 200 may include only the second pipeline 220, or the second pipeline 220 may be extended to multiple feature codec pipelines.
[0095] Meanwhile, in the first pipeline 210, the first stage 211 can be executed by an image / video encoder, and the second stage 212 can be executed by an image / video decoder. Additionally, in the second pipeline 220, the third stage 221 can be executed by a VCM encoder (or a feature / feature map encoder), and the fourth stage 222 can be executed by a VCM decoder (or a feature / feature map decoder). The encoder / decoder structure is described in detail below.
[0096] encoder
[0097] Figure 3This is a schematic diagram illustrating an image / video encoder to which embodiments of the present disclosure may be applied.
[0098] refer to Figure 3 The image / video encoder 300 may include an image partitioner 310, a predictor 320, a residual processor 330, an entropy encoder 340, an adder 350, a filter 360, and a memory 370. The predictor 320 may include an inter-frame predictor 321 and an intra-frame predictor 322. The residual processor 330 may include a transformer 332, a quantizer 333, a dequantizer 334, and an inverse transformer 335. The residual processor 330 may also include a subtractor 331. The adder 350 may be referred to as a reconstructor or a reconstruction block generator. According to one embodiment, the image partitioner 310, predictor 320, residual processor 330, entropy encoder 340, adder 350, and filter 360 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 370 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The aforementioned hardware component may further include the memory 370 as an internal / external component.
[0099] Image partitioner 310 can partition an input image (or picture, frame) input to image / video encoder 300 into at least one processing unit. As an example, a processing unit can be referred to as a coding unit (CU). Coding units can be recursively partitioned from coding tree units (CTUs) or largest coding units (LCUs) according to a quadtree-binary-trinary tree (QTBTTT) structure. For example, a coding unit can be partitioned into multiple coding units with greater depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, a quadtree structure can be applied first, and a binary tree structure and / or a ternary structure can be applied later. Alternatively, a binary tree structure can be applied first. The image / video encoding process according to this disclosure can be performed based on the final coding unit, which is no longer partitioned. In this scenario, the largest coding unit can be used as the final coding unit based on factors such as coding efficiency according to image features, or, if necessary, the coding unit can be recursively divided into deeper coding units, and the optimal-sized coding unit can be used as the final coding unit. Here, the coding process can include processes such as prediction, transformation, and reconstruction, as described later. As another example, the processing unit can further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can be partitioned or divided from the final coding unit described above, respectively. The prediction unit can be a unit for sample prediction, and the transformation unit can be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.
[0100] In some cases, a unit can be used interchangeably with terms such as block, region, etc. Typically, an MxN block can refer to a set of transform coefficients or a set of samples consisting of M columns and N rows. Samples can typically indicate pixels or pixel values, or they can indicate only pixel / pixel values of the luminance component, or only pixel / pixel values of the chrominance component. Samples can be used as items corresponding to pixels or cells.
[0101] The image / video encoder 300 generates a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 321 or intra-frame predictor 322 from the input image signal (original block, original sample array), and the generated residual signal is sent to the converter 332. In this case, as shown, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within the image / video encoder 300 can be called the subtractor 331. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block that includes the prediction samples of the current block. The predictor can determine whether intra-frame prediction or inter-frame prediction is applied to the current block or the unit of the CU. The predictor can generate various prediction-related information, such as prediction mode information, and send it to the entropy encoder 340. The prediction-related information can be encoded by the entropy encoder 340 and output in the form of a bitstream.
[0102] Intra-predictor 322 can predict the current block by referencing samples within the current image. In this case, the reference samples can be located in a neighboring region of the current block, or they can be located further away depending on the prediction mode. In intra-prediction, the prediction modes can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC modes and planar modes. Directional modes can include, for example, 33 or 65 directional prediction modes depending on the granularity of the prediction orientation. However, this is just an example, and more or fewer directional prediction modes can be used depending on the configuration. Intra-predictor 322 can also determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.
[0103] Inter-frame predictor 321 can derive the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted at the block, sub-block, or sample level based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may further include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing within the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a co-located reference block or co-located CU (colCU), and the reference image including the temporally neighboring block may be referred to as a co-located image (colPic). For example, inter-frame predictor 321 can construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes, and for example, in skip mode and merge mode, the inter-frame predictor 321 can use motion information of neighboring blocks as motion information of the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors of neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0104] Predictor 320 can generate prediction signals based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction for the prediction of a block, and can also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined inter-frame and intra-frame prediction (CIIP). Alternatively, the predictor can be based on an intra-block copy (IBC) prediction mode or a palette mode for the prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding, such as screen content coding (SCC). IBC essentially performs prediction within the current frame, but because it derives a reference block within the current frame, it can be similar to inter-frame prediction operation. In other words, IBC can use at least one of the inter-frame prediction methods described in this disclosure. The palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, sample values within the frame can be signaled based on information associated with the palette table and palette index.
[0105] The predicted signal generated by predictor 320 can be used to generate a reconstructed signal or a residual signal. Transformer 332 can generate transform coefficients by applying a transform method to the residual signal. For example, the transform method can include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), Graphical Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to the transform obtained from a graph when the relationship information between pixels is indicated as a graph. CNT refers to the transform obtained based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform process can be applied to pixel blocks of the same square size or non-square, variable-size blocks.
[0106] The quantizer 333 quantizes the transform coefficients and sends them to the entropy encoder 340, which encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. This information about the quantized transform coefficients can be referred to as residual information. The quantizer 333 can reorder the block-like quantized transform coefficients in a one-dimensional vector form based on the coefficient scan order, and can generate information about the quantized transform coefficients based on this one-dimensional vector form. The entropy encoder 340 can perform various encoding methods, such as Exponential Columbus, Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), etc. The entropy encoder 340 can encode the quantized transform coefficients together with or separately from information necessary for video / image reconstruction (e.g., values of syntax elements). The encoded information (e.g., encoded video / image information) can be sent or stored as a bitstream in a Network Abstraction Layer (NAL) unit. The image / video information may further include information about various parameter sets, such as Adaptation Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may further include general constraint information. Furthermore, the image / video information may further include methods for generating and using encoded information, their purpose, etc. In this disclosure, information and / or syntax elements sent / signed from the image / video encoder to the image / video decoder may be included in the image / video information. The image / video information may be encoded using the encoding process described above and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmission and / or a storage unit (not shown) for storing the signal output from the entropy encoder 340 may be configured as an internal / external element of the image / video encoder 300, or the transmitter may be included in the entropy encoder 340.
[0107] The quantization transform coefficients output from quantizer 333 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantization transform coefficients through dequantizer 334 and inverse transform 335. Adder 350 can generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from inter-frame predictor 321 or intra-frame predictor 322. When there is no residual for the processing target block, such as when a skip mode is applied, the prediction block can be used as a reconstructed block. Adder 350 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block within the current image, and, as described later, can also be used for inter-frame prediction of the next image by filtering.
[0108] Simultaneously, luminance mapping with chroma scaling can be applied during image encoding and / or reconstruction.
[0109] Filter 360 can apply filtering to the reconstructed signal to enhance subjective / objective image quality. For example, filter 360 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and the modified reconstructed image can be stored in memory 370, specifically in the DPB of memory 370. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. Filter 360 can generate various filtering-related information and send it to entropy encoder 340. The filtering-related information can be encoded by entropy encoder 340 and output as a bitstream.
[0110] The modified reconstructed image sent to memory 370 can be used as a reference image in inter-frame predictor 321. This avoids prediction mismatch between the encoder and decoder sides and improves coding efficiency.
[0111] The DPB of memory 370 can store modified reconstructed images for use as reference images in inter-frame predictor 321. Memory 370 can store motion information of blocks in which motion information within the current image and / or motion information of blocks within reconstructed images is derived (or encoded). The stored motion information can be sent to inter-frame predictor 321 as motion information for spatially or temporally neighboring blocks. Memory 370 can store reconstructed samples of reconstructed blocks in the current image and send the stored reconstructed samples to intra-frame predictor 322.
[0112] At the same time, the VCM encoder (or feature / feature map encoder) can have essentially the same reference. Figure 3The image / video encoder 300 described has the same / similar structure as the image / video encoder 300 because it performs a series of processes such as prediction, transformation, quantization, etc., to encode features / feature maps. However, the VCM encoder differs from the image / video encoder 300 because it targets the features / feature maps to be encoded, and therefore, its name in each unit (or component) (e.g., image partitioner 310, etc.) and its specific operational details derived from the image / video encoder 300 may differ. The specific operational details of the VCM encoder will be described in detail later.
[0113] decoder
[0114] Figure 4 This diagram schematically illustrates an image / video decoder to which embodiments of the present disclosure may be applied.
[0115] refer to Figure 4 The image / video decoder 400 may include an entropy decoder 410, a residual processor 420, a predictor 430, an adder 440, a filter 450, and a memory 460. The predictor 430 may include an inter-frame predictor 431 and an intra-frame predictor 432. The residual processor 420 may include a dequantizer 421 and an inverse transformer 422. According to one embodiment, the entropy decoder 410, residual processor 420, predictor 430, adder 440, and filter 450 may be configured by a single hardware component (e.g., a decoder chipset or processor). Additionally, the memory 460 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may also include the memory 460 as an internal / external component.
[0116] When the input includes a bitstream containing video / image information, the image / video decoder 400 can respond to... Figure 3 The image / video encoder 300 reconstructs the image / video through a process of processing image / video information. For example, the image / video decoder 400 can derive units / blocks based on block partitioning information obtained from the bitstream. The image / video decoder 400 can perform decoding using processing units applied in the image / video encoder. Therefore, the decoding processing unit can be, for example, an encoding unit, and the encoding unit can be partitioned according to a quadtree structure, binary tree structure, and / or ternary tree structure from the encoding tree unit or the maximum encoding unit. At least one transform unit can be derived from the encoding unit. Furthermore, the reconstructed image signal decoded and output by the image / video decoder 400 can be played back by a playback device.
[0117] The image / video decoder 400 is capable of transmitting data from a bitstream source. Figure 3The encoder in the process receives the signal output and can decode the received signal using the entropy decoder 410. For example, the entropy decoder 410 can parse the bitstream to derive the information necessary for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as adaptation parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), video parameter sets (VPS), etc. Additionally, the video / image information may further include general constraint information. Furthermore, the image / video information may include the method of generating the decoding information, its usage, purpose, etc. The image / video decoder 400 can further decode the picture based on information about the parameter sets and / or general constraint information. The signal notification / received information and / or syntax elements can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 410 can decode the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and can output the values of the syntax elements necessary for image reconstruction and the quantized values of the transform coefficients associated with the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model by using information about the target syntax element, the decoding information of neighboring and target blocks, or information about symbols / bins decoded in previous steps, predict the probability of bin occurrence based on the determined context model, and perform arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. In this case, the CABAC entropy decoding method can update the context model after determining it by using information about decoded symbols / bins of the context model for the next symbol / bin. In the information decoded by the entropy decoder 410, prediction-related information can be provided to the predictors (inter-frame predictor 432 and intra-frame predictor 431), and the residual values (i.e., quantization transform coefficients and related parameter information) entropied by the entropy decoder 410 can be input to the residual processor 420. The residual processor 420 can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, in the information decoded by the entropy decoder 410, filtering-related information can be provided to the filter 450. Meanwhile, the receiver (not shown) that receives the signal output from the image / video encoder can be additionally configured as an internal / external component of the image / video decoder 400, or the receiver can be a component of the entropy decoder 410. Furthermore, the image / video decoder according to this disclosure can also be referred to as an image / video decoding device, and the image / video decoder can be divided into an information decoder (image / video information decoder) and / or a sample decoder (image / video sample decoder).In this case, the information decoder may include an entropy decoder 410, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 440, a filter 450, a memory 460, an inter-frame predictor 432, and an intra-frame predictor 431.
[0118] Dequantizer 421 can dequantize the quantized transform coefficients and the output transform coefficients. Dequantizer 421 can reorder the quantized transform coefficients in the form of two-dimensional blocks. In this case, the reordering can be performed based on the coefficient scan order performed in the image / video encoder. Dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (i.e., quantization step size information) and obtain the transform coefficients.
[0119] The inverse transformer 422 can perform an inverse transformation on the transform coefficients to obtain the residual signal (residual block, residual sample array).
[0120] Predictor 430 can perform prediction for the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether intra-frame prediction or inter-frame prediction is applied to the current block based on prediction-related information output from entropy decoder 410, and can determine a specific intra-frame / inter-frame prediction mode (prediction method).
[0121] Predictor 420 can generate prediction signals based on various prediction methods. For example, the predictor can apply not only intra-frame prediction or inter-frame prediction, but also both intra-frame and inter-frame prediction simultaneously for the prediction of a block. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Alternatively, the predictor can be based on an intra-block copy (IBC) prediction mode or a palette mode for block prediction. The IBC prediction mode or palette mode can be used for content image / video encoding in games, such as screen content encoding (SCC). IBC essentially performs prediction within the current frame, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described in this document. The palette mode can be considered an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, information related to the palette table and palette index can be included in the image / video information and signaled.
[0122] The intra-predictor 431 can predict the current block by referencing samples within the current image. The reference samples can be located in the neighborhood of the current block or positioned far from the current block based on the prediction pattern. In intra-prediction, the prediction pattern can include multiple non-directional patterns and multiple directional patterns. The intra-predictor 431 can determine the prediction pattern applied to the current block by using prediction patterns applied to neighboring blocks.
[0123] Inter-frame predictor 432 can derive the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted at the block, sub-block, or sample level based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may further include information about the inter-frame prediction orientation (i.e., L0 prediction, L1 prediction, Bi prediction, etc.). In inter-frame prediction, neighboring blocks may include spatially neighboring blocks within the current image and temporally neighboring blocks in the reference image. For example, inter-frame predictor 432 can construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference image index of the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and prediction-related information may include information indicating the inter-frame prediction mode of the current block.
[0124] Adder 440 can generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictors (including inter-frame predictor 432 and / or intra-frame predictor 431). When there is no residual for the processing target block, such as when a skip mode is applied, the prediction block can be used as the reconstruction block.
[0125] Adder 440 can be referred to as a reconstructor or reconstruction block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current image, or, as described later, can be filtered out, or used for inter-frame prediction of the next image.
[0126] Simultaneously, luminance mapping with chroma scaling can be applied during image decoding.
[0127] Filter 450 can apply filtering to the reconstructed signal to enhance subjective / objective image quality. For example, filter 450 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and can send the modified reconstructed image to memory 460, specifically to the DPB in memory 460. Various filtering methods may include, for example, unblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.
[0128] The (modified) reconstructed image stored in the DPB of memory 460 can be used as a reference image in inter-frame predictor 432. Memory 460 can store motion information of blocks in which motion information within the current image is derived (or decoded) and / or the motion information of blocks in the reconstructed image is sent to inter-frame predictor 432 as motion information for spatially or temporally neighboring blocks. Memory 460 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 431.
[0129] Meanwhile, the VCM decoder (or feature / feature map decoder) can have the same features as described above (see reference). Figure 4 The image / video decoder 400 described below has the same / similar structure as the image / video decoder 400 because it performs a series of processes, such as prediction, inverse transform, dequantization, etc., to decode features / feature maps. However, the VCM decoder differs from the image / video decoder 400 because it targets the features / feature maps used for decoding, and therefore, its name in each unit (or component) (e.g., DPB, etc.) and its specific operational details from those of the image / video decoder 400 may differ. The operations of the VCM decoder can correspond to the operations of the VCM encoder, and their specific operational details will be described in detail later.
[0130] Feature / feature map encoding process
[0131] Figure 5 This is a flowchart illustrating a feature / feature map encoding process to which embodiments of the present disclosure can be applied.
[0132] refer to Figure 5 The feature / feature map encoding process may include a prediction process S510, a residual processing process S520, and an information encoding process S530.
[0133] The prediction process S510 can be referenced above. Figure 3 The predictor 320 is described and executed.
[0134] Specifically, the intra predictor 322 can predict the current block (i.e., the currently encoded set of feature elements) by referencing feature elements in the current feature / feature map. Intra-prediction can be performed based on the spatial similarity of feature elements in the configured feature / feature map. For example, it can be estimated that feature elements included in the same region of interest (RoI) within an image / video have similar data distribution characteristics. Therefore, the intra predictor 322 can predict the current block by referencing pre-reconstructed feature elements in the RoI that includes the current block. In this case, the referenced feature elements can be located near the current block or can be located separately from the current block depending on the prediction mode. The intra-prediction modes used for feature / feature map encoding can include multiple non-directional prediction modes and multiple directional prediction modes. Non-directional prediction modes can include, for example, prediction modes corresponding to DC mode and planar mode in the image / video encoding process. In addition, directional modes can include, for example, prediction modes corresponding to 33 directional modes or 65 directional modes in the image / video encoding process. However, this is only an example, and according to one embodiment, the type and number of intra-prediction modes can be configured / changed in various ways.
[0135] Inter-frame predictor 321 can predict the current block based on a reference block (i.e., a set of reference feature elements) specified by motion information on a reference feature / feature map. Inter-frame prediction can be performed based on the temporal similarity of feature elements in the configuration feature / feature map. For example, temporally consecutive features may have similar data distribution characteristics. Therefore, inter-frame predictor 321 can predict the current block by referencing pre-reconstructed feature elements of the current feature and temporally adjacent features. In this case, the motion information used to specify the reference feature elements may include motion vectors and reference feature / feature map indices. The motion information may further include information related to inter-frame prediction orientation (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current feature / feature map and temporally neighboring blocks existing in the reference feature / feature map. The reference feature / feature map including the reference block and the reference feature / feature map including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a co-located reference block, etc., and the reference feature / feature map including the temporally neighboring block may be referred to as a co-located feature / feature map. Inter-frame predictor 321 can configure a candidate list of motion information based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference feature / feature map index of the current block. Inter-frame prediction can be performed based on various prediction modes, and for example, for skip mode and merge mode, inter-frame predictor 321 can use the motion information of neighboring blocks as the motion information of the current block. For skip mode, unlike merge mode, residual signals may not be sent. For motion vector prediction (MVP) mode, motion vectors of neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference. In addition to the intra-frame prediction and inter-frame prediction described above, predictor 320 can also generate prediction signals based on various prediction methods.
[0136] The prediction signal generated by predictor 320 can be used to generate the residual signal (residual block, residual eigenvalue) S520. This can be seen from the above reference... Figure 3 The described residual processor 330 executes the residual processing procedure S520. Furthermore, the (quantization) transform coefficients can be generated through a transform and / or quantization process for the residual signal, and the entropy encoder 340 can encode the information associated with the (quantization) transform coefficients into residual information in the bitstream S530. In addition to the residual information in the bitstream, the entropy encoder 340 can also encode information necessary for feature / feature map reconstruction, such as prediction information (e.g., prediction pattern information, motion information, etc.).
[0137] Meanwhile, the feature / feature map encoding process may further include a process for generating a reconstructed feature / feature map for the current feature / feature map, and a process for applying in-loop filtering to the reconstructed feature / feature map (optional), as well as a process S530 for encoding information for feature / feature map reconstruction (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in the form of a bit stream.
[0138] The VCM encoder can derive (modified) residual features from the quantization transform coefficients through dequantization and inverse transform, and can generate reconstructed features / feature maps based on the predicted features as the output of the S510 and the (modified) residual features. The reconstructed features / feature maps generated in this way can be identical to those generated by the VCM decoder. When an in-loop filtering process is performed on the reconstructed features / feature maps, a modified reconstructed features / feature maps can be generated through this in-loop filtering process. The modified reconstructed features / feature maps can be stored in the decoded feature buffer (DFB) or memory, and then used as reference features / feature maps during feature / feature map prediction. Furthermore, (in-loop) filtering-related information (parameters) can be encoded and output as a bitstream. The in-loop filtering process removes noise that may occur during feature / feature map encoding and improves the performance of feature / feature map-based tasks. Additionally, the in-loop filtering process can be performed on both the encoder and decoder sides to ensure the recognition of prediction results, improve the reliability of feature / feature map encoding, and reduce the amount of data transmitted for feature / feature map encoding.
[0139] Feature / Feature Map Decoding Process
[0140] Figure 6 This is a flowchart illustrating a feature / feature map decoding process to which embodiments of the present disclosure can be applied.
[0141] refer to Figure 6The feature / feature map decoding process may include an image / video information acquisition process S610, feature / feature map reconstruction processes S620 to S640, and an intra-loop filtering process S650 for reconstructing the feature / feature map. The feature / feature map reconstruction process can be performed based on the prediction signal and residual signal obtained through inter-frame / intra-frame prediction S620, residual processing S630, and dequantization and inverse transform processes for the quantization transform coefficients described in this disclosure. A modified reconstructed feature / feature map can be generated through the intra-loop filtering process for reconstructing the feature / feature map, and the modified reconstructed feature / feature map can be output as a decoded feature / feature map. The decoded feature / feature map can be stored in a decoded feature buffer (DFB) or memory, and then used as a reference feature / feature map in the inter-frame prediction process when decoding the feature / feature map. In some cases, the above-described intra-loop filtering process may be omitted. In this case, the reconstructed features / feature maps can be output as decoded features / feature maps and stored in the decoded feature buffer (DFB) or memory, and then used as reference features / feature maps in the inter-frame prediction process when decoding the features / feature maps.
[0142] VCM coding layer and structure
[0143] Figure 7 This is a diagram illustrating an example of a VCM layered structure.
[0144] refer to Figure 7 The VCM layered structure may include a feature extraction layer 1310, a neural network (feature) abstraction layer 1320, and a feature encoding layer 1330.
[0145] The feature extraction layer 710 can refer to the layer used to extract features from the input source, and may also include the extracted results. The feature encoding layer 730 can refer to the layer used to compress the extracted features, and may also include the compressed results.
[0146] The neural network abstraction layer 720 can abstract the information generated from the feature extraction layer 710 (e.g., information about the extracted features / feature maps) and send it to the feature encoding layer 730. The neural network abstraction layer 720 can hide the internal structure of the feature extraction layer 710 and provide consistent feature interface functionality through information abstraction. Therefore, even when the compression target changes due to variations in tools (e.g., CNN, DNN, etc.), the feature encoding layer 730 can perform a consistent feature encoding process. In this disclosure, the neural network abstraction layer (NNAL) may also be referred to as the feature abstraction layer.
[0147] The interfaces between the feature extraction layer 710 and the neural network abstraction layer 720, as well as between the feature encoding layer 730 and the neural network abstraction layer 720, can be predefined, and the operations in the neural network abstraction layer 720 can be configured to be modified later.
[0148] Figure 8 A diagram illustrating an example of a VCM bitstream consisting of encoded abstract features and NNAL information.
[0149] like Figure 8 The bitstream configured as shown can be referred to as a Neural Network Abstraction Layer (NNAL) unit. An NNAL unit can be an independent feature reconstruction unit. The input features for a single NNAL unit can be extracted from the same layer within the neural network. Therefore, the input features for a single NNAL unit may be forced to have the same characteristics. For example, the same feature extraction method can be applied to the input features of a single NNAL unit.
[0150] An NNAL unit may include an NNAL unit header and an NNAL unit payload. The NNAL unit header may include all the information needed to perform the task using the encoded features. The NNAL unit payload may include abstract feature information. The NNAL unit payload may include a group header and group data. The group header may include configuration information for the feature group data, such as the temporal order, number, or common attributes of the feature channels constituting the feature group. A feature channel may refer to a unit that encodes a feature. The group data may include multiple feature channels and encoding indicators, and each feature channel may include type information, prediction information, auxiliary information, and residual information. In this case, the type information may indicate the encoding method, and the prediction information may indicate the prediction method. Furthermore, the auxiliary information may indicate additional information required for decoding (e.g., entropy coding, quantization-related information, etc.), and the residual information may include information about the encoded feature elements (i.e., a set of feature value information).
[0151] Figure 9 It is a diagram used to describe various problems in a VCM system.
[0152] refer to Figure 9The VCM system 900 may include a first device 910 and a second device 920. The first device 910 is a transmitting-side device and may include a source device 911 and an encoder 913. The source device 911 can acquire image information through processes such as image capture and synthesis. The encoder 913 can encode the image information acquired by the source device 911 to generate encoded image data (encoded image / feature information) and output it as a bitstream. The second device 920 is a receiving-side device and may include a decoder 921 and a task network 923. The decoder 921 can decode the data acquired from the first device 910 according to the receiving-side requirements (e.g., bandwidth, frame rate) to obtain image / feature information. The task network 923 can use the acquired image / feature information to perform predetermined tasks (e.g., machine tasks, human vision tasks).
[0153] Predetermined context information can be input to the source device 911. Here, the context information is information related to the environment when image information is acquired, and may include information related to the device environment (e.g., location, time) or task attributes. Due to environmental factors (e.g., weather, lighting, background), unwanted noise may be added to the context information. Additionally, due to the characteristics of the source device 911 (e.g., camera parameters, sensor noise), unwanted noise may also be added to the image information acquired based on the context information. Therefore, a preprocessing (or optimization) process is required to remove the aforementioned noise.
[0154] Furthermore, the information required for task execution can vary depending on the task's attributes or purpose. For example, the attributes of the information required for machine tasks and human vision tasks can differ. Additionally, the attributes of the required information can also vary depending on the type of machine task. Therefore, adaptive decoding or post-processing procedures that consider task attributes and purpose are necessary.
[0155] Therefore, this disclosure aims to provide various embodiments of optimization techniques and their expression methods for solving the above-mentioned problems. Embodiments of this disclosure can be implemented individually or in combination of two or more. Hereinafter, embodiments of this disclosure will be described in detail with reference to the accompanying drawings.
[0156] Example 1
[0157] Embodiment 1 of this disclosure relates to a method for eliminating noise from an input source and a method for improving information to be suitable for a specific task.
[0158] Figure 10 and Figure 11 This is a diagram illustrating a VCM system according to an embodiment of the present disclosure.
[0159] First, refer to Figure 10 The VCM system 1000 may include a first device 1010 and a second device 1020. The first device 1010 is a transmitting-side device, which may refer to a component that provides information (e.g., encodes and transmits image information), and the second device 1020 is a receiving-side device, which may refer to a component that consumes information (e.g., decodes and utilizes image information).
[0160] The first device 1010 may include a source device 1011 and an encoder 1013. Additionally, the first device 1010 may further include an optimizer 1012 for removing noise from the image information (or input source) obtained by the source device 1011. In this disclosure, the optimizer 1012 may be referred to by various terms such as preprocessor, noise remover, etc.
[0161] Optimizer 1012 can perform predetermined preprocessing operations (or optimizations) on image information to remove various noises added to the image information. Additionally, optimizer 1012 can improve the information required for task execution through preprocessing operations. The preprocessing method, i.e., the optimization method, needs to be adaptively defined by considering the task objective to improve task efficiency and accuracy. For this purpose, second device 1020 can request (B) first device 1010 to apply an appropriate optimization method, and first device 1010 can send (A) information related to the optimization method (e.g., whether an optimization method is applied, the type of optimization, and the purpose, etc.) (hereinafter, optimization information) to second device 1020. In this example, optimization information can be sent to second device 1020 via a bitstream. (See reference...) Figure 10 As described above, the bitstream can be configured in a unit for transmission (e.g., an NNAL unit), and the attribute information in the transmission unit (e.g., the NNAL unit type) can be used to transmit optimization information.
[0162] Simultaneously, optimization information can be generated by optimizer 1012 and sent to encoder 1013. For this purpose, an interface (INF) between optimizer 1012 and encoder 1013 can be defined. The interface (INF) can consist of general information about the input source and optimization information. Here, the general information about the input source is necessary for the basic operation of encoder 1013 and can include, for example, input bit depth (InputBitDepth), luminance to chrominance sample ratio (InputFormat), frame rate (FrameRete), input source width (SourceWidth), and source height (SourceHeight).
[0163] According to an embodiment, the first device 1010 described above can be expanded into multiple devices, and... Figure 10The difference is illustrated in the diagram. In other words, the second device 1020 can communicate with multiple first devices. In this case, the multiple first devices can be interconnected to form a single logical entity. For example, the multiple first devices can be a collection of CCTVs installed over a wide area, and the second device 1020 can be a central control device that receives CCTV images from the multiple first devices.
[0164] Because the first device 1010 is expanded into multiple devices, the number of decoders and task networks included in the second device 1020 can also be expanded to multiple decoders and task networks. In other words, the second device 1020 can also be a single logical entity in which multiple second devices are interconnected. In this case, as... Figure 11 As illustrated, the second device 1120 may further include an optimization manager 1123, which controls the decoder 1121 and the task network 1122 and manages optimization information. The optimization manager 1123 can receive multiple output signals (A) from the decoder 1121 and the task network 1122, request optimization information corresponding to each signal from the first device 1110, and send a request (B) on the receiving side. Figure 10 In this case, the information transmitted between the first device 1110 and the second device 1120 may be all or part of the same as (B).
[0165] Meanwhile, optimization information refers to the information required for preprocessing (or optimization) operations on the input source, and can be such as Figure 12 and Figure 13 The diagram illustrates the step-by-step (or hierarchical) definition.
[0166] First, refer to Figure 12 The first-level interface can define whether optimization is applied to the input source and the scope of the application. The optimization attributes and types, or optimization purposes, can be defined in the next second-level interface. Additional information based on the optimization attributes and types can be defined in the final third-level interface. Depending on the embodiment, only some steps or all steps can be defined. For example, only the first-level interface can be defined, the first and second-level interfaces can be defined, or all of the first to third-level interfaces can be defined.
[0167] Table 1 below shows examples where only the first-level interface is defined.
[0168] [Table 1]
[0169] Referring to Table 1, optimization_info() can include the syntax elements optimization_enable_flag, level_2_information_enable_flag, and level_2_information.
[0170] The syntax element `optimization_enable_flag` can indicate whether optimization is applied to the input source. For example, a first value (e.g., 1) for `optimization_enable_flag` indicates that optimization is applied to the input source. Conversely, a second value (e.g., 0) for `optimization_enable_flag` indicates that optimization is not applied to the input source. When `optimization_enable_flag` is absent, it can be assumed that optimization is not applied to the input source, and the value of `optimization_enable_flag` can be inferred to be the second value (e.g., 0). The value of `optimization_enable_flag` can be used to confirm whether the encoded bitstream is related to the original input source or the optimized input source. Also, unlike Table 1, two syntax elements can be used to indicate whether optimization is applied. For example, when the optimization enable flag (e.g., `optimization_enable_flag`) indicates whether optimization is available and optimization is available (e.g., `optimization_enable_flag=1`), the optimization flag (e.g., `optimization_flag`) can indicate whether optimization is applied to the input source.
[0171] The syntax elements `level_2_infomation_enable_flag` and `level_2_information` can respectively indicate whether a second-level interface is defined and the corresponding information. In the example, the second-level interface information may include information identifying optimization attributes and purposes.
[0172] Table 2 below shows examples of how first and second level interfaces are defined.
[0173] [Table 2]
[0174] Referring to Table 2, optimization_info() can include the syntax elements optimization_enable_flag, optimization_type_flag, optimization_type, optimization_purpose_flag, and optimization_purpose.
[0175] The syntax element `optimization_enable_flag` is first-level interface information, which is the same as described above with reference to Table 1.
[0176] The syntax element `optimization_type_flag` can indicate whether `optimization_type` is defined. For example, a first value (e.g., 1) for `optimization_type_flag` indicates that `optimization_type` is defined. Conversely, a second value (e.g., 0) for `optimization_type_flag` indicates that `optimization_type` is not defined.
[0177] When `optimization_type_flag` is the first value (e.g., 1), the syntax element `optimization_type` can be defined. `optimization_type` can represent optimization attributes applied to the input source. Optimization attributes can be represented by the type of noise added to the input source and the method used to remove / reduce it, the type of information to be improved and the method used to improve it, etc.
[0178] The syntax element `optimization_purpose_flag` can indicate whether `optimization_purpose` is defined. For example, a first value (e.g., 1) for `optimization_purpose_flag` indicates that `optimization_purpose` is defined. Conversely, a second value (e.g., 0) for `optimization_purpose_flag` indicates that `optimization_purpose` is not defined.
[0179] When `optimization_purpose_flag` is the first value (e.g., 1), the syntax element `optimization_purpose` can be defined. `optimization_purpose` can represent the type of task that can be performed, i.e., the target application.
[0180] Furthermore, the requirements of the receiving side may differ depending on the device's performance and operating environment. Therefore, even when the purpose is the same, the applied optimization attributes may differ, and the optimization attributes may not accurately represent the executable application. Therefore, it is necessary to define the optimization attribute (optimization_type) and optimization purpose (optimization_purpose) as second-level interface information. Through this second-level interface information, users can select and request bitstreams suitable for the device's performance and operating environment.
[0181] The specific examples used to express the optimization attribute (optimization_type) are the same as those in Tables 3 and 4 below.
[0182] First, Table 3 shows the application of various optimization methods.
[0183] [Table 3]
[0184] Referring to Table 3, when "optimization_type==0" is true, the optimization method determined by the application can be used as is. When "optimization_type>0 &(optimization_type &0x01)==0" is true, temporal resampling can be excluded from the optimization method. When "(optimization_type &0x01)!=0" is true, temporal resampling can be used as the optimization method. When "optimization_type>0 && (optimization_type &0x02)==0" is true, spatial resampling can be excluded from the optimization method. When "(optimization_type &0x02)!=0" is true, spatial resampling can be used as the optimization method. When "optimization_type>0 && (optimization_type &0x04)==0" is true, the region of interest (ROI) based optimization method can be omitted. When "(optimization_type &0x04)!=0" is true, the ROI-based optimization method can be used. Spatial quality optimization can be omitted when "optimization_type > 0 && (optimization_type & 0x08) == 0" is true. Quality optimization can be used when "(optimization_type & 0x08) != 0" is true. When using quality optimization, images can be preprocessed to reduce unnecessary information or improve the quality of necessary information (e.g., reduce noise and remove speckles).
[0185] Table 4 covers the case where only one optimization method is applied.
[0186] [Table 4]
[0187] Referring to Table 4, when the optimization_type value is "000", temporal resampling can be used as an optimization method. When the optimization_type value is "001", spatial resampling can be used as an optimization method. When the optimization_type value is "010", ROI-based optimization methods can be used. When the optimization_type value is "011", quality optimization can be used. When using spatial quality optimization, images can be preprocessed to reduce unnecessary information or improve the quality of necessary information (e.g., reduce noise and remove speckles).
[0188] Furthermore, the optimization methods described above (e.g., time resampling, etc.) are merely examples, and the embodiments of this disclosure are not limited thereto. In other words, the number of bits or the expression method of optimization_type in Tables 3 and 4 can be extended to express various optimization methods.
[0189] The specific examples used to express the optimization purpose are the same as those in Tables 5 and 6 below.
[0190] First, Table 5 shows the application of various optimization methods.
[0191] [Table 5]
[0192] Referring to Table 5, when "optimization_purpose == 0" is true, the optimization purpose determined by the application can be applied as is. When "optimization_purpose > 0 && (optimization_purpose & 0x01) == 0" is true, object detection may not be included in the optimization purpose. When "(optimization_purpose & 0x01) != 0" is true, object detection can be included in the optimization purpose. When "optimization_purpose > 0 && (optimization_purpose & 0x02) == 0" is true, object tracking may not be included in the optimization purpose. When "(optimization_purpose & 0x02) != 0" is true, object tracking can be included in the optimization purpose. When "optimization_purpose > 0 && (optimization_purpose & 0x04) == 0" is true, object segmentation may not be included in the optimization purpose. When "(optimization_purpose & 0x04) != 0" is true, object segmentation can be included in the optimization purpose. When "optimization_purpose > 0 && (optimization_purpose & 0x08) == 0" is true, face recognition can be excluded from the optimization purpose. When "(optimization_purpose & 0x08) != 0" is true, face recognition can be included in the optimization purpose. When "optimization_purpose > 0 && (optimization_purpose & 0x10) == 0" is true, human viewing can be excluded from the optimization purpose. When "(optimization_purpose & 0x10) != 0" is true, human viewing can be included in the optimization purpose. When "optimization_purpose > 0 && (optimization_purpose & 0x20) == 0" is true, machine analysis can be excluded from the optimization purpose. When "(optimization_purpose & 0x20 ) != 0" is true, machine analysis can be included in the optimization purpose.
[0193] Table 6 covers the case where only one optimization method is applied.
[0194] [Table 6]
[0195] Referring to Table 6, when the optimization_purpose value is "000", object detection can be included in the optimization purpose. When the optimization_purpose value is "001", object tracking can be included in the optimization purpose. When the optimization_purpose value is "010", object segmentation can be included in the optimization purpose. When the optimization_purpose value is "011", face recognition can be included in the optimization purpose. When the optimization_purpose value is "100", human viewing can be included in the optimization purpose. When the optimization_purpose value is "101", machine analysis can be included in the optimization purpose.
[0196] Furthermore, the aforementioned optimization objectives (e.g., object detection, etc.) are merely examples, and the embodiments of this disclosure are not limited thereto. In other words, the number of bits for the optimization_purpose or expression method in Tables 5 and 6 can be expanded to express various optimization objectives.
[0197] Referring again to Table 2, which defines the first and second level interfaces, optimization_info() can further include the syntax elements level_3_information_enable_flag and level_3_information.
[0198] The syntax elements `level_3_information_enable_flag` and `level_3_information` can respectively indicate whether a third-level interface is defined and the corresponding information. In the example, the third-level interface information may include details of the optimization methods applied to the input source.
[0199] General information about the original input source (e.g., frame rate, source width / height, etc.) can be altered through optimization, and depending on the application, information related to these general changes may be required. For example, when it's necessary to restore the spatial resolution reduced through optimization to the original resolution, the encoder / decoder must know the original input source resolution or the rate of resolution reduction. Alternatively, when applying ROI-based optimization, the encoder / decoder must know the ROI and non-ROI regions. Therefore, detailed information including information related to general changes can be defined separately as a third-level interface.
[0200] Tables 7 through 9 below show examples of defining all first-level to third-level interfaces.
[0201] [Table 7]
[0202] [Table 8]
[0203] [Table 9]
[0204] Table 7 is an example of expressing optimization attributes (optimization_type) and optimization purposes (optimization_purpose) based on Tables 3 and 5 above. Table 8 is an example of expressing multiple optimization attributes and optimization purposes based on Tables 4 and 6 above, and Table 9 is an example of expressing one optimization attribute and optimization purpose based on Tables 4 and 6 above. In Table 8, to express multiple optimization attributes and optimization purposes applied to the input source, the syntax element num_of_optimization_type representing the number of optimization attributes and the syntax element num_of_optimization_purpose representing the optimization purpose can be additionally defined.
[0205] For example, through reference Figure 12 As described in the related tables above, it is possible to individually determine, at each step, which level of interface will be responsible for defining the optimization information. For example, it is possible to determine whether to define a second-level interface (e.g., level_2_information_enable_flag) within the first-level interface, and whether to define a third-level interface (e.g., level_3_information_enable_flag) within the second-level interface. Conversely, it is possible to immediately determine, within the first-level interface (the highest level), which will be responsible for defining the optimization information. Figure 13 The diagram is shown in the figure. Table 10 shows the data based on... Figure 13 An example of the optimization_info() function of the interface structure.
[0206] [Table 10]
[0207] Referring to Table 10, `optimization_info()` can include the syntax element `optimization_info_level`, which indicates the level of the interface for which optimization information will be defined. `optimization_info()` can further include corresponding level interface information based on the value of `optimization_info_level`. For example, when the value of `optimization_info_level` is 0, `optimization_info()` only includes first-level interface information. Conversely, when the value of `optimization_info_level` is 1, `optimization_info()` further includes `optimization_type` and `optimization_purpose` as second-level interface information. Conversely, when the value of `optimization_info_level` is 2, `optimization_info()` further includes third-level interface information such as `ROI_info_flag`, in addition to the aforementioned second-level interface information.
[0208] Additionally, in Tables 7 through 10 above, the syntax element ROI_info_flag can indicate whether the syntax ROI_info() representing ROI-related information is defined. For example, a first value (e.g., 1) for ROI_info_flag indicates that ROI_info() is defined. Conversely, a second value (e.g., 0) for ROI_info_flag indicates that ROI_info() is not defined. ROI_info_flag can be defined when the optimization attribute (optimization_type) includes ROI-based optimizations. Examples of the ROI_info() syntax structure are the same as those in Table 11.
[0209] [Table 11]
[0210] Referring to Table 11, ROI_info() can use the syntax element num_of_ROI.
[0211] The syntax element `num_of_ROI` can represent the number of ROIs within the input source (e.g., an image). As many `region_info` values representing the location and size of ROI regions as `num_of_ROI` indicate the number of ROIs can be defined.
[0212] Examples of expressing the location and size of the optimized ROI region within the input source in a pixel unit and... Figure 14 The same as the illustration shown in the image. Figure 14In this context, the variable `topX` represents the horizontal position of the top-left point of the ROI (shadow area) within a pixel unit, and the variable `topY` represents the vertical position of the top-left point of the ROI within a pixel unit. Additionally, `variable width` represents the width of the ROI within a pixel unit, and `variable height` represents the height of the ROI within a pixel unit. Furthermore, unlike... Figure 12 As illustrated in the diagram, the position and size of the ROI region within the input source can also be expressed in a predetermined cell larger than a pixel. For this purpose, the example of the ROI_info() syntax structure is the same as in Table 12.
[0213] [Table 12]
[0214] Referring to Table 12, ROI_info() can further include the syntax elements unit_width and unit_height, which represent the expression units of the ROI region, unlike Table 11.
[0215] The syntax element `unit_width` can represent the width of a pre-determined unit obtained by partitioning the input source into regions of interest (ROIs) of a specific size. Additionally, the syntax element `unit_height` can indicate the height of a pre-determined unit. When defining appropriate partitioning units in advance, taking into account the processing units of the encoder and decoder, it is not necessary to define `unit_width` and `unit_height` to represent the size of the partitioning units.
[0216] Examples illustrating the location and size of the ROI region within the optimized input source in the aforementioned predetermined partitioning units are as follows: Figure 15 The same as the illustration shown in the image. Figure 15 In this context, the variable `topUnitX` represents the horizontal position of the top-left point of the ROI region (shaded area) within a predefined partition unit, and the variable `topUnitY` represents the vertical position of the top-left point of the ROI region within the partition unit. For example, the `topUnitX` of the ROI region can be 1 and `topUnitY` can be 2. Additionally, the variable `num_of_horizontal_unit` represents the width of the ROI region within the partition unit, and the variable `num_of_vertical_unit` represents the height of the ROI region within the partition unit. For example, the `num_of_horizontal_unit` of the ROI region can be 2 and `num_of_vertical_unit` can be 1.
[0217] Additionally, ROI_info() can further include attribute information of predefined regions within the input source. An example of the ROI_info() syntax structure for further including region attribute information is the same as in Table 13.
[0218] [Table 13]
[0219] Referring to Table 13, ROI_info() can further include the syntax element region_property, unlike Tables 11 and 12.
[0220] The syntax element `region_property` can represent the properties of a predefined region within the input source. The properties of the predefined region can be divided into non-ROI or ROI. In other words, `region_property` can have a first value (e.g., 0) representing a non-ROI (e.g., background) or a second value (e.g., 1) representing an ROI (e.g., foreground), as shown in Table 14 below.
[0221] [Table 14]
[0222] Another example of the ROI_info() syntax structure that further includes region attribute information is the same as that in Table 15.
[0223] [Table 15]
[0224] Referring to Table 15, ROI_info() can further include the syntax element region_property, unlike Tables 11 and 12.
[0225] The syntax element `region_property` can represent an attribute of a Region of Interest (ROI) within the input source. ROI attributes can be categorized by importance. In other words, `region_property` can have one of a first value (e.g., 0) representing importance level 0 to a fourth value (e.g., 4) representing importance level 3, as shown in Table 16 below.
[0226] [Table 16]
[0227] Referring again to Tables 7 to 10 above, optimization_info() can further include the syntax element temporal_optimization_info_flag.
[0228] The syntax element `temporal_optimization_info_flag` indicates whether `temporal_optimization_info()`, which includes information related to time resampling, is defined. For example, a first value (e.g., 1) of `temporal_optimization_info_flag` indicates that `temporal_optimization_info()` is defined. Conversely, a second value (e.g., 0) of `temporal_optimization_info_flag` indicates that `temporal_optimization_info()` is not defined. `temporal_optimization_info_flag` can only be defined if the optimization attribute (`optimization_type`) includes time resampling. Examples of `temporal_optimization_info()` are the same as those in Table 17.
[0229] [Table 17]
[0230] Referring to Table 17, temporal_optimization_info() can include the syntax elements original_temporal_info and optimization_method_idc.
[0231] The syntax element `origal_temporal_info` can be the necessary information to reconstruct the temporal information of the optimized input source from the temporal information of the original input source, or to compare the two. For example, `original_temporal_info` can include the frame rate and / or temporal resampling ratio of the original input source.
[0232] The syntax element `optimization_method_idc` can be an identifier for the time optimization method applied to the input source. `optimization_method_idc` can be used when, for a specific purpose, the decoded information needs to be reconstructed to the time level of the original input source.
[0233] Referring again to Tables 7 through 10, optimization_info() can further include the syntax element spatial_optimization_info_flag.
[0234] The syntax element `spatial_optimization_info_flag` can indicate whether `spatial_optimization_info()` including information related to spatial resampling (or, spatial resampling) is defined. For example, a first value (e.g., 1) of `spatial_optimization_info_flag` indicates that `spatial_optimization_info()` is defined. Conversely, a second value (e.g., 0) of `spatial_optimization_info_flag` indicates that `spatial_optimization_info()` is not defined. `spatial_optimization_info_flag` can only be defined if the optimization attribute (`optimization_type`) includes spatial resampling. Examples of `spatial_optimization_info()` are the same as those in Table 18.
[0235] [Table 18]
[0236] Referring to Table 18, spatial_optimization_info() can include the syntax elements origami_spatial_info and optimization_method_idc.
[0237] The syntax element `origal_spatial_info` can be the necessary information to reconstruct the spatial level of the optimized input source to the spatial level of the original input source, or to compare the two. For example, `original_spatial_info` can include the resolution and / or spatial resampling ratio of the original input source.
[0238] The syntax element `optimization_method_idc` can be identifying information about the spatial optimization method applied to the input source. `optimization_method_idc` can be used when, for a specific purpose, the decoded information needs to be reconstructed at the spatial level of the original input source.
[0239] Referring again to Tables 7 through 10, optimization_info() can further include the syntax element quality_optimization_info_flag.
[0240] The syntax element `quality_optimization_info_flag` indicates whether `quality_optimization_info()` containing information related to quality optimization is defined. For example, a first value (e.g., 1) of `quality_optimization_info_flag` indicates that `quality_optimization_info()` is defined. Conversely, a second value (e.g., 0) of `quality_optimization_info_flag` indicates that `quality_optimization_info()` is not defined. `quality_optimization_info_flag` can only be defined if the optimization attribute (`optimization_type`) includes quality optimization. Examples of `quality_optimization_info()` are the same as those in Table 19.
[0241] [Table 19]
[0242] Referring to Table 19, quality_optimization_info() can contain the syntax elements optimization_type and optimization_method_idc.
[0243] The syntax element `optimization_type` can be attribute information of the optimization method applied. The purpose of optimization can be to remove noise or enhance or weaken specific attribute information. Furthermore, optimization methods can alter the semantic information of the input source. Therefore, information consumers can use the attribute information of the optimization method to determine whether the corresponding optimization method is suitable for the task objective and its impact.
[0244] The syntax element `optimization_method_idc` can be identification information for the applied optimization method. It's even possible to implement and apply optimization methods with the same purpose and similar properties in different ways. For example, optimization methods can be defined differently based on the operating conditions and performance of the device to which they are applied. Therefore, information consumers can use the identification information of the optimization method to specifically identify the tool used for the corresponding optimization method. Examples of `optimization_method_idc` are the same as those in Table 20.
[0245] [Table 20]
[0246] Referring to Table 20, when the optimization_method_idc value is "00", a de-noising filter can be used to optimize the input source. When the optimization_method_idc value is "01", an edge preserving filter can be used to optimize the input source. When the optimization_method_idc value is "10", the image enhancement process can be used to optimize the input source. When the optimization_method_idc value is "11", a filter bank (an array of bandpass filters) can be used to optimize the input source.
[0247] In order to send the above-mentioned optimization information to the receiving side 1020 and 1120, it is necessary to include the optimization information sent to encoder 1013 and 1113 in the bit stream through the definition of the interface (INF) between optimizer 1012 and 1112 and encoder 1013 and 1113.
[0248] Figure 16 This is a diagram illustrating the process of optimizing the input source and sending it to the encoder. Figure 16 In this context, 0 to 8 represent the frame numbers of the input source, which can be the same as the input / output order of each frame. Additionally, time point A refers to the time point when optimization method A is applied, and time point B refers to the time point when the new optimization method B is applied.
[0249] refer to Figure 16 At time point A, optimization method A, corresponding to optimization_info_A, is applied to frame 0. Furthermore, optimization method A is also applied in input / output order to frames 1 through 3 following frame 0. Subsequently, at time point B, the optimization method is changed, and optimization method B, corresponding to optimization_info_B, is applied to frame 4. Furthermore, optimization method B is also applied in input / output order to frames 5 through 8 following frame 4.
[0250] Frames 0 through 8, optimized by optimization method A or B, are sent sequentially from the optimizer to the encoder.
[0251] Figure 17 It means to Figure 16 A diagram illustrating the process of encoding the optimized input source.
[0252] refer to Figure 17 The encoding order (and decoding order) of the input source can differ from the order in which it is sent from the optimizer. For example, in Figure 17In this process, the optimized input source can be encoded in the order of "frame 0 → frame 4 → frame 2 → frame 1 → frame 3 → frame 8 → frame 6 → frame 5 → frame 7". Therefore, because the encoding order of the input source may differ from the input / output order, the timing of changing the optimization method may also differ between the optimization and encoding processes. Specifically, in Figure 16 In the example, the discontinuity of the optimization method only occurs at time point B (i.e., the time point from frame 3 to frame 4), but in Figure 17 In the example, discontinuities in the optimization method may occur at each of the three time points (i.e., the time points of frame 0 → frame 4, frame 4 → frame 2, and frame 3 → frame 8).
[0253] When frame 4 is encoded, frame 0 can be referenced only. However, since the optimization methods applied to the two frames are different, the properties of the information can also be different. Therefore, referencing frame 0 to facilitate encoding frame 4 may be inappropriate. Furthermore, when encoding frame 2, referencing frame 0, which applies the same optimization method, is fine, but referencing frame 4, which applies a different optimization method, may be unsuitable. Therefore, the encoder must be able to determine whether the optimization methods are the same between the target frame and the reference candidate frames. Thus, information can be defined to identify the optimization method applied to each frame of the input source. Examples of syntax structures including this identification information are the same as those in Tables 21 and 22.
[0254] [Table 21]
[0255] [Table 22]
[0256] First, referring to Table 21, `input_optimization_info()` can include the syntax element `info_id`. The syntax element `info_id` can be information used to identify the preprocessing information `optimization_info()` sent from the optimizer (see Tables 1-2 and 7-10). In the examples, the value of `info_id` can be defined based on the order in which `optimization_info()` is received. For example, the value of `info_id` for the first `optimization_info()` received by the encoder can be 0, and the value of `info_id` for the second `optimization_info()` received can be 1. In another example, the value of `info_id` can also be defined based on whether the information in `optimization_info()` has been changed. For example, when the first `optimization_info()` received by the encoder and the second `optimization_info()` received contain the same information, the value of `info_id` can be 0 for both. Conversely, when the second `optimization_info()` received by the encoder and the third `optimization_info()` received contain different information, the value of `info_id` for the third `optimization_info()` received can be 1.
[0257] Next, referring to Table 22, `picture_parameter_set_rbsp()` can include the syntax element `info_id`. Table 22 shows the methods for including the aforementioned `info_id` in existing syntax structures, without defining new syntax structures as shown in Table 21. According to Table 22, identification information for optimization methods can be defined in picture units. Of course, unlike Table 22, identification information can also be defined in specific units other than pictures (e.g., frame groups).
[0258] Simultaneously, the same optimization method can be applied until new identification information is received. In this case, optimization information for consecutive frames, in input / output order, can be obtained (or inferred) from the most recently received optimization_info() until a new optimization_info() is received.
[0259] Tables 23 and 24 show examples of methods used to express the input_optimization_info() method in Table 22 above.
[0260] [Table 23]
[0261] [Table 24]
[0262] input_optimization_info() can be defined as a new NAL cell type, as shown in Table 23. Alternatively, input_optimization_info() can also be defined via a Supplemental Enhancement Information (SEI) message, as shown in Table 24. In Table 24, the payload type value (i.e., 205) for input_optimization_info() is merely an example and can be determined to be any value.
[0263] Meanwhile, when each frame of the input source has a different info_id, pre-defined identification information can be defined to minimize inefficiencies during the reference process.
[0264] In the example, additionally defined identification information may include `optimization_compensation_flag`. `optimization_compensation_flag` can indicate whether to compensate for attribute differences between the target and reference frames due to different optimization methods (e.g., when the `info_id` value is different or when different methods are applied to each ROI region). For example, a first value (e.g., 1) for `optimization_compensation_flag` can indicate that compensation is performed between the target and reference frames. Conversely, a second value (e.g., 0) for `optimization_compensation_flag` can indicate that no compensation is performed between the target and reference frames. `optimization_compensation_flag` can be defined in a frame unit via a picture header or picture parameter set, or in a frame group unit via a sequence parameter set, picture group header, etc.
[0265] In another example, additionally defined identification information may include `constrained_prediction_flag`. `constrained_prediction_flag` can indicate whether a predefined reference constraint is applied. For example, a `constrained_prediction_flag` with a first value (e.g., 1) can indicate that a reference constraint is applied. Examples of reference constraints here may include priority restrictions (e.g., referencing candidate frames for which different optimization methods are applied using lower priority references), restricted references (e.g., the texture information of candidate frames for which different optimization methods are applied is not referenced), etc. Conversely, a `constrained_prediction_flag` with a second value (e.g., 0) can indicate that a reference constraint is not applied. `constrained_prediction_flag` can be defined in a frame unit via a picture header or picture parameter set, or in a frame group unit via a sequence parameter set, picture group header, etc.
[0266] According to Embodiment 1 of this disclosure, a preprocessing operation based on a predetermined optimization method can be performed on the input source of a VCM. Various information and interfaces for expressing, identifying, and sending the optimization method can be defined. Additionally, an encoding method (e.g., reference constraints) considering the optimization method can be provided. Through such preprocessing (optimization process), noise included in the input source can be effectively removed, and the necessary information can be further enhanced to match task attributes and objectives.
[0267] Example 2
[0268] Multiple optimization methods can be applied to a single input source. In this case, inefficiencies in the optimization process can arise when the application order of these methods is not defined. For example, when temporal resampling is performed after spatial resampling, unnecessary spatial resampling can occur, even for information scheduled to be removed via temporal resampling. Furthermore, the interplay between different optimization methods needs to be considered during the optimization process. For instance, region information defined during ROI-based optimization can be reused for adaptive processing of each region during quality optimization. Additionally, when post-processing is required (e.g., conversion to the original input frame rate, conversion to the original input resolution, etc.), the optimization application order needs to be defined for more efficient post-processing.
[0269] Therefore, according to Embodiment 2 of this disclosure, when multiple optimization methods are applied to an input source, various methods can be provided for defining the application order among the multiple optimization methods.
[0270] Figure 18This is a diagram showing the application sequence of the optimization method according to an embodiment of the present disclosure.
[0271] refer to Figure 18 Different optimization methods can be assigned to each bit of the aforementioned `optimization_type`. For example, temporal resampling can be assigned to the least significant bit (LSB) of the `optimization_type`, spatial resampling to the second LSB, ROI-based optimization to the third LSB, and quality optimization to the fourth LSB. In this case, multiple optimization methods can be applied in ascending bit order, starting with the temporal resampling assigned to the LSB. Alternatively, multiple optimization methods can be applied in descending bit order, starting with the quality optimization method assigned to the fourth LSB. Thus, according to the embodiment, the application order among the multiple optimization methods can be determined, and the `optimization_type` can be defined based on the application order. In this case, it is not necessary to define information related to the application order separately.
[0272] Meanwhile, when defining the `optimization_type` based on the application order is impossible or the application order can be changed, it is necessary to define information related to the application order. Therefore, in the embodiment, information related to the application order can be defined separately in the preprocessing information `optimization_info()` (see Tables 1-2 and 7-10). Specific examples are the same as those in Tables 25 to 27.
[0273] [Table 25]
[0274] [Table 26]
[0275] First, referring to Table 25, optimization_info() can include the syntax elements num_of_optimization and ordering_criterion_idc.
[0276] The syntax element `num_of_optimization` can represent the number of optimization methods applied to the input source. In an embodiment, `num_of_optimization` can be derived from a predefined `optimization_type`. In this case, for example, `num_of_optimization` can be set to the same number of bits with a value of 1 in `optimization_type`.
[0277] The syntax element `ordering_criterion_idc` can represent criterion information related to the application order of multiple optimization methods. Table 26 shows examples of `ordering_criterion_idc`. Referring to Tables 26 and 25, when `ordering_criterion_idc` is "00", multiple optimization methods can be applied in LSB order of `optimization_type`. When `ordering_criterion_idc` is "01", multiple optimization methods can be applied in MSB order of `optimization_type`. When `ordering_criterion_idc` is "10", a separate application order (`alternative_ordering`) can be defined. When `ordering_criterion_idc` is "11", multiple optimization methods can be applied in any order.
[0278] [Table 27]
[0279] Next, referring to Table 27, `optimization_info()` may include the syntax elements `num_of_optimization`, `ordering_criterion_idc`, and `alternative_ordering_flag`. The syntax elements `num_of_optimization` and `ordering_criterion_idc` are the same as those described above with reference to Table 25.
[0280] The syntax element `alternative_ordering_flag` can indicate whether the application order of multiple optimization methods is defined separately. For example, a first value (e.g., 1) for `alternative_ordering_flag` indicates that the application order is defined separately. Conversely, a second value (e.g., 0) for `alternative_ordering_flag` indicates that the application order is not defined separately. Unlike Table 28, `alternative_ordering_flag` can only be defined when multiple optimization methods are applied to the input source (i.e., `num_of_optimization > 1`). Additionally, `alternative_ordering_flag` may not be defined if the basic application order is predefined based on the MSB or LSB.
[0281] Table 28 shows examples of information utilization among optimization methods.
[0282] [Table 28]
[0283] Referring to Table 28, the quality_optimization_info() function for quality optimization methods can include the syntax elements adaptive_quality_optimization_flag, optimization_type[i], optimization_methods[i], optimization_type, and optimization_method.
[0284] The syntax element `adaptive_quality_optimization_flag` indicates whether an adaptive quality optimization method is applied to the ROI region. For example, a first value (e.g., 1) for `adaptive_quality_optimization_flag` indicates that the adaptive quality optimization method is applied to the ROI region. Conversely, a second value (e.g., 0) for `adaptive_quality_optimization_flag` indicates that the adaptive quality optimization method is not applied to the ROI region. `adaptive_quality_optimization_flag` can only be defined if `optimization_type` includes ROI-based optimization (i.e., `(optimization_type & 0x04) != 0)`). Additionally, `adaptive_quality_optimization_flag` can only be defined if ROI-related details are defined (i.e., `ROI_info_flag = true`).
[0285] The syntax element `optimization_methods[i]` can represent the optimization method applied to each ROI region. Alternatively, the value of `optimization_methods[i]` can be defined using the `region_property` included in `ROI_info()` in Table 13 above. In this case, different optimization methods can be applied based on the attributes (e.g., ROI vs. non-ROI) or importance of each region.
[0286] The syntax element `optimization_type[i]` can represent an attribute of the optimization method applied to each ROI region. The value of `optimization_type[i]` can also be defined using the `region_property` included in `ROI_info()` in Table 13 above. In this case, different optimization attributes can be applied based on the attributes (e.g., ROI vs. non-ROI) or importance of each region. For example, optimization attributes designed to remove unnecessary information (or noise) can be applied to regions with low importance, while optimization attributes designed to enhance necessary information can be applied to regions with high importance.
[0287] The syntax element `optimization_type` can be attribute information of the optimization method applied. The purpose of optimization can be to remove noise or enhance or weaken specific attribute information. Furthermore, optimization methods can alter the semantic information of the input source. Therefore, information consumers can use the attribute information of the optimization method to determine whether the corresponding optimization method is suitable for the task objective and its impact.
[0288] The syntax element `optimization_method_idc` can be identification information of the applied optimization method. It is even possible to implement and apply optimization methods with the same purpose and similar properties in different ways. For example, optimization methods can be defined differently based on the operating conditions and performance of the device to which they are applied. Therefore, information consumers can use the identification information of the optimization method to specifically identify the tool used for the corresponding optimization method. Examples of `optimization_method_idc` are the same as those described above with reference to Table 19.
[0289] According to Embodiment 2 of this disclosure, various methods can be provided for defining the application order among multiple optimization methods. Therefore, the efficiency of the optimization process can be further improved, and the post-processing on the receiving side can be further improved.
[0290] Simultaneously, a post-processing process corresponding to the optimization (or pre-processing) processes of Embodiments 1 and 2 described above can be executed in the receiving-side device. For this purpose, the receiving-side device may further include a post-processor 1922, such as... Figure 19 As illustrated in the figure, the post-processor 1922 can identify the optimization methods and attributes applied to the input source based on optimization information obtained from the bitstream, and determine a specific post-processing method that is the same as or corresponds to the identified optimization methods and attributes. Furthermore, the post-processor 1922 can perform post-processing on the decoded image data based on the determined post-processing method.
[0291] Example 3
[0292] The optimization methods in Embodiments 1 and 2 described above can be defined differently depending on the purpose of the receiving device. This is because the necessary information can vary depending on the purpose of information consumption in the VCM system. Accordingly, an interface needs to be defined for the receiving device to request a suitable optimization method from the transmitting device. Furthermore, since the optimization method and its expression method (optimization information) can be changed according to the request of the receiving device, the interface must correspond to the interface defined between the optimizer and encoder of the transmitting device. Accordingly, Embodiment 3 of this disclosure provides an interface between the receiving device and the transmitting device for optimization operations.
[0293] Figure 20 This is a diagram illustrating the setup process of an optimization method according to an embodiment of the present disclosure.
[0294] refer to Figure 20 The second device 2020 (e.g., a client) must be able to obtain information related to the optimization method that can be applied to the first device 2010 (e.g., a server). To this end, the second device 2020 can request the first device 2010 to send information related to the applicable optimization method via a description function (a). In this case, the first device 2010 can send the information related to the applicable optimization method to the second device 2020 via a response function A (a').
[0295] Furthermore, the second device 2020 must be able to request the first device 2010 to apply an optimization method (including change / application cancellation, etc.). To this end, the second device 2020 can request the first device 2010 to apply an optimization method via a setting function (b). In this case, the first device 2010 can apply the optimization method and send information related to the applied optimization method to the second device 2020 via a response function B (b').
[0296] The above functions can be implemented through the interface (INF) between the first device 2010 and the second device 2020. Alternatively, the above functions can also be implemented by extending / applying existing protocols (e.g., RTSP / RTCP / WebRTC, etc.) or by defining new protocols.
[0297] The example of information sent via the response function A is the same as in Table 29.
[0298] [Table 29]
[0299] Referring to Table 29, the `available_optimization_info()` function used in response to function A can include the syntax element `available_optimization_type`. `available_optimization_type` represents the available optimization attributes. Examples of available optimization attributes are the same as those in Table 30.
[0300] [Table 30]
[0301] Referring to Table 30, when "available_optimization_type==0" is true, the optimization method determined by the application can be used as is. When "available_optimization_type>0 &&(available_optimization_type & 0x01)==0" is true, temporal resampling as an optimization method is not available. When "(available_optimization_type & 0x01)!=0" is true, temporal resampling is available as an optimization method. When "available_optimization_type>0 &&(available_optimization_type & 0x02)==0" is true, spatial resampling is not available as an optimization method. When "(available_optimization_type & 0x02)!=0" is true, spatial resampling is available as an optimization method. When "available_optimization_type>0 &&(available_optimization_type & 0x04)==0" is true, optimization methods based on the region of interest (ROI) are not available. When "(available_optimization_type & 0x04) != 0" is true, ROI-based optimization methods are available. When "available_optimization_type > 0 && (available_optimization_type & 0x08) == 0" is true, quality optimization may not be available. When "(available_optimization_type & 0x08) != 0" is true, quality optimization is available.
[0302] According to the embodiment, unlike Tables 29 and 30 above, the information sent via the response A function can be defined using existing protocols such as RTSP or RTCP. Specific examples are the same as in Table 31.
[0303] [Table 31]
[0304] Meanwhile, the example of the information required to disable the optimization settings is the same as in Table 32.
[0305] [Table 32]
[0306] Unlike activation requests, information related to the deactivation target may be unnecessary for deactivation requests. Therefore, as shown in Table 32, only the deactivation request message can be defined. In this case, the application of all optimization methods applied to the input source can be canceled.
[0307] Another example of the information required when requesting the deactivation of optimization settings is the same as in Table 33.
[0308] [Table 33]
[0309] Referring to Table 33, deactivation_request() can include the syntax elements deactivation_all_optimization_flag and deactivation_optimization_type.
[0310] The syntax element `deactivate_all_optimization_flag` can indicate whether optimization settings are selectively disabled. For example, a first value (e.g., 1) for `deactivate_all_optimization_flag` can indicate a request to disable all optimization settings. Conversely, a second value (e.g., 0) for `deactivate_all_optimization_flag` can indicate a request to partially (selectively) disable optimization settings.
[0311] The syntax element `deactivate_optimization_type` can represent optimization settings to be disabled. `deactivate_optimization_type` can be defined only when a selective deactivation request exists, i.e., when `deactivate_all_optimization_flag` has a second value (e.g., 0). Examples of selective deactivation requests are the same as those in Table 34.
[0312] [Table 34]
[0313] Referring to Table 34, when "deactivate_optimization_type==0" is true, the optimization method determined by the application can be used as is. When "deactivate_optimization_type>0 &&(deactivate_optimization_type & 0x01)==0" is true, there may be no change in the current state. When "(deactivate_optimization_type & 0x01)!=0" is true, the deactivation request may be related to temporal resampling. When "deactivate_optimization_type>0 &&(deactivate_optimization_type & 0x02)==0" is true, there may be no change in the current state. When "(deactivate_optimization_type & 0x02)!=0" is true, the deactivation request may be related to spatial resampling. When "deactivate_optimization_type>0 &&(deactivate_optimization_type & 0x04)==0" is true, there may be no change in the current state. When "(deactivate_optimization_type & 0x04) != 0" is true, the deactivation request may be related to ROI-based optimization methods. When "deactivate_optimization_type > 0 && (deactivate_optimization_type & 0x08) == 0" is true, there may be no change in the current state. When "(deactivate_optimization_type & 0x08) != 0" is true, the deactivation request may be related to quality optimization.
[0314] According to the embodiment, the information required when requesting deactivation can also be defined using existing protocols such as RTSP or RTCP, unlike Tables 32 to 34 above. In this case, the optimization methods to be deactivated are listed, as in Table 35.
[0315] [Table 35]
[0316] Additionally, it is possible to request selective activation of optimization settings. The example of information required when requesting selective activation is the same as in Table 36.
[0317] [Table 36]
[0318] According to embodiment 3 of this disclosure, an interface for optimization control can be defined between the transmitting-side device and the receiving-side device in the VCM. In this case, control (specifically, a deactivation request) can be executed individually for each optimization method, or it can be executed together for all optimization methods. Therefore, appropriate optimization control based on the purpose of the receiving-side device is possible.
[0319] In the following text, see references Figure 21 and Figure 22 The image data processing method according to embodiments of the present disclosure will be described in detail.
[0320] Figure 21 This is a flowchart illustrating an image data processing method according to an embodiment of the present disclosure. Figure 21 The image data processing method can be executed by the receiving device.
[0321] refer to Figure 21 The receiving device can decode the image data obtained from the bitstream S2110. Additionally, the receiving device can obtain information related to the optimization method of the image data S2120, and post-process the decoded image data based on the obtained information related to the optimization method S2130. In this case, the information related to the optimization method may include optimization attribute information indicating the type of at least one optimization method applied to the image data.
[0322] In the embodiments, the types of optimization methods may include temporal resampling, spatial resampling, region of interest (ROI)-based optimization, and quality optimization.
[0323] In an embodiment, the information related to the optimization method may further include information indicating the type number of the optimization method applied to the image data.
[0324] In an embodiment, the information related to the optimization method may further include optimization objective information representing the purpose of at least one optimization method applied to the image data. Here, the objective of the optimization method may include object detection, object tracking, object segmentation, face recognition, human viewing, and machine analysis.
[0325] In an embodiment, the information related to the optimization method may further include information indicating the quantity of the purpose of the optimization method applied to the image data.
[0326] In an embodiment, when an optimization method based on a region of interest (ROI) is applied to image data, the information related to the optimization method may further include information representing the attributes of each region included in the original image of the image data.
[0327] In an embodiment, the information related to the optimization method further includes information related to general information changes in the image data, and the information related to general information changes can be used to reconstruct the decoded image data.
[0328] In an embodiment, when multiple optimization methods are applied to image data, the information related to the optimization methods may further include information indicating the type of tool used for each of the multiple optimization methods.
[0329] In an embodiment, when multiple optimization methods are applied to image data, the information related to the optimization methods may further include information related to the application order of the multiple optimization methods.
[0330] In an embodiment, the information related to the optimization method may further include information indicating the type of optimization method that can be used with the image data. In this case, the available optimization methods can be controlled individually based on a selective deactivation request from the receiving device.
[0331] Figure 22 This is a flowchart illustrating an image data processing method according to an embodiment of the present disclosure. Figure 22 The image data processing method in the image can be executed by the transmitting device.
[0332] refer to Figure 22 The transmitting device can determine an optimization method for the image data (S2210) and preprocess the image data based on the determined optimization method (S2220). Furthermore, the transmitting device can encode the preprocessed image data and information related to the optimization method (S2230). In this case, the information related to the optimization method may include optimization attribute information indicating the type of at least one optimization method applied to the image data.
[0333] Simultaneously, information related to the optimization method encoded by the transmitting device can be included in the bitstream. This bitstream can be stored on a non-transitory computer-readable recording medium and transmitted to the receiving device via communication means (e.g., wired / wireless network).
[0334] Although the exemplary methods of this disclosure are expressed as a series of operations for clarity of explanation, this is not intended to limit the order in which the steps are performed, and each step may be performed simultaneously or in a different order if necessary. To implement the methods according to this disclosure, another step may be additionally included in the exemplary steps, or remaining steps may be included while excluding some steps, or another additional step may be included while excluding some steps.
[0335] In this disclosure, the image encoding device or image decoding device that performs a predetermined operation (step) can perform an operation (step) for checking conditions or conditions for performing the corresponding operation (step). For example, when it is stated that a predetermined operation will be performed when a predetermined condition is met, the image encoding device or image decoding device can perform an operation for checking whether the predetermined condition is met, and then perform the predetermined operation.
[0336] The various embodiments of this disclosure do not list all possible combinations, but are intended to describe representative aspects of this disclosure, and the matters described in the various embodiments may be applied independently or in combination of at least two.
[0337] The embodiments described in this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information for implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.
[0338] Furthermore, the decoders (decoding devices) and encoders (encoding devices) used in the embodiments of this disclosure can be included in multimedia broadcasting and receiving devices, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video communication devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, OTT (Over-The-Top) video devices, Internet streaming service providers, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony devices, transportation terminals (e.g., vehicle (including autonomous vehicles) terminals, robot terminals, aircraft terminals, ship terminals, etc.), medical video devices, etc., and can be used to process video signals or data signals. For example, OTT (Over-The-Top) video devices can include game consoles, Blu-ray players, Internet-connected televisions, home theater systems, smartphones, tablets, digital video recorders (DVRs), etc.
[0339] Furthermore, the processing methods applying embodiments of this disclosure can be generated in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having data structures according to embodiments of this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices for storing computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Additionally, computer-readable recording media include media implemented in carrier wave form (e.g., transmission via the Internet). Furthermore, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired or wireless communication networks.
[0340] Furthermore, the embodiments of this disclosure can be implemented as computer program products by program code, and the program code can be executed on a computer by the embodiments of this disclosure. The program code can be stored on a computer-readable medium.
[0341] Figure 23 This is a diagram illustrating an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0342] refer to Figure 23 The content streaming system using embodiments of this disclosure can broadly include encoding servers, streaming servers, web servers, media storage, user equipment, and multimedia input devices.
[0343] An encoding server compresses content input from multimedia input devices such as smartphones, cameras, or camcorders into digital data, generates a bitstream, and sends it to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, or camcorders directly generate bitstreams, the encoding server can be omitted.
[0344] The bitstream can be generated by the image encoding method and / or image encoding apparatus of the embodiments of this disclosure, and the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.
[0345] A streaming server can send multimedia data to a user's device based on a user request via a web server, and the web server can act as a medium to notify the user of available services. When a user requests a service from the web server, the web server can send the request to the streaming server, and the streaming server can transmit the multimedia data to the user. In this scenario, the content streaming system may include a separate control server, which in this case can control the commands / responses between devices within the content streaming system.
[0346] A streaming server can receive content from media storage and / or encoding servers. For example, content can be received in real time when it is received from an encoding server. In this case, to provide a seamless streaming service, the streaming server can store a bitstream for a certain period of time.
[0347] Examples of user devices may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet computers, ultrabooks, wearable devices (i.e., smartwatches, smart glass, head-mounted displays (HMDs)), digital televisions, desktop computers, digital signage, etc.
[0348] In a content streaming system, each server can operate as a distributed server, in which case the data received by each server can be processed in a distributed manner.
[0349] Figure 24 This is a diagram illustrating another example of a content streaming system to which embodiments of the present disclosure can be applied.
[0350] refer to Figure 24 In embodiments such as VCM, the task can be executed by a user terminal, or by an external device (e.g., a streaming server, analytics server, etc.) based on the device's performance, the user's request, the characteristics of the task to be performed, etc. In this way, in order to send the information necessary for executing the task to the external device, the user terminal can generate a bitstream directly or through an encoding server, which includes the information necessary for executing the task (e.g., information such as the task, neural network, and / or information used).
[0351] An analysis server can execute tasks requested by the user after decoding encoded information sent from a user terminal (or an encoding server). The analysis server can then send the results obtained from executing the tasks back to the user terminal or another linked service server (e.g., a web server). For example, the analysis server can send the results obtained from executing a task to determine a fire to a fire-related server. The analysis server may include a separate control server, in which case the control server can play a role in controlling the commands / responses between the analysis server and each device associated with it. Additionally, the analysis server can request desired information from a web server based on information about the tasks the user device wants to perform and the tasks the user device can perform. When the analysis server requests the desired service from the web server, the web server can forward it to the analysis server, and the analysis server can then send its data to the user terminal. In this scenario, the control server of the content streaming system can play a role in controlling the commands / responses between each device within the streaming system.
[0352] [Industrial Applicability]
[0353] The embodiments of this disclosure can be used to process image data.
Claims
1. An image data processing method performed by a receiving device, comprising: Decode image data obtained from a bitstream; Obtain information related to the optimization method for the image data; as well as Based on the information related to the optimization method, the decoded image data is post-processed. The information related to the optimization method includes optimization attribute information indicating the type of at least one optimization method applied to the image data.
2. The method according to claim 1, wherein, Based on the regional attributes of the image data, the type of optimization method is determined for each region.
3. The method according to claim 1, wherein, Decoding the image data based on inter-frame prediction, and The inter-frame prediction is performed based on whether the same optimization method is applied to the current frame and the reference frame.
4. The method according to claim 1, wherein, The information associated with the optimization method further includes information indicating the type number of the optimization method applied to the image data.
5. The method according to claim 1, wherein, The information associated with the optimization method further includes optimization objective information indicating the purpose of the at least one optimization method applied to the image data.
6. The method according to claim 5, wherein, The information associated with the optimization method further includes information indicating the amount of optimization purpose applied to the image data.
7. The method according to claim 1, wherein, An optimization method based on the region of interest (ROI) is applied to the image data, and the information associated with the optimization method further includes information representing the attributes of each region included in the image data.
8. The method according to claim 1, wherein, The information associated with the optimization method further includes information related to general information changes in the image data, and the information related to the general information changes is used to reconstruct the decoded image data.
9. The method according to claim 1, wherein, Since multiple optimization methods are applied to the image data, the information related to the optimization methods further includes information indicating the tool type used for each of the multiple optimization methods.
10. The method according to claim 1, wherein, Since multiple optimization methods are applied to the image data, the information related to the optimization methods further includes information related to the application order of the multiple optimization methods.
11. The method according to claim 1, wherein, The information associated with the optimization method further includes information indicating the type of optimization method that can be used for the image data.
12. The method according to claim 11, wherein, Based on the selective deactivation request of the receiving device, the available optimization methods are controlled individually.
13. An image data processing method performed by a transmitting device, comprising: Determine the optimization method for image data; Image data is preprocessed based on the optimization method described above; as well as The preprocessed image data and information related to the optimization method are encoded. The information related to the optimization method includes optimization attribute information indicating the type of at least one optimization method applied to the image data.
14. The method of claim 13, wherein a computer-readable recording medium stores a bit stream generated by the processing method.
15. A method for transmitting a bitstream generated by an image data processing method, the image data processing method comprising: Determine the optimization method for image data; Image data is preprocessed based on the optimization method described above; as well as The preprocessed image data and information related to the optimization method are encoded. The information related to the optimization method includes optimization attribute information indicating the type of at least one optimization method applied to the image data.