Image encoding / decoding method and apparatus using image segmentation and recording medium storing bitstream
Through an image splitting method adapted to device resources and peak memory, the problem of inefficient image encoding/decoding in the prior art is solved, and efficient image encoding/decoding and bitstream storage in neural networks are realized.
Patent Information
- Application Number
- CN202380083565.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-11
- Filing Date
- 2023-10-11
- Publication Date
- 2025-07-11
AI Technical Summary
The existing image compression technology is not suitable for artificial intelligence services, and it fails to effectively use limited resources for image encoding/decoding, resulting in ineffective encoding/decoding.
Through an image splitting method adapted to the available resources and peak memory of the device, the splitting information of the image is derived, and image encoding/decoding is performed in a neural network, and a bit stream is generated and stored.
It improves the efficiency of image encoding/decoding, adapts to device performance under different resource conditions, and reduces peak memory requirements.
Smart Images

Figure CN120303930A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding / decoding method and apparatus using image splitting, and a recording medium storing a bitstream, and more particularly, to an image encoding / decoding technique using image splitting in consideration of resources. Background Art
[0002] With the development of machine learning technology, the demand for artificial intelligence services based on image processing is increasing day by day. In order to effectively process a large amount of image data required for artificial intelligence services within limited resources, an image compression technology for optimizing machine task performance is crucial. However, existing image compression technologies have been developed for high-resolution and high-quality image processing for the human visual system, and there are problems that are not suitable for artificial intelligence services. Accordingly, research and development of a new type of machine-oriented image compression technology suitable for artificial intelligence services is actively underway. Summary of the Invention
[0003] Technical Problem
[0004] The present disclosure aims to provide an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.
[0005] In addition, the present disclosure aims to provide an image splitting and split region representation technique that is adaptive to the available resources of a device.
[0006] In addition, the present disclosure aims to provide an image splitting and split region representation technique that is adaptive to peak memory.
[0007] In addition, the present disclosure aims to provide a neural network-based image encoding / decoding technique.
[0008] In addition, the present disclosure aims to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0009] In addition, the present disclosure aims to provide a recording medium that stores a bitstream received and decoded by an image decoding device according to the present disclosure and used for reconstructing an image.
[0010] The technical objects of the present disclosure are not limited to the above technical objects, and those skilled in the art to which the present disclosure pertains will clearly understand other technical objects not described above based on the following description.
[0011] Technical Solution
[0012] An image decoding method performed by an image decoding device according to an aspect of the present disclosure may include the following steps: deriving the size of a processing unit of an image based on available resources of a decoder and peak memory, deriving a splitting method of the image based on the derived size of the processing unit, deriving splitting information of the image based on the splitting method, and connecting and reconstructing the processing unit of the image based on the derived splitting information of the image, wherein the splitting information of the image includes position information of the processing unit.
[0013] An image encoding method performed by an image encoding device according to an aspect of the present disclosure may include the following steps: determining the size of a processing unit of an image, determining a splitting method of the image based on the determined size of the processing unit, and
[0014] generating splitting information of the image based on the splitting method, wherein the splitting information of the image includes position information of the processing unit.
[0015] A bitstream sending method according to an aspect of the present disclosure may include the following steps: sending a bitstream generated by an image encoding method, wherein the image encoding method includes the following steps: determining the size of a processing unit of an image, determining a splitting method of the image based on the determined size of the processing unit, and generating splitting information of the image based on the splitting method, and the splitting information of the image includes position information of the processing unit.
[0016] A recording medium according to another aspect of the present disclosure may store a bitstream generated by the image encoding method or the image encoding device of the present disclosure.
[0017] A bitstream sending method according to still another aspect of the present disclosure may send a bitstream generated by the image encoding method or the image encoding device of the present disclosure to an image decoding device.
[0018] The features of the present disclosure briefly outlined above are merely illustrative aspects of the detailed description of the present disclosure and do not limit the scope of the present disclosure.
[0019] Beneficial Effects
[0020] According to the present disclosure, an image encoding / decoding method and device with improved encoding / decoding efficiency can be provided.
[0021] In addition, according to the present disclosure, an image encoding / decoding technique based on a neural network can be provided.
[0022] In addition, according to the present disclosure, an image encoding / decoding technique adaptable to the resources of a device can be provided.
[0023] In addition, according to the present disclosure, an image encoding / decoding technique considering peak memory can be provided.
[0024] In addition, according to the present disclosure, the availability of a device can be improved by enabling encoding / decoding according to the available resources of the device.
[0025] In addition, according to the present disclosure, a recording medium can be provided that stores a bitstream received and decoded by an image decoding device according to the present disclosure and used for reconstructing an image.
[0026] The effects of the present disclosure are not limited to the above effects, and those skilled in the art to which the present disclosure pertains will clearly understand other effects not described above from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a view schematically showing a VCM system to which an embodiment of the present disclosure is applicable.
[0028] Figure 2 is a diagram schematically showing a VCM pipeline structure to which an embodiment of the present disclosure is applicable.
[0029] Figure 3 is a diagram schematically showing an image / video encoder to which an embodiment of the present disclosure is applicable.
[0030] Figure 4 is a diagram schematically showing an image / video decoder to which an embodiment of the present disclosure is applicable.
[0031] Figure 5 is a flowchart schematically illustrating a feature / feature map encoding process to which an embodiment of the present disclosure is applicable.
[0032] Figure 6 is a flowchart schematically illustrating a feature / feature map decoding process to which an embodiment of the present disclosure is applicable.
[0033] Figure 7 is a diagram illustrating an example of a feature extraction and reconstruction method to which an embodiment of the present disclosure is applicable.
[0034] Figure 8 is a diagram illustrating an example of an image splitting method to which an embodiment of the present disclosure is applicable.
[0035] Figure 9 and Figure 10 is a diagram illustrating an example of an image encoding / decoding system including a machine video / image coding (VCM) image encoder and decoder.
[0036] Figure 11 is a diagram illustrating an example of a VCM hierarchical structure.
[0037] Figure 12 A diagram illustrating an example of a VCM bitstream composed of encoded abstract features and neural network abstraction layer (NNAL) information.
[0038] Figure 13 A diagram illustrating an example of the operation of an image encoder according to an embodiment of the present disclosure.
[0039] Figure 14 A diagram illustrating an example of the operation of an image decoder according to an embodiment of the present disclosure.
[0040] Figure 15 A diagram for describing an example of split input data according to an embodiment of the present disclosure.
[0041] Figure 16 Illustrates an example of an image splitting method according to an embodiment of the present disclosure.
[0042] Figures 17 to 19 A diagram for describing an example of the generation and configuration of image splitting information according to an embodiment of the present disclosure.
[0043] Figure 20 A diagram illustrating an example of bitstream configuration according to an embodiment of the present disclosure.
[0044] Figure 21 A diagram illustrating an image decoding method according to an embodiment of the present disclosure. Figure 22 A diagram illustrating an image encoding method according to an embodiment of the present disclosure.
[0045] Figure 23 A view illustrating an example of a content stream transmission system to which an embodiment of the present disclosure is applicable.
[0046] Figure 24 A view showing another example of a content stream transmission system to which an embodiment of the present disclosure is applicable. Detailed Description of the Embodiments
[0047] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. However, the present disclosure can be implemented in various different forms and is not limited to the embodiments described herein.
[0048] When describing the present disclosure, if the detailed description of related known functions or configurations makes the scope of the present disclosure unnecessarily ambiguous, the detailed description thereof will be omitted. In the drawings, parts irrelevant to the description of the present disclosure are omitted, and similar reference numerals are attached to similar parts.
[0049] In the present disclosure, when a component is "connected", "coupled", or "linked" to another component, it may include not only a direct connection relationship but also an indirect connection relationship with an intermediate component present. Additionally, when a component "includes" or "has" other components, unless otherwise stated, it means that other components may be further included rather than excluding other components.
[0050] In the present disclosure, terms such as first, second, etc. may be used only for the purpose of distinguishing one component from other components and do not limit the order or importance of the components, unless otherwise stated. Thus, within the scope of the present disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0051] In the present disclosure, components that are distinguishable from each other are intended to clearly describe each feature and do not mean that these components necessarily need to be separated. That is, multiple components may be integrated and implemented in one hardware or software unit, or one component may be distributed and implemented in multiple hardware or software units. Therefore, embodiments in which these components are integrated or a component is distributed are included within the scope of the present disclosure even without additional specification.
[0052] In the present disclosure, components described in various embodiments do not necessarily mean essential components, and some components may be optional components. Therefore, embodiments consisting of a subset of the components described in an embodiment are also included within the scope of the present disclosure. Additionally, embodiments including other components in addition to the components described in various embodiments are also included within the scope of the present disclosure.
[0053] The present disclosure relates to the encoding and decoding of images, and the terms used in the present disclosure may have the general meanings commonly used in the technical field to which the present disclosure pertains, unless newly defined in the present disclosure.
[0054] The present disclosure may be applied to the methods disclosed in the Versatile Video Coding (VVC) standard and / or the Machine Video Coding (VCM) standard. Additionally, the present disclosure may be applied to the methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Second Generation Audio Video Coding Standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268, etc.).
[0055] The present disclosure provides various embodiments related to video / image encoding, and these embodiments can be executed in combination with each other unless otherwise specified. In the present disclosure, "video" refers to a collection of a series of images over time. An "image" can be information generated by artificial intelligence (AI). The input information used in the process of performing a series of tasks by AI, the information generated during the information processing, and the output information can be used as images. In the present invention, a "picture" generally refers to a unit representing one image within a specific time period, and a slice / tile is an encoding unit that forms part of a picture during encoding. A picture can be composed of one or more slices / tiles. Additionally, a slice / tile can include one or more coding tree units (CTUs). A CTU can be divided into one or more CUs. A tile is a rectangular area in a specific tile row and a specific tile column existing in a picture, and can be composed of multiple CTUs. A tile column can be defined as a rectangular area of CTUs, can have the same height as the picture, and can have a width specified by a syntax element signaled from a bitstream part such as a picture parameter set. A tile row can be defined as a rectangular area of CTUs, can have the same width as the picture, and can have a height specified by a syntax element signaled from a bitstream part such as a picture parameter set. Tile scanning is a certain consecutive sorting method of CTUs that divides a picture. Here, the CTUs can be sorted sequentially according to the CTU raster scan order within a tile, and the tiles in a picture can be sorted sequentially according to the raster scan order of the tiles in the picture. A slice can contain an integer number of complete tiles, or can contain a consecutive integer number of complete CTU rows within one tile of a picture. A slice can be specifically contained in a single NAL unit. A picture can be composed of one or more tile groups. A tile group can include one or more tiles. A patch can indicate a rectangular area of CTU rows within a tile in a picture. A tile can include one or more patches. A patch can refer to a rectangular area of CTU rows in a tile. A tile can be split into multiple patches, and each patch can include one or more CTU rows belonging to one tile. A tile that is not split into multiple patches can also be regarded as a patch.
[0056] In the present disclosure, a "pixel" or "picture element" can mean the smallest unit that constitutes a picture (or image). Additionally, "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or the value of a pixel, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0057] In an embodiment, particularly when applied to a VCM, when there is an image composed of a set of components with different characteristics and meanings, a pixel / pixel value may represent the pixel / pixel value of a component generated through independent information or the combination, synthesis, and analysis of individual components. For example, in an RGB input, the pixel / pixel value of only R may be represented, the pixel / pixel value of only G may be represented, or the pixel / pixel value of only B may be represented. For example, the pixel / pixel value of only the luminance component synthesized using the R, G, and B components may be represented. For example, the pixel / pixel value of an image and information extracted by analyzing the R, G, and B components from the components may be represented.
[0058] In the present disclosure, a "unit" may represent a basic unit of image processing. The unit may include at least one of a specific region of an image and information related to the region. A unit may include one luminance block and two chrominance (e.g., Cb and Cr) blocks. In some cases, the unit may be used interchangeably with terms such as "sample array", "block", or "region". In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients in M columns and N rows. In an embodiment, particularly when applied to a VCM, the unit may represent a basic unit containing information for performing a specific task.
[0059] In the present disclosure, a "current block" may mean one of "current coding block", "current coding unit", "encoding target block", "decoding target block", or "processing target block". When performing prediction, a "current block" may mean "current prediction block" or "prediction target block". When performing transform (inverse transform) / quantization (dequantization), a "current block" may mean "current transform block" or "transform target block". When performing filtering, a "current block" may mean "filtering target block".
[0060] In addition, in the present disclosure, a "current block" may mean the luminance block of the "current block" unless explicitly stated as a chrominance block. The chrominance block of the "current block" may be expressed by including an explicit description of a chrominance block such as "chrominance block" or "current chrominance block".
[0061] In the present disclosure, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expressions "A / B" and "A,B" may mean "A and / or B". In addition, "A / B / C" and "A / B / C" may mean "at least one of A, B, and / or C".
[0062] In the present disclosure, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" may include 1) only "A", 2) only "B", and / or 3) both "A and B". In other words, in the present disclosure, the term "or" should be interpreted as indicating "additionally or alternatively".
[0063] The present disclosure relates to machine video / image coding (VCM).
[0064] VCM refers to a compression technique for encoding / decoding a part of a source image / video or information obtained from the source image / video for the purpose of machine vision. In VCM, the encoding / decoding target may be referred to as a feature. The feature may refer to information extracted from the source image / video based on task purposes, requirements, surrounding environments, etc. The feature may have an information form different from that of the source image / video, and accordingly, the compression method and expression format of the feature may also be different from those of the video source.
[0065] VCM can be applied to a variety of application fields. For example, in a surveillance system for identifying and tracking objects or people, VCM can be used to store or transmit object recognition information. In addition, in intelligent transportation or an intelligent transportation system, VCM can be used to send vehicle position information collected from GPS, sensing information collected from LIDAR, radar, etc., and various vehicle control information to other vehicles or infrastructure. In addition, in the field of smart cities, VCM can be used to perform individual tasks of interconnected sensor nodes or devices.
[0066] The present disclosure provides various embodiments of feature / feature map coding. Unless otherwise specified, the embodiments of the present disclosure can be implemented individually or in combination of two or more.
[0067] Overview of the VCM system
[0068] Figure 1 is a diagram schematically showing a VCM system to which the embodiments of the present disclosure are applicable.
[0069] Refer to Figure 1 , the VCM system may include an encoding device 10 and a decoding device 20.
[0070] The encoding device 10 may compress / encode the feature / feature map extracted from the source image / video to generate a bitstream, and send the generated bitstream to the decoding device 20 through a storage medium or a network. The encoding device 10 may also be referred to as a feature encoding device. In the VCM system, features / feature maps may be generated at each hidden layer of the neural network. The channel size and number of the generated feature maps may vary depending on the type of the neural network or the position of the hidden layer. In the present disclosure, the feature map may be referred to as a feature set, and the feature or feature map may be referred to as "feature information".
[0071] The encoding device 10 may include a feature acquisition unit 11, an encoding unit 12, and a transmission unit 13.
[0072] The feature acquisition unit 11 can acquire features / feature maps for the source image / video. Depending on the implementation, the feature acquisition unit 11 can acquire features / feature maps from an external device, e.g., a feature extraction network. In this case, the feature acquisition unit 11 performs the function of a feature receiving interface. Alternatively, the feature acquisition unit 11 can acquire features / feature maps by using the source image / video as input to execute a neural network (e.g., CNN, DNN, etc.). In this case, the feature acquisition unit 11 performs the function of a feature extraction network.
[0073] According to an implementation, the encoding device 10 can further include a source image generator (not shown) for acquiring the source image / video or instead of the feature acquisition unit 11. The source image generator can be implemented by using an image sensor, a camera module, etc., and can acquire the source image / video through an image / video capture, synthesis, or generation process. In this case, the generated source image / video can be sent to the feature extraction network and used as input data for extracting features / feature maps.
[0074] The encoding unit 12 can encode the features / feature maps acquired by the feature acquisition unit 11. The encoding unit 12 can perform a series of processes such as prediction, transformation, and quantization to increase the encoding efficiency. The encoded data (encoded feature / feature map information) can be output in the form of a bitstream. The bitstream containing the encoded feature / feature map information can be referred to as a VCM bitstream.
[0075] The transmission unit 13 can acquire the feature / feature map information or data output in the form of a bitstream, which can be delivered to the decoding device 20 or other external objects in the form of a file or a stream via a digital storage medium or a network. Here, the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmission unit 13 can include an element for generating a media file with a predetermined file format or an element for sending data through a broadcast / communication network. The transmission unit 13 can include a transmission device separate from the encoding unit 12. In this case, the transmission device can include at least one processor for acquiring the feature / feature map information or data output in the form of a bitstream and a transmission part for delivering the feature / feature map information or data in the form of a file or a stream.
[0076] The decoding device 20 can acquire the feature / feature map information from the encoding device 10 and reconstruct the features / feature maps based on the acquired information.
[0077] The decoding device 20 can include a receiving unit 21 and a decoding unit 22.
[0078] The receiving unit 21 can receive a bitstream from the encoding device 10, obtain feature / feature map information from the received bitstream, and send it to the decoding unit 22.
[0079] The decoding unit 22 can decode the feature / feature map based on the obtained feature / feature map information. The decoding unit 22 can perform a series of processes corresponding to the operations of the encoding unit 12, such as dequantization, inverse transformation, and prediction, to improve the decoding efficiency.
[0080] Depending on the implementation, the decoding device 20 can further include a task analysis / rendering unit 23.
[0081] The task analysis / rendering unit 23 can perform task analysis based on the decoded feature / feature map. Additionally, the task analysis / rendering unit 23 can render the decoded feature / feature map into a form suitable for task execution. Various machine (oriented) tasks can be performed based on the task analysis results and the rendered feature / feature map.
[0082] As described above, the VCM system can encode / decode the features extracted from the source image / video according to user and / or machine requests, task purposes, and the surrounding environment, and perform various machine (oriented) tasks based on the decoded features. The VCM system can be implemented by extending / redesigning the video / image encoding system and can perform various encoding / decoding methods defined in the VCM standard.
[0083] VCM pipeline
[0084] Figure 2 is a diagram schematically showing the VCM pipeline structure to which the embodiments of the present disclosure are applicable.
[0085] Refer to Figure 2 , the VCM pipeline 200 can include a first pipeline 210 for encoding / decoding an image / video and a second pipeline 220 for encoding / decoding a feature / feature map. In the present disclosure, the first pipeline 210 can be referred to as a video codec pipeline, and the second pipeline 220 can be referred to as a feature codec pipeline.
[0086] The first pipeline 210 can include a first stage 211 for encoding the input image / video and a second stage 212 for decoding the encoded image / video to generate a reconstructed image / video. The reconstructed image / video can be used for human viewing, i.e., human vision.
[0087] The second pipeline 220 may include a third stage 221 for extracting features / feature maps from an input image / video, a fourth stage 222 for encoding the extracted features / feature maps, and a fifth stage 223 for decoding the encoded features / feature maps to generate reconstructed features / feature maps. The reconstructed features / feature maps may be used for machine (vision) tasks. Here, a machine (vision) task may refer to a task in which an image / video is consumed by a machine. Machine (vision) tasks may be applied to service scenarios such as, for example, surveillance, intelligent transportation, smart cities, smart industries, smart content, etc. Depending on the implementation, the reconstructed features / feature maps may be used for human vision.
[0088] Depending on the implementation, the features / feature maps encoded in the fourth stage 222 may be transmitted to the first stage 221 and used for encoding the image / video. In this case, an additional bitstream may be generated based on the encoded features / feature maps, and the generated additional bitstream may be transmitted to the second stage 222 and used for decoding the image / video.
[0089] Depending on the implementation, the features / feature maps decoded in the fifth stage 223 may be transmitted to the second stage 222 and used for decoding the image / video.
[0090] Figure 2 The case where the VCM pipeline 200 is shown to include the first pipeline 210 and the second pipeline 220 is merely an example and the embodiments of the present disclosure are not limited thereto. For example, the VCM pipeline 200 may include only the second pipeline 220, or the second pipeline 220 may be extended to multiple feature codec pipelines.
[0091] Meanwhile, in the first pipeline 210, the first stage 211 may be performed by an image / video encoder, and the second stage 212 may be performed by an image / video decoder. Additionally, in the second pipeline 220, the third stage 221 may be performed by a VCM encoder (or a features / feature maps encoder), and the fourth stage 222 may be performed by a VCM decoder (or a features / feature maps encoder). The encoder / decoder structures will be described in detail below.
[0092] Encoder
[0093] Figure 3 is a diagram schematically showing an image / video encoder to which embodiments of the present disclosure are applicable.
[0094] Reference Figure 3, the image / video encoder 300 may further include an image splitter 310, a predictor 320, a residual processor 330, an entropy encoder 340, an adder 350, a filter 360, and a memory 370. The predictor 320 may include an inter-predictor 321 and an intra-predictor 322. The residual processor 330 may include a transformer 332, a quantizer 333, an inverse quantizer 334, and an inverse transformer 335. The residual processor 330 may further include a subtractor 331. The adder 350 may be referred to as a reconstructor or a reconstruction block generator. Depending on the implementation, the image splitter 310, the predictor 320, the residual processor 330, the entropy encoder 340, the adder 350, and the filter 360 may be configured by one or more hardware components (e.g., an encoder chipset or a processor). Additionally, the memory 370 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The above hardware components may further include the memory 370 as an internal / external component.
[0095] The image splitter 310 may split an input image (or picture, frame) input to the image / video encoder 300 into one or more processing units. As an example, the processing unit may be referred to as a coding unit (CU). The coding unit may be recursively split from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, one coding unit may be split into multiple coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or the ternary tree structure may be applied later. Alternatively, the binary tree structure may be applied first. The image / video coding process according to the present disclosure may be performed based on the final coding unit that is no longer split. In this case, the largest coding unit may be used as the final coding unit based on the coding efficiency according to the image characteristics, or if necessary, the coding unit may be recursively split into coding units with a deeper depth to use the coding unit with the optimal size as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, both the prediction unit and the transformation unit may be divided or split from the above final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.
[0096] In some cases, the unit can be used interchangeably with terms such as block or region. In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. Samples can typically represent pixels or pixel values, and can represent only the pixels / pixel values of the luminance component, or only the pixels / pixel values of the chrominance component. Samples can be used as a term corresponding to pixels or picture elements.
[0097] The image / video encoder 300 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 321 or the intra-frame predictor 322 from the input image signal (original block, original sample array), and send the generated residual signal to the transformer 332. In this case, as shown, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, prediction sample array) within the image / video encoder 300 can be called the subtractor 331. The predictor can perform prediction on the processing target block (hereinafter referred to as the current block) and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction in the current block or CU unit. The predictor can generate various prediction-related information such as prediction mode information and transmit it to the entropy encoder 340. The information about the prediction can be encoded in the entropy encoder 340 and output in the form of a bitstream.
[0098] The intra-frame predictor 322 can predict the current block by referring to samples in the current picture. At this time, depending on the prediction mode, the samples referred to can be located in the neighbors of the current block or can be located at a position far from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The non-directional modes can include, for example, the DC mode and the planar mode. Depending on the level of detail of the prediction direction, the directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes can be used depending on the settings. The intra-frame predictor 322 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0099] The inter - frame predictor 321 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter - frame predictor 321 can construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter - frame predictor 321 can use the motion information of neighboring blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of a neighboring block can be used as a motion vector predictor, and the motion vector difference can be signaled to indicate the motion vector of the current block.
[0100] The predictor 320 can generate a prediction signal based on various prediction methods. For example, for the prediction of a block, the predictor can apply not only intra - frame prediction or inter - frame prediction but also both intra - frame prediction and inter - frame prediction simultaneously. This can be referred to as combined inter - frame and intra - frame prediction (CIIP). Additionally, the predictor can be based on the intra - block copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter - frame prediction because a reference block is derived within the current picture. That is, IBC can use at least one of the inter - frame prediction techniques described in the present disclosure. The palette mode can be regarded as an example of intra - frame coding or intra - frame prediction. When the palette mode is applied, the sample values within the picture can be signaled based on information about the palette table and the palette index.
[0101] The prediction signal generated by the predictor 320 can be used to generate a reconstructed signal or to generate a residual signal. The transformer 332 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by the graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Additionally, the transform processing can be applied to square pixel blocks of the same size, or can be applied to non-square blocks of variable size.
[0102] Quantizer 333 may quantize the transform coefficients and send them to entropy encoder 340. Entropy encoder 340 may encode the quantized signal (information regarding the quantized transform coefficients) and output a bitstream. The information regarding the quantized transform coefficients may be referred to as residual information. Quantizer 333 may reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and generate information regarding the quantized transform coefficients based on the quantized transform coefficients in one-dimensional vector form. Entropy encoder 340 may perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. Entropy encoder 340 may encode, together or separately, information necessary for video / image reconstruction in addition to the quantized transform coefficients (e.g., values of syntax elements, etc.). The encoded information (e.g., encoded video / image information) may be sent or stored in the form of a bitstream in units of network abstraction layer (NAL). The video / image information may further include information regarding various parameter sets such as adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). Additionally, the video / image information may further include general constraint information. Additionally, the video / image information may further include methods, purposes, etc. for generating and using the encoded information. In the present disclosure, the information and / or syntax elements transmitted / signaled from the image / video encoder to the image / video decoder may be included in the image / video information. The image / video information may be encoded through the above encoding process and included in the bitstream. The bitstream may be sent through a network or may be stored in a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for sending the signal output from entropy encoder 340 and / or a storage unit (not shown) for storing the signal may be configured as an internal / external element of image / video encoder 300, or the transmitter may be included in entropy encoder 340.
[0103] The quantized transform coefficients output from the quantizer 333 can be used to generate a prediction signal. For example, the quantized transform coefficients can be dequantized and inverse-transformed by the dequantizer 334 and the inverse-transformer 335 to reconstruct the residual signal (residual block or residual samples). The adder 350 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 321 or the intra-frame predictor 322 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). In the case where there is no residual for the target block to be processed, such as when the skip mode is applied, the predicted block can be used as the reconstructed block. The adder 350 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next target block in the current picture, and can be used for inter-frame prediction of the next picture through filtering as described below.
[0104] Meanwhile, Luminance Mapping and Chrominance Scaling (LMCS) is applicable during picture encoding and / or reconstruction.
[0105] The filter 360 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 360 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 370, specifically, in the DPB of the memory 370. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 360 can generate various information related to filtering and send the generated information to the entropy encoder 190. The information related to filtering can be encoded by the entropy encoder 340 and output in the form of a bitstream.
[0106] The modified reconstructed picture sent to the memory 370 can be used as a reference picture in the inter-frame predictor 321. Through this, the prediction mismatch between the encoder and the decoder can be avoided and the coding efficiency can be improved.
[0107] The DPB of the memory 370 can store the modified reconstructed picture to be used as a reference picture in the inter-frame predictor 321. The memory 370 can store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the block in the already reconstructed picture. The stored motion information can be transmitted to the inter-frame predictor 321 to be used as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 370 can store the reconstructed samples of the reconstructed blocks in the current picture and can transmit the stored reconstructed samples to the intra-frame predictor 322.
[0108] Meanwhile, the VCM encoder (or feature / feature map encoder) basically performs a series of processes such as prediction, transformation, and quantization to encode the feature / feature map, and thus can basically have the same as the referenceFigure 3 The VCM encoder has the same / similar structure as the described image / video encoder 300. However, the VCM encoder is different from the image / video encoder 300 in that the feature / feature map is the encoding target, and thus may be different from the image / video encoder 300 in the name of each unit (or component) (e.g., image segmenter 310, etc.) and its specific operation content. The specific operation of the VCM encoder will be described in detail later.
[0109] Decoder
[0110] Figure 4 FIG. is a diagram schematically showing an image / video decoder to which embodiments of the present disclosure are applicable.
[0111] Reference Figure 4 , the image / video decoder 400 may include an entropy decoder 410, a residual processor 420, a predictor 430, an adder 440, a filter 450, and a memory 460. The predictor 430 may include an inter-frame predictor 431 and an intra-frame predictor 432. The residual processor 420 may include a dequantizer 421 and an inverse transformer 422. Depending on the embodiment, the entropy decoder 410, the residual processor 420, the predictor 430, the adder 440, and the filter 450 may be configured by one hardware component (e.g., a decoder). Chipset or processor). In addition, the memory 460 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory 460 as an internal / external component.
[0112] When receiving a bitstream containing video / image information, the image / video decoder 400 may reconstruct the image / video corresponding to the process of processing the image / video information in the Figure 3 image / video encoder 300. For example, the image / video decoder 400 may derive units / blocks based on the block segmentation-related information obtained from the bitstream. The image / video decoder 400 may perform decoding using the processing units applied in the image / video encoder. Accordingly, the decoding processing unit may be, for example, an encoding unit, and the encoding unit may be segmented from the coding tree unit or the maximum coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. In addition, the reconstructed image signal decoded and output by the image / video decoder 400 may be played by a playback device.
[0113] The image / video decoder 400 may receive, in the form of a bitstream, from Figure 3The signal output by the encoder, and the received signal is decoded by the entropy decoder 410. For example, the entropy decoder 410 can parse the bitstream to derive information necessary for image reconstruction (or picture reconstruction) (e.g., image / video information). This image / video information can further include information about various parameter sets, such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Additionally, the image / video information can further include general constraint information. Additionally, the image / video information can include methods, purposes, etc. for generating and using the decoded information. The image / video decoder 400 can further decode the picture based on the information about the parameter set and / or the general constraint information. The information and / or syntax elements signaled / received can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 410 can decode the information in the bitstream based on coding methods such as Exponential Golomb coding, CAVLC, or CABAC, and output the values of the syntax elements necessary for image reconstruction and the quantization values of the transform coefficients related to the residuals. More specifically, in the CABAC entropy decoding method, bins corresponding to each syntax element can be received in the bitstream, a context model can be determined using the decoding target syntax element information and the decoding information of neighboring blocks and the decoding target block or the information about the symbols / bins decoded in the previous step, the occurrence probability of the bin can be predicted according to the determined context model, and the bin can be arithmetically decoded to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can use the information about the decoded symbols / bins for the context model of the next symbol / bin to update the context model after determining the context model. Among the information decoded in the entropy decoder 410, the information about prediction is provided to the predictors (inter-frame predictor 432 and intra-frame predictor 431), and the residual values obtained by performing entropy decoding in the entropy decoder 410, that is, the quantized transform coefficients and the related parameter information can be input to the residual processor 420. The residual processor 420 can derive a residual signal (residual block, residual sample, residual sample array). Additionally, the information about filtering among the information decoded by the entropy decoder 410 can be provided to the filter 450. Meanwhile, a receiver (not shown) that receives the signal output by the image / video encoder can be further configured as an internal / external component of the image / video decoder 400, or the receiver can be a component of the entropy decoder 410. Meanwhile, the image / video decoder according to the present disclosure can be referred to as an image / video decoding device, and the image / video decoder can be divided into an information decoder (image / video information decoder) and a sample decoder (image / video sample decoder).In this case, the information decoder may include an entropy decoder 410, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 440, a filter 450, a memory 460, an inter-frame predictor 432, or an intra-frame predictor 431.
[0114] The dequantizer 421 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 421 may rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scan order executed in the image / video encoder. The dequantizer 321 may dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.
[0115] The inverse transformer 422 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0116] The predictor 430 may perform prediction on a current block and generate a prediction block including prediction samples for the current block. The predictor may determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from the entropy decoder 410, and may determine a specific intra-frame / inter-frame prediction mode.
[0117] The predictor 420 may generate a prediction signal based on various prediction methods. For example, the predictor may apply not only intra-frame prediction or inter-frame prediction to the prediction of a block, but also apply intra-frame prediction and inter-frame prediction simultaneously. This may be referred to as combined inter-frame and intra-frame prediction (CIIP). Additionally, the predictor may be based on an intra-block copy (IBC) prediction mode or a palette mode for the prediction of a block. The IBC prediction mode or the palette mode may be used, for example, for image / video coding of content such as games, such as screen content coding (SCC). In IBC, the prediction is basically performed within the current picture, but may be performed similarly to inter-frame prediction because a reference block is derived within the current picture. That is, IBC may use at least one of the inter-frame prediction techniques described in this document. The palette mode may be regarded as an example of intra-frame coding or intra-frame prediction. When the palette mode is applied, information about the palette table and palette index may be included in the image / video information and signaled.
[0118] The intra-frame predictor 431 may predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference samples may be located in the neighbors of the current block, or may be located at a position far from the current block. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. The intra-frame predictor 431 may use the prediction mode applied to an adjacent block to determine the prediction mode applied to the current block.
[0119] The inter-frame predictor 432 may derive a prediction block for a current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, the adjacent blocks may include spatial adjacent blocks present in the current picture and temporal adjacent blocks present in the reference picture. For example, the inter-frame predictor 432 may construct a motion information candidate list based on adjacent blocks, and derive a motion vector and / or a reference picture index of the current block based on the received candidate selection information. The inter-frame prediction may be performed based on various prediction modes, and the information about the prediction may include information indicating the mode of the inter-frame prediction for the current block.
[0120] The adder 440 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (prediction block, prediction sample array) output from a predictor (including the inter-frame predictor 432 and / or the intra-frame predictor 431). If there is no residual for the target block to be processed, such as when the skip mode is applied, the prediction block may be used as the reconstructed block.
[0121] The adder 440 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next target block to be processed in the current picture, may be output after filtering as described later, or may be used for inter-frame prediction of the next picture.
[0122] Meanwhile, luminance mapping and chrominance scaling (LMCS) are applicable during picture decoding.
[0123] The filter 450 is capable of improving the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 450 may generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and send the modified reconstructed picture to the memory 460, specifically to the DPB of the memory 460. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0124] The (modified) reconstructed picture stored in the DPB of the memory 460 can be used as a reference picture in the inter - frame predictor 432. The memory 460 can store the motion information of the blocks from which the motion information in the current picture is derived (decoded) and / or the motion information of the blocks in the already reconstructed pictures. The stored motion information can be transmitted to the inter - frame predictor 432 to be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 460 can store the reconstructed samples of the reconstructed blocks in the current picture and transmit them to the intra - frame predictor 431.
[0125] Meanwhile, the VCM decoder (or feature / feature - map decoder) performs a series of processes such as prediction, inverse transformation, and de - quantization to decode the feature / feature - map, and can basically have the same / similar structure as the image / video decoder 400 described above with reference to Figure 4 However, the VCM decoder is different from the image / video decoder 400 in that the feature / feature - map is the decoding target, and may be different from the image / video decoder 400 in the name of each unit (or component) (e.g., DPB, etc.) and its specific operations. The operations of the VCM decoder can correspond to the operations of the VCM encoder, and the specific operations will be described in detail later.
[0126] Feature / feature map encoding process
[0127] Figure 5 is a flowchart schematically illustrating the feature / feature - map encoding process applicable to the embodiments of the present disclosure.
[0128] Reference Figure 5 shows that the feature / feature - map encoding process may include a prediction process (S510), a residual processing process (S520), and an information encoding process (S530).
[0129] The prediction process (S510) can be performed by the predictor 320 described above with reference to Figure 3 description.
[0130] Specifically, the intra predictor 322 may predict a current block (i.e., a set of currently encoded target feature elements) by referring to feature elements in a current feature / feature map. Intra prediction may be performed based on the spatial similarity of the feature elements constituting the feature / feature map. For example, feature elements included in the same region of interest (RoI) within an image / video may be estimated to have similar data distribution characteristics. Accordingly, the intra predictor 322 may predict the current block by referring to the feature elements that have been reconstructed within the RoI including the current block. At this time, depending on the prediction mode, the referred feature elements may be located at positions adjacent to the current block or may be located at positions far from the current block. The intra prediction modes for feature / feature map encoding may include a plurality of non-directional prediction modes and a plurality of directional prediction modes. The non-directional prediction modes may include, for example, prediction modes corresponding to the DC mode and the planar mode of the image / video encoding process. Additionally, the directional modes may include prediction modes corresponding to, for example, 33 directional modes or 65 directional modes of the image / video encoding process. However, this is an example, and the types and numbers of intra prediction modes may be set / changed in various ways depending on the implementation.
[0131] The inter-frame predictor 321 may predict a current block based on a reference block (i.e., a set of reference feature elements) specified by motion information regarding a reference feature / feature map. Inter-frame prediction may be performed based on the temporal similarity of the feature elements constituting the feature / feature map. For example, temporally consecutive features may have similar data distribution characteristics. Thus, the inter-frame predictor 321 may predict the current block by referring to the already reconstructed feature elements of the feature that is temporally adjacent to the current feature. At this time, the motion information for specifying the reference feature elements may include a motion vector and a reference feature / feature map index. The motion information may further include information regarding the inter-frame prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing within the current feature / feature map and temporally neighboring blocks existing within the reference feature / feature map. The reference feature / feature map including the reference block and the reference feature / feature map including the temporally neighboring blocks may be the same or different. The temporally neighboring blocks may be referred to as collocated reference blocks, etc., and the reference feature / feature map including the temporally neighboring blocks may be referred to as a collocated feature / feature map. The inter-frame predictor 321 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or the reference feature / feature map index of the current block. Inter-frame prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter-frame predictor 321 may use the motion information of neighboring blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be sent. In the case of the motion vector prediction (MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector of the current block may be indicated by signaling a motion vector difference. In addition to the above intra-frame prediction and inter-frame prediction, the predictor 320 may also generate a prediction signal based on various prediction methods.
[0132] The prediction signal generated by the predictor 320 may be used to generate a residual signal (residual block, residual feature element) (S520). The residual processing procedure (S520) may be performed by the residual processor 330 described above with reference to Figure 3 In addition, (quantized) transform coefficients may be generated through a transform and / or quantization process for the residual signal, and the entropy encoder 340 may encode information regarding the (quantized) transform coefficients in the bitstream as residual information (S530). In addition, in addition to the residual information, the entropy encoder 340 may also encode information necessary for feature / feature map reconstruction, such as prediction information (e.g., prediction mode information, motion information, etc.) in the bitstream.
[0133] Meanwhile, the feature / feature map encoding process may further include not only a process (S530) for encoding information for feature / feature map reconstruction (e.g., prediction information, residual information, segmentation information, etc.) and outputting it in the form of a bitstream, a process for generating a reconstructed feature / feature map for the current feature / feature map, and an optional process for applying loop filtering to the reconstructed feature / feature map.
[0134] The VCM encoder can derive (modified) residual features from the quantized transform coefficients through dequantization and inverse transformation, and generate a reconstructed feature / feature map based on the predicted features output in step S510 and the (modified) residual features. The reconstructed feature / feature map generated in this way can be the same as the reconstructed feature / feature map generated in the VCM decoder. When performing the loop filtering process on the reconstructed feature / feature map, a modified reconstructed feature / feature map can be generated by performing the loop filtering process on the reconstructed feature / feature map. The modified reconstructed feature / feature map can be stored in the decoded feature buffer (DFB) or memory and used as a reference feature / feature map in a later feature / feature map prediction process. Additionally, (loop) filtering-related information (parameters) can be encoded and output in the form of a bitstream. Through the loop filtering process, noise that may occur during feature / feature map encoding can be removed, and the task performance based on the feature / feature map can be improved. Additionally, by performing the loop filtering process in both the encoder stage and the decoder stage, the identity of the prediction results can be guaranteed, the reliability of feature / feature map encoding can be improved, and the data transmission volume for feature / feature map encoding can be reduced.
[0135] Feature / feature map decoding process
[0136] Figure 6 is a flowchart schematically illustrating a feature / feature map decoding process to which embodiments of the present disclosure are applicable.
[0137] Reference Figure 6, The feature / feature map decoding process may include an image / video information acquisition process (S610), a feature / feature map reconstruction process (S620 to S640), and a loop filtering process (S650) for the feature / feature map to be reconstructed. The feature / feature map reconstruction process may be performed on the prediction signal and the residual signal obtained through the inter-frame / intra-frame prediction (S620) and residual processing (S630) described in the present disclosure, and the dequantization and inverse transformation processes for the quantized transform coefficients described in the present disclosure. The modified reconstructed feature / feature map may be generated through the loop filtering process for the reconstructed feature / feature map, and the modified reconstructed feature / feature map may be output as the decoded feature / feature map. The decoded feature / feature map may be stored in a decoded feature buffer (DFB) or a memory and used as a reference feature / feature map in the inter-frame prediction process when decoding the feature / feature map. In some cases, the above loop filtering process may be omitted. In this case, the reconstructed feature / feature map may be output as the decoded feature / feature map without change, stored in the decoded feature buffer (DFB) or the memory, and then used as a reference feature / feature map in the inter-frame prediction process when decoding the feature / feature map.
[0138] Feature extraction method and data distribution characteristics
[0139] Embodiments of the present disclosure propose a prediction process and a method for generating a related bitstream required for compressing activation (feature) maps generated in hidden layers of a deep neural network.
[0140] The input data input to the deep neural network undergoes computational processing of several hidden layers, and according to the type of the deep neural network used and the position of the hidden layer in the deep neural network, the computational results of each hidden layer are output as feature / feature maps with various sizes and numbers of channels.
[0141] Figure 7 is a diagram illustrating an example of a feature extraction and reconstruction method to which embodiments of the present disclosure are applicable.
[0142] Reference Figure 7 , the feature extraction network 710 may extract the intermediate layer activation (feature) map of the deep neural network from the source image / video and output the extracted feature map. The feature extraction network 710 may be a set of consecutive hidden layers from the input of the deep neural network.
[0143] The encoding device 720 may compress the output feature map and output the compressed feature map in the form of a bitstream, and the decoding device 730 may reconstruct the (compressed) feature map from the output bitstream. The encoding device 720 may correspond to Figure 1 the encoder 12, and the decoding device 720 may correspond toFigure 1 The decoder 22. The task network 740 may perform a task based on the reconstructed feature map.
[0144] The number of channels of the feature map that is the compression target of the VCM may vary according to the network used for feature extraction and the extraction location, and may be greater than the number of channels of the input data.
[0145] Figure 8 is a diagram illustrating an example of an image splitting method to which the embodiments of the present disclosure are applicable. As an example, Figure 8 is a diagram illustrating CTUs, slices, and tiles in an image.
[0146] The video / image encoding method according to this document may be performed based on the following segmentation structure. It may be based on the CTUs, CUs (and / or TUs or PUs) derived from the segmentation structure according to Figure 8 to perform processes such as prediction, residual processing ((inverse) transformation, (inverse) quantization, etc.), syntax element encoding, filtering, etc. The block segmentation process may be performed in the encoding device, and the segmentation-related information may be encoded and sent to the decoding device in the form of a bitstream. The decoding device may derive the block segmentation structure of the current picture based on the segmentation-related information obtained from the bitstream, and perform a series of processes for image decoding (e.g., prediction, residual processing, block / picture reconstruction, loop filtering, etc.) based on the derived block segmentation structure. The size of the CU may be the same as the size of the TU, and multiple TUs may exist in the CU region. In addition, the size of the CU may generally represent the size of the luminance component (samples) CB. The size of the TU may generally represent the size of the luminance component (samples) TB. The size of the chrominance component (samples) CB or TB may be derived based on the size of the luminance component (samples) CB or TB according to the component ratio according to the color format (chrominance format) of the picture / video (e.g., 4:4:4, 4:2:2, 4:2:0, etc.), and the transformation / inverse transformation may be performed based on the TU (TB).
[0147] In addition, in the encoding of video / images according to this document, the image processing unit may have a hierarchical structure. A picture may be split into one or more CUs, and one or more CUs may be grouped and classified into one or more tiles, patches, slices, and / or tile groups. A slice may include one or more patches. A patch may include one or more CTU rows in a tile. A slice may include an integer number of patches of a picture. A tile group may include one or more tiles. A tile may include one or more CTUs. A CTU may be split into one or more CUs. According to the raster scan of tiles in a picture, a tile group may include an integer number of tiles. A slice header may carry information / parameters applicable to the corresponding slice (blocks in the slice). A picture header may carry information / parameters applicable to the corresponding picture (or blocks in the picture). When the encoding / decoding device has a multi-core processor, the encoding / decoding process for tiles, slices, patches, and / or tile groups may be processed in parallel. In this document, a slice or a tile group may be used interchangeably. That is, a tile group header may be referred to as a slice header. Here, a slice may have one of the slice types including I slice, P slice, and B slice.
[0148] In the encoding device, the sizes of tiles / tile groups, patches, slices, and the maximum and minimum coding units may be determined according to the characteristics of the video image (e.g., resolution) or considering the efficiency of encoding or parallel processing, and information about them may be included in the bitstream or information from which they can be derived may be included.
[0149] In the decoding device, information indicating whether the CTUs in the tiles / tile groups, patches, slices, or tiles of the current picture have been split into multiple coding units may be obtained. The efficiency may be improved by obtaining (sending) such information only under specific conditions.
[0150] A slice header (slice header syntax) may include information / parameters that can be commonly applied to a slice. APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more pictures. SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. DPS (DPS syntax) may include information / parameters that can be commonly applied to the entire video. DPS may include information / parameters related to the concatenation of the encoded video sequence (CVS).
[0151] In this document, the upper-layer syntax may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.
[0152] In addition, for example, information regarding the splitting, configuration, etc. of tiles / tile groups / patches / slices can be configured during the encoding phase through upper-layer syntax and sent to the decoding device in the form of a bitstream.
[0153] Figure 9 is a diagram illustrating an example of a VCM image encoding / decoding system, and Figure 10 is a diagram illustrating another example of a VCM image encoding / decoding system. As an example, in Figure 9 and Figure 10 the image encoding system (e.g., Figure 1 ) can be extended / redesigned such that, according to the requests, purposes, and surrounding environments of users or machines, only a part of the video source can be used, or necessary parts / information can be used by obtaining them from the video source. That is, Figure 9 and Figure 10 can be related to VCM.
[0154] VCM can refer to encoding / decoding the entire image and / or a part of the image and / or necessary information (features) obtained from the image according to the requests, purposes, and surrounding environments of users and / or machines. In VCM, the encoding target can be the image itself, or it can be information called features extracted from the image according to the requests, purposes, and surrounding environments of users and / or machines, and can be a collection of a series of information over time.
[0155] Referring to Figure 9 , the VCM system can include an image encoder and an image decoder for VCM. The source device (see Figure 1 ) can send the encoded image information to the receiving device via a storage medium or a network. The users of the device can be humans and / or machines.
[0156] Referring to Figure 10 , the VCM system can include an image encoder and an image decoder for VCM. The source device can send the encoded image information (features) to the receiving device via a storage medium or a network. The users of the device can be humans and / or machines.
[0157] As an example, the process of extracting information (i.e., features) from an image can be called feature extraction. Feature extraction can be performed by both video / image capture devices and / or video / image generation devices. Features can be information extracted / processed from an image according to the requests, purposes, or surrounding environments of users and / or machines, and can be a collection of a series of information over time.
[0158] An image encoder for VCM can perform a series of processes, such as prediction, transformation, quantization, etc., for the compression and encoding efficiency of the entire image and / or a part of the image and / or features. The encoded data can be output in the form of a bitstream.
[0159] An image decoder for VCM can perform a series of processes (such as inverse quantization, inverse transformation, prediction, etc.) corresponding to the operations of an encoding device (i.e., an encoder) to decode the video / image.
[0160] The decoded image and / or features can be rendered. Additionally, the decoded image and / or features can be used to perform tasks of a user or a machine. Examples of tasks can include artificial intelligence (AI), computer vision tasks, etc., such as face recognition, behavior recognition, lane recognition, etc.
[0161] This disclosure presents various embodiments regarding obtaining and encoding all and / or a part of an image for VCM, and these embodiments can be performed in combination with each other unless otherwise specified. The methods / embodiments of this disclosure can be applied to the methods disclosed in the VCM standard.
[0162] VCM hierarchical structure
[0163] VCM can be based on a hierarchical structure composed of a feature encoding layer, a neural network (feature) abstraction layer, and a feature extraction layer. Figure 11 is a diagram illustrating an example of the VCM hierarchical structure, and Figure 12 is a diagram illustrating an example of a VCM bitstream composed of encoded abstract features and neural network abstraction layer (NNAL) information.
[0164] As an example, referring to Figure 11 , the VCM hierarchical structure can be composed of a feature extraction layer 1110, a neural network (feature) abstraction layer 1120, and a feature encoding layer 1130.
[0165] The feature extraction layer 1110 can be a layer for extracting features from an input source and can also include the extraction results. The feature encoding layer 1130 can be a layer for compressing the extracted features and can also include the compression results.
[0166] The neural network abstraction layer 1120 can abstract the information generated in the feature extraction layer 1110 (e.g., information about the extracted features / feature maps) and transmit the abstracted information to the feature encoding layer 1130. The neural network abstraction layer 1120 can hide the interior of the feature extraction layer 1110 through the abstraction of information and provide a consistent feature interface function. Thus, even when the compression target changes due to changes in tools (e.g., convolutional neural network (CNN), deep neural network (DNN), etc.), the feature encoding layer 1130 can perform a consistent feature encoding process. In the present disclosure, the neural network abstraction layer (NNAL) can be referred to as the feature abstraction layer.
[0167] The interfaces between the feature extraction layer 1110 and the neural network abstraction layer 1120 and between the feature encoding layer 1130 and the neural network abstraction layer 1120 can be predefined, and the operations in the neural network abstraction layer 1120 can be set to be changed later.
[0168] Reference Figure 12 , the bitstream configured as illustrated can be referred to as a neural network abstraction layer (NNAL) unit. The NNAL unit can be an independent feature reconstruction unit. The input features for one NNAL unit can be extracted from the same layer in the neural network. Thus, the input features for one NNAL unit can be forced to have the same characteristics. For example, the same feature extraction method can be applied to the input features for one NNAL unit.
[0169] The NNAL unit can include an NNAL unit header and an NNAL unit payload. The NNAL unit header can include all the information required to use the encoded features according to the task. The NNAL unit payload can include the abstracted feature information. The NNAL unit payload can include a group header and group data. The group header can include the configuration information of the feature group data, such as information about the temporal order, quantity, and common nature of the feature channels that make up the feature group. The feature channel can refer to the encoded feature unit. The group data can include multiple feature channels and a coding demonstrative, and each feature channel can include type information, prediction information, side information, and residual information. In this case, the type information can represent the coding method, and the prediction information can represent the prediction method. Additionally, the side information can include additional information required for decoding (e.g., entropy coding, quantization-related information, etc.), and the residual information can include information about the encoded feature elements (i.e., a set of eigenvalue information).
[0170] In addition, the image encoding for VCM can be based on a neural network, and in the neural network-based image encoding technology, the memory required for the convolution process (e.g., the increase in the size of the feature map of the intermediate layer) can increase more than the size of the input source (i.e., the input image). Therefore, the size of the peak memory of the encoding / decoding process can increase.
[0171] This application proposes an image encoding technology that can solve the above problems. By splitting the input source and encoding / decoding the split input source to perform encoding / decoding using limited resources, it can reduce the size of the peak memory.
[0172] Therefore, this application proposes the number of split images and the splitting method to reconstruct the split regions into their original form. According to this application, the unit of splitting (i.e., the unit of the processing unit) is not fixed but can vary according to the available resources of the device. That is, according to this application, an image encoding technology using adaptive splitting of the image to the available resources of the device can be provided. According to an embodiment of this application, when there are more resources, splitting processing may not be required, while when the available resources are less, the number of split regions may need to be increased.
[0173] In addition, the embodiments described in this disclosure may relate to Figure 3 the splitting process of the input data of the image splitter 310 of the encoder and the method of representing the input data for processing. The image splitter 310 can be used to perform the splitting process of the input data. According to an embodiment of this disclosure, the input data of the image splitter 310 can be split according to resource constraints or service purposes. In addition, information for reconstructing the split input data can be defined in Figure 11 and Figure 12 the encoding layers and structures (including the feature encoding layer, the neural network (feature) abstraction layer, the feature extraction layer, etc.). According to this application, the size of the peak memory can be reduced.
[0174] Table 1 below can show an example of the size of the features of each block of Resnet, which is a representative CNN. More specifically, Table 1 is a table for describing an example of the structure of Resnet as a CNN, the input / output size, and the size of the features. The input image consists of 3 RGB channels with a width of 244 and a height of 244, and its size (i.e., width * height) is assumed to be 150528.
[0175] [Table 1]
[0176]
[0177] According to Table 1, during the CNN process, the size of the output features of each block can be larger than the input image. For example, Convolutional Block 1 is a single convolution, and its size is 802816 (which is 5.6 times the size of the input image). Generally, as the layer gets deeper, the size of the output of the block decreases, and the size of the fully connected (FC) block as the final result can be smaller than the size of the input image. However, the size of the output of the block may increase during the intermediate process and may require a lot of resources (e.g., memory) during processing. According to the embodiments described in the present disclosure, data splitting processing can be performed to perform neural network-based encoding / decoding using limited resources (e.g., memory).
[0178] According to a conventional video codec, profiles, tiers, levels (PTL) can be defined for appropriate designs and applications according to its purpose (i.e., service type) and application environment (i.e., device performance, available bandwidth, etc.). The profile can define the encoding tools (e.g., algorithms) used. Additionally, the profile can represent constraints on the encoding tools (e.g., algorithms) used, and the level can define constraints on the available memory and buffer sizes (such as the maximum number of applicable pixels, the maximum number of pixels per second, the maximum size of the encoded bitstream per second, etc.). Similar to a conventional video codec, a neural network-based video codec (such as VCM) may also need to perform information for the role of PTL. Through such information, the amount of resources allowed during the decoding process can be known.
[0179] The peak memory during the decoding process can be determined according to the size of the input data. To process the input data smoothly, preferably, the size of the peak memory of the decoder determined according to the size of the input data does not exceed the maximum amount of resources that can be known through the PTL information. That is, the maximum processing unit value can be defined through the PTL information. When the size of the input data is larger than the size of the maximum processing unit, the input data can be processed only when split into segments smaller than the size of the maximum processing unit.
[0180] Figure 13 and Figure 14 is a diagram illustrating an example of the input data splitting processing procedure according to an embodiment of the present disclosure. More specifically, Figure 13 can illustrate an example of the operation of an encoder according to the present disclosure, and Figure 14 can illustrate an example of the operation of a decoder according to the present disclosure.
[0181] According to Figure 13The image encoding method according to an example of the present disclosure as shown can determine the PTL (S1310), can determine the size of the processing unit (S1320), can determine the number of splits and the splitting method (S1330), and can generate splitting information (S1340).
[0182] According to Figure 14 The image decoding method according to an example of the present disclosure as shown can determine the PTL (S1410), and can derive the splitting information (S1420).
[0183] Hereinafter, each operation will be described in detail.
[0184] First, according to Figure 13 the example of the present disclosure as shown, the encoder can determine the PTL (S1310), and during this process, can determine information about resources (resource information, for example, information about a random access memory (RAM), a graphics processing unit (GPU), a central processing unit (CPU), etc.) (which is device-specific information) and all PTL information according to services, purposes, etc. Additionally, during this process, the corresponding information can be determined considering the resources of the decoder.
[0185] Thereafter, the size of the processing unit can be determined (S1320). The size of the processing unit can be determined based on the determined profile level information and resource information. Hereinafter, an embodiment of the method for determining the size of the processing unit will be described in detail.
[0186] As an example, the size of the processing unit can be based on the size of the maximum processing unit and can be based on the size of the maximum processing unit defined in the PTL. For example, the size of the maximum processing unit can be defined in the PTL. Based on the PTL, the maximum input resolution and the size of the required memory are determined. Additionally, the size of the maximum processing unit can be defined. Table 2 below can show an example of defining the maximum processing unit (processing unit) in the PTL.
[0187] [Table 2]
[0188]
[0189] In Table 2, the type of level, the value of the size of the largest picture, and the value of the size of the largest processing unit are examples, which may mean that the corresponding information can be defined in the PTL information (i.e., profile / hierarchy / level information), and the size of the largest luminance picture may refer to the size of the largest input data. In this case, the size of the largest processing unit may be included in the PTL information and thus may not be encoded as separate information. This is because the size of the largest processing unit is predefined in the PTL. For example, when the size of the predefined largest processing unit is stored, the encoder may indicate the size of the largest processing unit included in the PTL (e.g., level) information as an index, and the decoder may derive the size of the largest processing unit using the corresponding index. The width and height of the processing units within the size of the largest processing unit may be defined together with the size of the largest processing unit or may be defined independently of the size of the largest processing unit. For example, the height may be determined discretely or within a continuous range with respect to the width and height of the processing unit, and the width or height of the processing unit may be determined independently as long as it is within the size of the largest processing unit.
[0190] As another example, the size of the largest processing unit may be determined according to the required resource information. As an example, the size of the largest processing unit may be derived from the resource information. In addition, for example, when the requirements for the size of the largest processing unit are not clearly defined, the amount of required resources may be calculated based on given information (e.g., the size of the largest buffer, the maximum resolution, and the maximum frame rate). Once the amount of required resources is derived, the size of the largest processing unit may be derived based on the derived amount. According to the present embodiment, the size of the largest processing unit may not be explicitly defined in the PTL, but the size of the largest processing unit may be derived by applying predefined information and methods. For example, when the minimum memory requirement for driving the encoder / decoder is defined as 8GB, the size of the largest processing unit may be derived such that the peak memory of the encoder / decoder does not exceed 8GB.
[0191] According to the above embodiment, since the size of the processing unit is derived in a predefined manner at both the encoder and the decoder, the size of the processing unit may not be encoded because the derived size of the processing unit is applied to the encoder / decoder in the same way, and thus the decoder may derive information about the processing unit without having separate signaling.
[0192] In addition, according to another example, since the size of the processing unit can be determined based on the resources of the encoder / decoder and / or the service purpose, in this case, the size or shape of the processing unit can be explicitly signaled. However, as an example, even in this case, the settable range of the processing unit may be limited. For example, the size of the maximum processing unit can be limited according to the minimum memory requirement. Alternatively, the size of the processing unit can be defined as less than the size of the minimum memory requirement, but can be no greater than the size of the minimum memory requirement. Each of Tables 3 to 8 below shows an example of the definition of the processing unit. First, Table 3 is as follows.
[0193] [Table 3]
[0194]
[0195] The syntax element processing_unit_info_present_flag (e.g., the first syntax) can be processing unit information presence / absence information indicating the presence of information about the processing unit, and can be signaled to the bitstream. As an example, when the value of the first syntax element is a first value (e.g., 1), it can mean that processing unit information (e.g., it can include information about the size of the maximum processing unit) is present, and when the value of the first syntax element is a second value (e.g., 0), it can mean that there is no information about the maximum processing unit.
[0196] The syntax element processing_unit_size (e.g., the second syntax) can indicate the size of the processing unit and can be signaled to the bitstream. As an example, the second syntax can indicate the size of the maximum processing unit. The second syntax can be signaled based on the first syntax and is signaled only when the value of the first syntax is a specific value (e.g., 1).
[0197] Table 4 below shows another example of the definition of the processing unit. Table 4 is as follows.
[0198] [Table 4]
[0199]
[0200] The syntax element processing_unit_info_present_flag (e.g., the first syntax) can be processing unit information presence / absence information (flag), and its repeated description will be omitted.
[0201] Syntax elements processing_width (e.g., the third syntax) and processing_height (e.g., the fourth syntax) may indicate the width and height of the processing unit, respectively. As an example, the third syntax and the fourth syntax may be signaled to the bitstream based on the first syntax. For example, when the first syntax is a specific value (e.g., 1), the third syntax and the fourth syntax may be signaled. In addition, Table 3 is written in a way that signals the third syntax first and then signals the fourth syntax, but this is only an embodiment of the present disclosure. Thus, the fourth syntax may be signaled before the third syntax, and the third syntax and the fourth syntax may be signaled as one syntax (e.g., processing_width_height) (this also corresponds to an embodiment of the present disclosure), etc.
[0202] Table 5 below shows another example of the definition of the processing unit. Table 5 is as follows.
[0203] [Table 5]
[0204] Descriptor processing_unit_size u(8)
[0205] Different from Table 3, according to Table 5, the syntax element processing_unit_size (e.g., the second syntax) may be signaled regardless of the first syntax (e.g., processing_unit_info_present_flag). In this case, the first syntax may not be signaled or defined. The repeated description of the second syntax will be omitted.
[0206] Table 6 below shows yet another example of the definition of the processing unit. Table 6 is as follows.
[0207] [Table 6]
[0208] Descriptor processing_width u(4) processing_height u(4)
[0209] Syntax elements processing_width (e.g., the third syntax) and processing_height (e.g., the fourth syntax) may indicate the width and height of the processing unit, respectively, and the repeated description thereof will be omitted. However, the third syntax and the fourth syntax may be signaled regardless of the first syntax (e.g., processing_unit_info_present_flag). In this case, the first syntax may not be signaled or defined.
[0210] In addition, as an example, when information about the splitting quantity and splitting method is derived from information about the splitting quantity and splitting method, based on the above information, information about the processing unit can be derived. And when information about the processing unit is derived and the processing unit is specified, based on the specified processing unit, information about the splitting quantity and splitting method can be derived. That is, when there is no information about the splitting quantity and / or splitting method, that is, only when the value of the information about the splitting quantity of the processing unit is a specific value (for example, 0), information about the processing unit can be signaled. Tables 7 and 8 below illustrate these embodiments of the present disclosure.
[0211] Table 7 below shows another example of the definition of the processing unit. Table 7 is as follows.
[0212] [Table 7]
[0213]
[0214] The syntax element num_processing_unit_info_present_flag (for example, the fifth syntax) is information indicating whether there is information about the splitting quantity of the processing unit, and can be information indicating the presence of the splitting quantity of the processing unit. For example, when the value of the fifth syntax is the first value (for example, 1), it can mean that there is information about the splitting quantity, and when the value of the fifth syntax is the second value (for example, 0), it can mean that there is no information about the splitting quantity.
[0215] The syntax element processing_unit_info_present_flag (for example, the first syntax) can be information indicating the presence / absence of processing unit information (for example, a flag). However, the first syntax can be signaled to the bitstream based on the fifth syntax. As an example, when the fifth syntax is a specific value (for example, 0), the first syntax can be signaled to the bitstream, and when the fifth syntax is another specific value (for example, 1), the first syntax can be not signaled. The repeated description of the first syntax will be omitted.
[0216] The syntax element processing_unit_size (for example, the second syntax) can be signaled to the bitstream based on the first syntax (for example, processing_unit_info_present_flag). As an example, when the first syntax is a specific value (for example, 1), the second syntax can be signaled to the bitstream, and when the first syntax is another specific value (for example, 0), the second syntax can be not signaled.
[0217] Table 8 below shows another example of the definition of the processing unit. Table 8 is as follows.
[0218] [Table 8]
[0219]
[0220] The syntax element num_processing_unit_info_present_flag (e.g., the fifth syntax) can be information indicating whether there is information about the splitting number of the processing unit, and its repeated description will be omitted.
[0221] The syntax element processing_unit_info_present_flag (e.g., the first syntax) can be a processing unit information presence / absence flag, and its repeated description will be omitted. The first syntax can be signaled based on the fifth syntax.
[0222] The syntax elements processing_width (e.g., the third syntax) and processing_height (e.g., the fourth syntax) can respectively indicate the width and height of the processing unit, and their repeated descriptions will be omitted. The third syntax and the fourth syntax can be signaled based on the first syntax.
[0223] When the size of the processing unit is determined, the splitting number and the splitting method can be determined based on the size of the processing unit (S1320). The number of splitting regions (i.e., the splitting number) can also be derived in a predefined manner based on the size of the processing unit, and the splitting number and the splitting method can be determined by information about the splitting number and the splitting method explicitly signaled based on the size of the processing unit.
[0224] As an example, in order to derive the splitting number from other given information (e.g., information about the processing unit, etc.), the value obtained by substantially dividing the size of the input data by the size of the processing unit can be derived as the splitting number. The following Expression 1 shows an example of defining the splitting number.
[0225] [Expression 1]
[0226]
[0227] In Expression 1, the number of processing units (i.e., the number of split regions) can be expressed as an element (e.g., a syntax element) num_processing_unit_minus1 and can indicate a value obtained by subtracting 1 from the actual number of split regions. The number of split regions can be derived by rounding up the value obtained by dividing the size of the input data by the size of the processing unit, e.g., using the value of the ceil() function. As an example, when the processing unit is defined in terms of width units and height units as in Tables 4, 6, and 8, the split unit (i.e., the processing unit) can also be expressed in terms of width units and height units. In this case, the width of the input picture can be divided by the width of the processing unit, and the height of the input picture can be divided by the height of the processing unit. For example, the element num_processing_column_unit_minus1 (e.g., a syntax element) represents the number of splits in the width direction, which can be a value obtained by subtracting 1 from the value (e.g., applying rounding up) obtained by dividing the width of the input image by the width of the processing unit (e.g., processing_width). Additionally, the element num_processing_row_unit_minus1 (e.g., a syntax element) represents the number of splits in the height direction, which can be a value obtained by subtracting 1 from the value (e.g., applying rounding up) obtained by dividing the height of the input image by the height of the processing unit (e.g., processing_height). In this case, the number of split regions can be derived based on the value obtained by dividing the width of the input picture by the width of the processing unit and the value obtained by dividing the height of the input picture by the height of the processing unit. For example, the number of split regions can be derived by multiplying the two values.
[0228] For example, when the size of the input image is 2073600 (1920×1080) and the size of the processing unit is 518400, according to Expression 1, since the total number of processing units is 4, num_processing_unit_minus1 can be 3. This will be described in more detail with reference to Figure 15 More specifically.
[0229] Figure 15 is a diagram for describing an example in which input data is split according to an embodiment of the present disclosure. Figure 15 Example (a) of Figure 15 illustrates an example of splitting input data in the height direction, and Figure 15 Example (b) of Figure 15 illustrates an example of splitting input data in the width direction. Figure 15(b) Examples of splitting the input data in both the width and height directions. As an example, the splitting method can split the input data evenly in the width or height direction according to a predefined method, and the splitting direction can be encoded and signaled. In this case, the number of columns and rows of the processing units (e.g., num_processing_column_unit_minus1, num_processing_row_unit_minus1) can be derived. In this case, the value obtained by multiplying the number of columns of the processing units by the number of rows of the processing units (e.g., (num_processing_column_unit_minus1 + 1)×(num_processing_row_unit_minus + 1)) can be derived as a value larger than the total number of processing units (e.g., num_processing_unit_minus1 + 1). For example, when the total number of processing units (e.g., num_processing_unit_minus1 + 1) is 7, the information about the number of columns of the processing units (e.g., num_processing_column_unit_minus1) and the information about the number of rows of the processing units (e.g., num_processing_row_unit_minus1) can be calculated as 3 and 1 or 1 and 3 respectively, so that the value obtained by multiplying the number of columns by the number of rows (e.g., (num_processing_column_unit_minus1 + 1)×(num_processing_row_unit_minus + 1)) is 8. That is, when the total number of processing units num_processing_unit_minus1 + 1 is a prime number, the input data cannot be split into two-dimensional data, so it can be split into units that are not prime numbers. In this case, the splitting amounts in the width and height directions can be determined by comparing the sizes of the input data in the width and height directions. That is, the larger the width of the input image, the more the input image is split in the width direction. Such a rule can be predefined by the encoder / decoder so that the rule can operate in the same way, and in addition to the above method, various other methods can also be applied.
[0230] As another example, the number of splits of an image and the split method can be defined and explicitly signaled. That is, as described above, information about the number of processing units, information about the number of columns of the processing units, information about the number of rows of the processing units (e.g., num_processing_unit_minus1, num_processing_column_unit_minus1, num_processing_row_unit_minus1), and the value of a syntax element (e.g., split_method) indicating the split method can be signaled to the bitstream. Elements including the syntax element will be described in detail below.
[0231] In addition, each of Tables 9 to 16 below is an example that explicitly defines the number of splits and the split method. First, Table 9 is as follows.
[0232] [Table 9]
[0233] Descriptor num_processing_unit_minus1 u(8)
[0234] The syntax element num_processing_unit_minus1 (e.g., the sixth syntax element) can be information about the number of splits (i.e., the total number of processing units), and when 1 is added to the value of the sixth syntax element, it can mean the actual total number of splits. The sixth syntax element can be signaled independently as in Table 9.
[0235] Table 10 below shows another example that explicitly defines the number of splits and the split method.
[0236] [Table 10]
[0237] Descriptor num_processing_unit_minus1 u(8) Split_method u(2)
[0238] In addition to the above syntax element num_processing_unit_minus1 (e.g., the sixth syntax element), the syntax element split_method (e.g., the seventh syntax element) can also be signaled. The seventh syntax element can indicate the split method of the input data. Figure 16 Examples of split methods according to embodiments of the present disclosure are illustrated. In Figure 16In the example, when the split_method value is the first value (e.g., 0), it may mean that the splitting method along the width direction has been applied to the input image, and when the split_method value is the second value (e.g., 1), it may mean that the splitting method along the height direction has been applied. Additionally, if the value of the seventh syntax element is the third value (e.g., 2), it may mean that the splitting method along both the width direction and the height direction has been applied. That is, the splitting method can be defined according to the value of the seventh syntax element. However, since Figure 16 corresponds to an embodiment of the present disclosure, the splitting method can be defined differently according to the actual defined value, and the seventh syntax element can have a fourth value to an nth value.
[0239] Table 11 below shows another example of clearly defining the splitting number and the splitting method.
[0240] [Table 11]
[0241] Descriptor num_processing_column_unit_minus1 u(4) num_processing_row_unit_minus1 u(4)
[0242] The syntax element num_processing_column_unit_minus1 (e.g., the eighth syntax element) can be information about the number of columns of the splitting processing unit. That is, the syntax element num_processing_column_unit_minus1 represents the number of splitting regions (i.e., processing units) of the input image along the width direction, and can be a value obtained by subtracting 1 from the value (e.g., applying ceiling) obtained by dividing the width of the input image by the width of the processing unit (e.g., processing_width). The eighth syntax element can be explicitly signaled to the bitstream.
[0243] The syntax element num_processing_row_unit_minus1 (e.g., the ninth syntax element) can be information about the number of rows of the splitting processing unit. That is, the syntax element num_processing_row_unit_minus1 represents the number of splitting regions (i.e., processing units) of the input image along the height direction, and can be a value obtained by subtracting 1 from the value (e.g., applying ceiling) obtained by dividing the height of the input image by the height of the processing unit (e.g., processing_height). The ninth syntax element can be explicitly signaled to the bitstream. As an example, the eighth syntax element and the ninth syntax element can be signaled independently.
[0244] Table 12 below shows another example of clearly defining the splitting number and the splitting method.
[0245] [Table 12]
[0246]
[0247] The syntax element num_processing_unit_info_present_flag (e.g., the fifth syntax element) is information regarding whether there is information about the number of processing units and can be signaled to the bitstream, and its repeated description will be omitted.
[0248] The syntax element num_processing_unit_minus1 (e.g., the sixth syntax element) is information about the total number of processing units, and when added with 1, it can mean the split total. The sixth syntax element can be signaled based on another syntax element (e.g., the fifth syntax element) as shown in Table 12. As an example, when the value of the fifth syntax element is a specific value (e.g., 1), the sixth syntax element can be signaled. The repeated description of the sixth syntax element will be omitted.
[0249] [Table 13]
[0250]
[0251] The syntax element num_processing_unit_info_present_flag (e.g., the fifth syntax element) is information regarding whether there is information about the number of processing units and can be signaled to the bitstream, and its repeated description will be omitted.
[0252] The syntax element num_processing_unit_minus1 (e.g., the sixth syntax element) is information about the split number, and when added with 1, it can mean the split number. The repeated description of the sixth syntax element will be omitted.
[0253] The syntax element split_method (e.g., the seventh syntax element) is information about the split method and can be signaled based on another syntax element (e.g., the fifth syntax element). As an example, when the value of the fifth syntax element is a specific value (e.g., 1), the seventh syntax element can be signaled. The repeated description of the seventh syntax element will be omitted.
[0254] Table 14 below shows another example that clearly defines the split number and the split method.
[0255] [Table 14]
[0256]
[0257] The syntax element num_processing_unit_info_present_flag (e.g., the fifth syntax element) is information on whether there is information on the number of processing units, and its repeated description will be omitted.
[0258] The syntax element num_processing_column_unit_minus1 (e.g., the eighth syntax element) may be information on the number of columns for splitting the processing units. The eighth syntax element may be signaled based on another syntax element (e.g., the fifth syntax element). As an example, when the value of the fifth syntax element is a specific value (e.g., 1), the eighth syntax element may be signaled. The repeated description of the eighth syntax element will be omitted.
[0259] The syntax element num_processing_row_unit_minus1 (e.g., the ninth syntax element) may be information on the number of rows for splitting the processing units. The ninth syntax element may be signaled based on another syntax element (e.g., the fifth syntax element). As an example, when the value of the fifth syntax element is a specific value (e.g., 1), the ninth syntax element may be signaled. The repeated description of the ninth syntax element will be omitted.
[0260] Table 15 below shows another example that clearly defines the splitting quantity and the splitting method.
[0261] [Table 15]
[0262]
[0263] The syntax element num_processing_unit_info_present_flag (e.g., the fifth syntax element) is information on whether there is information on the number of processing units, and the fifth syntax element may be signaled based on another syntax element (e.g., the first syntax element processing_unit_info_present_flag). As an example, when the value of the first syntax element is a specific value (e.g., 1), the fifth syntax element may be signaled. The repeated description of the fifth syntax element will be omitted.
[0264] The syntax element num_processing_unit_minus1 (e.g., the sixth syntax element) is information about the split count, and when added with 1, it can mean the split count. The sixth syntax element can be signaled based on another syntax element (e.g., the fifth syntax element num_processing_unit_info_present_flag). As an example, when the value of the fifth syntax element is a specific value (e.g., 1), the sixth syntax element can be signaled. The repeated description of the sixth syntax element will be omitted.
[0265] The syntax element split_method (e.g., the seventh syntax element) is information about the split method, and it can be signaled based on another syntax element (e.g., the fifth syntax element). As an example, when the value of the fifth syntax element is a specific value (e.g., 1), the seventh syntax element can be signaled. The repeated description of the seventh syntax element will be omitted.
[0266] Table 16 below shows another example that clearly defines the split count and the split method.
[0267] [Table 16]
[0268]
[0269] The syntax element num_processing_unit_info_present_flag (e.g., the fifth syntax element) is information about whether there is information about the number of processing units, and the fifth syntax element can be signaled based on another syntax element (e.g., the first syntax element processing_unit_info_present_flag). As an example, when the value of the first syntax element is a specific value (e.g., 0), the fifth syntax element can be signaled. The repeated description of the fifth syntax element will be omitted.
[0270] The syntax element num_processing_column_unit_minus1 (e.g., the eighth syntax element) can be information about the number of columns of the split processing units. The eighth syntax element can be signaled based on another syntax element (e.g., the fifth syntax element). As an example, when the value of the fifth syntax element is a specific value (e.g., 1), the eighth syntax element can be signaled. The repeated description of the eighth syntax element will be omitted.
[0271] The syntax element num_processing_row_unit_minus1 (e.g., the ninth syntax element) may be information regarding the number of rows of the splitting processing units. The ninth syntax element may be signaled based on another syntax element (e.g., the fifth syntax element). As an example, when the value of the fifth syntax element is a specific value (e.g., 1), the ninth syntax element may be signaled. The repeated description of the ninth syntax element will be omitted.
[0272] Once the splitting number and the splitting method are determined, splitting information may be generated (S1340). As an example, the splitting information includes information regarding the splitting number and the splitting method, and may further include information regarding each input data split according to the splitting number and the splitting method (e.g., position information of the split data in the entire input data, additional index information from which the position can be identified, etc.). This process may be performed in the same manner during the encoding process of the encoder and the image decoding process to be described, and operation S1340 may further include defining the size of each split region, defining the position of each split region, and defining the splitting information of each region. Figure 14 As an example, the definition of the size of the split region (e.g., the width of a column, the height of a row, etc.) may be performed based on the width and height of the input data (e.g., PicWidth and PicHeight). Each of Tables 17 to 26 below shows an example of a method for defining the size of each split region.
[0273] As an example, the definition of the size of the split region (e.g., the width of a column, the height of a row, etc.) may be performed based on the width and height of the input data (e.g., PicWidth and PicHeight). Each of Tables 17 to 26 below shows an example of a method for defining the size of each split region.
[0274] [Table 17]
[0275]
[0276] [Table 18]
[0277]
[0278] [Table 19]
[0279]
[0280] [Table 20]
[0281]
[0282] [Table 21]
[0283]
[0284] [Table 22]
[0285]
[0286] [Table 23]
[0287]
[0288] [Table 24]
[0289]
[0290] [Table 25]
[0291]
[0292]
[0293] Table 17 shows examples of cases where only one of num_processing_unit_info_present_flag (e.g., the fifth syntax element) and processing_unit_info_present_flag (e.g., the first syntax element) can have a first value (e.g., 1) or both can have a second value (e.g., 0). Table 17 also shows examples of cases where the split unit and size are defined in the width and height directions and can be associated with Tables 15 and 16.
[0294] Table 18 is an example of a method for defining the size of a split region when information about the number of splits and the size of the split unit is not encoded and the relevant information is predefined. This can be associated with the case where the size of the processing unit can be specified in operation S1320 or calculated from given information. The syntax elements ProUnitWidth(general_level_idc) and ProUnitHeight(general_level_idc) can respectively mean functions that return the width and height of a processing unit predefined according to a level identifier (e.g., general_level_idc). The level identifier general_level_idc can be information for identifying the technology for configuring the corresponding encoder / decoder and the amount of required resources (e.g., PTL information).
[0295] Table 19 can be associated with Table 15. In Table 19, the size of the split region can be defined according to the presence or absence of information about the number of splits.
[0296] Table 20 can be associated with Table 12. In this case, information about the number of splits can always be defined, and the size of the split region can be defined accordingly.
[0297] Table 21 can be associated with Table 11. In this case, the size of the split region can always be defined according to the presence or absence of split size information.
[0298] Table 22 can be associated with Table 13. In this case, the split unit size information can always be defined, and the size of the split region can be defined accordingly.
[0299] Table 23 and Table 24 can be associated with the example of Table 10. More specifically, Table 23 can show the case of width splitting of a predefined input image, and Table 24 can show the case of height splitting of a predefined input image.
[0300] Table 25 can be associated with the example of Table 16. ProUnitSize(general_level_idc) can represent a function that returns the split size according to general_level_idc, and NumUnit(PicWith, PicHeight, processing_unit_size, num_processing_row_unit_minus1, num_processing_column_unit_minus1) can represent a function that receives the width and height of the input image and the split size as factors and returns the number of width splits and the number of height splits.
[0301] In addition, as an example, the definition of the position of each split region can be performed by specifying the position in the input image and is based on the boundary position of the row and the boundary position of the column. The following Table 26 and Table 27 are examples of the process of defining the position of each split region.
[0302] [Table 26]
[0303]
[0304] [Table 27]
[0305]
[0306] In Tables 26 and 27, SplitColBdVal and SplitRowBdVal can represent the boundary positions of columns and rows, respectively. Table 26 can show an example of a method for defining the boundary position of a column, and Table 27 can show an example of a method for defining the boundary position of a row. As an example, according to the method of splitting an input image, only one or both of the two splitting methods disclosed in Tables 26 and 27 can be applied. For example, when splitting the input image along the column direction, only the boundary position of the column can be defined, and when splitting the input image along the row direction, only the boundary position of the row can be defined. In addition, when splitting the input image in both directions, it may be necessary to define the boundary positions of both the column and the row. For example, when splitting the input image only along the column direction, SplitRowBdVal[0] can be 0, and when splitting the input image only along the row direction, SplitColBdVal[0] can be 0.
[0307] In addition, as an example, defining the splitting information for each region may include a process of deriving the index of each splitting region. Tables 28 and 29 below can show examples of the process of calculating the position index of each splitting region.
[0308] [Table 28]
[0309]
[0310] [Table 29]
[0311]
[0312] In Tables 28 and 29, SplitX can show the width position index of the splitting region (the index of SplitColBdVal and SplitColBdWidth), and SplitY can show the height position index of the splitting region (the index of SplitRowBdVal and SplitRowBdHeight). The position index of the splitting region can be defined as Split X and Split Y respectively, or by combining the two values into one index. Expression 2 below shows an example of an expression for expressing the two values as one index.
[0313] [Expression 2]
[0314] SplitIndex = SplitX + (num_processing_row_unit_minus1 + 1) * SplitY
[0315] As described above, an index SplitIndex can represent the width position index of a split region (i.e., a processing unit) (the index of SplitColBdVal and SplitColBdWidth), SplitY can represent the height position index of a processing unit (the index of SplitRowBdVal and SplitRowBdHeight), and the repeated description of num_processing_row_unit_minus1 will be omitted.
[0316] Finally, split information (S1340) can be generated, which will be described with reference to Figure 17 、 Figure 18 、 Figure 19 and Figure 20 . More specifically, Figures 17 to 19 is a diagram illustrating an example of generating and configuring split information, and Figure 20 is a diagram illustrating an example of configuring a bitstream accordingly. First, the split information can include information about the split region and also include information about the position of the split region. As an example, the information about the position of the split region can include a split region position index. Referring to Figures 17 to 19 , Figures 17 to 19 , the boundaries of the rows, the height of the rows, the boundaries of the columns, and the widths of the columns can be represented by SplitRowBdVal, SplitRowHeightVal, SplitColBdVal, and SplitColWidthVal respectively, and their repeated descriptions will be omitted. According to Figure 17 , when an input image is split into processing units for multiple rows and multiple columns, indexes can be assigned to each split region (processing unit) from the upper left to the lower right by assigning numbers from the left side to the right side and from the top to the bottom of the input image. The index numbers can be integer values starting from 0 and incrementing by 1. According to Figure 18 , when an input image is split into processing units for multiple rows and a single column, indexes can be assigned to each split region (processing unit) by assigning numbers from the top to the bottom of the input image. According to Figure 19 , when an input image is split into processing units for a single row and multiple columns, indexes can be assigned to each split region (processing unit) by assigning numbers from the left side to the right side of the input image. The following Tables 30 to 34 can show examples of the definition of the indexes of the split regions. First, Table 30 is as follows.
[0317] [Table 30]
[0318] Descriptor SplitX u(4) SplitY u(4)
[0319] The syntax element SplitX (e.g., the tenth syntax element) may represent the index of the split region (i.e., the processing unit) in the width direction.
[0320] The syntax element SplitY (e.g., the eleventh syntax element) represents the index of the split region in the height direction. SplitX and SplitY can be derived from Table 28, Table 29, etc. Additionally, SplitX and SplitY can be signaled independently.
[0321] [Table 31]
[0322] Descriptor SplitIndex u(8)
[0323] The syntax element SplitIndex (e.g., the twelfth syntax element) may represent the unified index defined by SplitX and SplitY. This can be derived from Expression 2 and signaled to the bitstream.
[0324] [Table 32]
[0325] Descriptor split_index_present_flag u(1)
[0326] The syntax element split_index_present_flag (e.g., the thirteenth syntax element) may indicate whether the index of the split region is defined. That is, the syntax element split_index_present_flag can be the split index presence information. When the value of the thirteenth syntax element is the first value (e.g., 1), it may mean that the split region has been defined in the bitstream, and when the value of the thirteenth syntax element is not the first value, that is, any other value, it may mean that the index of the split region is not defined in the bitstream.
[0327] [Table 33]
[0328]
[0329] SplitX (e.g., the tenth syntax element) and SplitY (e.g., the eleventh syntax element) can be signaled based on another syntax element (e.g., the thirteenth syntax element split_index_present_flag). As an example, when the value of the thirteenth syntax element is the first value (e.g., 1), the tenth syntax element and the eleventh syntax element can be signaled. The repetitive description of each syntax element will be omitted.
[0330] [Table 34]
[0331]
[0332] The SplitIndex (e.g., the twelfth syntax element) may be signaled based on another syntax element (e.g., the thirteenth syntax element split_index_present_flag). As an example, when the value of the thirteenth syntax element is the first value (e.g., 1), the twelfth syntax element may be signaled. The repetitive description of each syntax element will be omitted.
[0333] The split information generated in this way may be encoded as Figure 20 a bitstream and signaled to the decoder side. Refer to Figure 20 (a) of Figure 20 (b) of Figure 20 The bitstream may include a sequence header and a sequence unit payload, and the sequence unit payload may include a plurality of region headers and a plurality of region data, as shown in Figure 20 (a) of Figure 20 (b), the region headers may be omitted. Here, the region data may include the encoded data of each split region (i.e., processing unit), and the region headers may include information (e.g., syntax elements SplitX, SplitY, and / or SplitIndex) from which the encoded data can be identified. As an example, when the encoding / decoding order of the split regions (processing units) is predefined (e.g., raster scan order, etc.), the storage order in the bitstream may be index information, so in this case, the index information may not be encoded. In this case, the bitstream may not include region headers as in
[0334] (b) of Figure 14 When describing the above embodiments of the present disclosure in terms of an image decoder, the operation S1410 of determining the PTL may include the process of decoding or deriving and obtaining the information defined in operations S1320 and S1330 and the information defined in operation S1310 of Figure 13 . That is, since all descriptions of the information described in operations S1310, S1320, and S1330 and the corresponding operations can be applied, the repetitive description thereof will be omitted.
[0335] In addition, the operation S1420 of generating split information on the decoder side may correspond to Figure 13 operation S1340 of
[0336] [Table 35]
[0337]
[0338] When the bitstream includes position index information in the split information and the position index can be defined as SplitX and SplitY, this information can be used to derive SplitRowBdVal, SplitColBdVal, SplitColBdWidth, and SplitRowBdHeight. As another example, when the position index information is derived as a single index (e.g., SplitIndex), SplitX and SplitY can be obtained by applying Expression 2 in reverse. The DecodedSplitPic included in Table 35 can represent the decoded split region, and ConcatPic can represent each split region merged into its original form. As described in the reference Figure 20 stated, in Figure 20 (a) of, the region data can be the encoded data of each split region, and the region header can be the information from which the encoded data can be identified (e.g., SplitX, SplitY, and / or SplitIndex). Different from Figure 20 (a) of, in Figure 20 (b) of, the region header may not be defined. This may indicate that when the encoding / decoding order of the split regions is predefined (e.g., raster scan order, etc.), the storage order in the bitstream can be the index information. Therefore, the index information may not be included in the split information. That is, the index information may not be signaled to the bitstream. Since this is described with reference to other tables, its repeated description will be omitted.
[0339] According to the present disclosure, the input data can be effectively split considering resources and peak memory, and the split regions can be reconstructed into their original forms.
[0340] Figure 21 is a diagram illustrating an image decoding method according to an embodiment of the present disclosure. Figure 21 The image decoding method of can be executed by an image decoding device and can be based on the above embodiments. Therefore, the repeated description of the above embodiments will be omitted hereinafter.
[0341] As an example, the size of the processing unit of an image can be derived (S2110) based on the available resources and peak memory of a decoder for image decoding. As described above, the available resources of the decoder (i.e., the available resources and peak memory) can vary according to the performance of the decoder and can also vary according to the hardware configuration of the decoder (such as GPU, CPU, RAM, etc.). During the decoding process, the peak memory can be determined according to the size of the input data, and the size of the processing unit can vary according to the performance of the decoder. The process of deriving the size of the processing unit based on the peak memory and resources is as described above. Thereafter, an image splitting method can be derived (S2120) based on the derived size of the processing unit. The image splitting method can include the number of processing units, i.e., the number of splitting regions, etc. Thereafter, image splitting information can be derived (S2130) based on the splitting method. The image splitting information can include the position information of the processing unit. Thereafter, the image can be reconstructed (S2140) by connecting the processing units of the image based on the derived image splitting information. This process can include the process of determining the position of the processing unit in the image based on the image splitting information. As an example, the position of the processing unit can be determined based on the position index of the processing unit. Here, the position index can include a width index and a height index, and the width index and height index can be represented by the corresponding indices or represented as a single index. Additionally, the position index can be derived based on the boundary position information about rows and the boundary position information about columns. As an example, the position information of the processing unit can include the boundary position information about rows and the boundary position information about columns. Further, the size of the processing unit can be derived based on the number of width splits and the number of height splits of the image. Here, the number of width splits and the number of height splits of the image can be obtained from the bitstream. As an example, the size of the processing unit can be additionally derived based on the size of the maximum processing unit included in the PTL information.
[0342] In addition, as an example, the splitting method, the size of the processing unit, the size of the maximum processing unit, etc. can be signaled to the bitstream or derived.
[0343] In addition, since Figure 21 it corresponds to an embodiment of the present disclosure, it is obvious that the order of some operations can be changed, some operations can be omitted, or some operations can be added.
[0344] Figure 22 is a diagram illustrating an image encoding method according to an embodiment of the present disclosure.
[0345] Figure 22 The image encoding method can be executed by an image encoding device and can be based on the above embodiments. Therefore, hereinafter, the repeated description of the above embodiments will be omitted.
[0346] First, for image encoding, the size of the processing unit of the image can be determined (S2210). The size of the processing unit can be adaptively determined based on the available resources and peak memory of the device. That is to say, the available resources (that is, the available resources and peak memory) can vary according to the performance of the device and can also vary according to the hardware configuration of the device (such as GPU, CPU, RAM, etc.). Thereafter, an image splitting method can be determined (S2220) based on the determined size of the processing unit. The image splitting method can include the number of processing units, that is, the number of splitting regions, etc. Image splitting information can be generated (S2230) based on the splitting method. In this case, the image splitting information can include the position information of the processing unit. The image splitting information can include the position information of the processing unit. This process can include the process of determining the position of the processing unit in the image based on the image splitting information. In addition, the position information of the processing unit can include the position index of the processing unit, and the position index can include the above-mentioned width index and height index. The width index and height index can be composed of separate indexes or represented by a single index. Additionally, the position index can be derived based on the boundary position information about rows and the boundary position information about columns. As an example, the position information of the processing unit can include the boundary position information about rows and the boundary position information about columns. Furthermore, the size of the processing unit can be derived based on the number of width splits and the number of height splits of the image. Here, the number of width splits and the number of height splits of the image can be obtained from the bitstream. As an example, the size of the processing unit can be additionally derived based on the size of the maximum processing unit included in the PTL information. During the decoding process, the peak memory can be determined according to the size of the input data, and the size of the processing unit can vary according to the performance of the decoder. Since the process of deriving the size of the processing unit based on the peak memory and resources is as described above, its repeated description will be omitted.
[0347] All the names of the above syntax elements are arbitrarily specified for clarity and do not limit the names of the corresponding syntax elements. Additionally, the first syntax element to the thirteenth syntax element can be respectively referred to as the first information to the thirteenth information. Additionally, the first syntax element to the thirteenth syntax element can be obtained from the bitstream or can be derived from other syntax elements, and the other syntax elements can also be included in the embodiments of the present disclosure.
[0348] Additionally, the bitstream generated by the image encoding method can be stored in a non-transitory computer-readable recording medium.
[0349] Additionally, as another example, the bitstream generated by the image encoding method can be sent to another device (such as an image decoding device, etc.). In this case, the method of sending the bitstream can include the process of sending the bitstream.
[0350] Although, for clarity of description, the exemplary methods of the present disclosure above are represented as a series of operations, it is not intended to limit the order in which the steps are performed, and, when necessary, these steps may be performed simultaneously or in a different order. To implement the methods according to the present disclosure, the described steps may further include other steps, may include the remaining steps except for some steps, or may include other additional steps except for some steps.
[0351] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) may perform an operation (step) of confirming the execution conditions or circumstances of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device may perform the predetermined operation after determining whether the predetermined condition is satisfied.
[0352] The various embodiments of the present disclosure are not a list of all possible combinations, but are intended to describe the representative aspects of the present disclosure, and the content described in the various embodiments may be applied independently or in combinations of two or more.
[0353] The embodiments described in the present disclosure may be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in the respective figures may be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, the information for implementation (e.g., information about instructions) or algorithms may be stored in a digital storage medium.
[0354] In addition, a decoder (decoding device) and an encoder (encoding device) to which the embodiments of the present disclosure are applied may be included in a multimedia broadcast transmission and reception device, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camera, a video on demand (VoD) service providing device, an OTT video (over-the-top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, an augmented reality (AR) device, a videophone video device, a transportation terminal (e.g., a vehicle (including an autonomous vehicle) terminal, a robot terminal, an aircraft terminal, a ship terminal, etc.), and a medical video device, etc., and may be used to process video signals or data signals. For example, an OTT video device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0355] In addition, the processing method according to an embodiment of the present disclosure can be generated in the form of a program executed by a computer and stored in a computer-readable recording medium. Multimedia data having a data structure according to an embodiment of this document can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store computer-readable data. The computer-readable recording medium includes, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmitted via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0356] In addition, an embodiment of the present disclosure can be implemented as a computer program product by program code, and the program code can be executed by a computer according to an embodiment of the present disclosure. The program code can be stored on a computer-readable carrier.
[0357] Figure 23 is a view illustrating an example of a content streaming system to which an embodiment of the present disclosure is applicable.
[0358] Reference Figure 23 , the content streaming system according to an embodiment of the present disclosure may mainly include an encoding server, a streaming server, a network server, a web storage, a user device, and a multimedia input device.
[0359] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, video camera, etc. into digital data to generate a bitstream and sends the bitstream to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, video camera, etc. directly generates a bitstream, the encoding server can be omitted.
[0360] The bitstream can be generated by an image encoding method or an image encoding device according to an embodiment of the present disclosure, and the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.
[0361] The streaming server sends multimedia data to the user device via the network server based on the user's request, and the network server serves as a medium for notifying the user of the service. When the user requests a desired service from the network server, the network server can forward it to the streaming server, and the streaming server can send the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server serves as a command / response control between devices in the content streaming system.
[0362] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming media server can store the bitstream within a predetermined time.
[0363] Examples of user devices can include mobile phones, smartphones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc.
[0364] Each server in the content streaming system can operate as a distributed server, in which case the data received from each server can be distributed.
[0365] Figure 24 FIG. is another example of a content streaming system to which an embodiment of the present disclosure is applicable.
[0366] Reference Figure 24 , in an embodiment such as VCM, depending on the performance of the device, the user's request, the characteristics of the task to be performed, etc., a task can be executed in the user terminal or can be executed in an external device (e.g., a streaming media server, an analysis server, etc.). In this way, in order to send the information necessary to execute the task to the external device, the user terminal can generate a bitstream including the information necessary to execute the task directly or through the encoding server (e.g., information such as tasks, neural networks, and / or usage).
[0367] In an embodiment, the analysis server may perform a task requested by a user terminal (or from an encoding server) after decoding the encoded information received from the user terminal (or from the encoding server). At this time, the analysis server may send the result obtained through task execution back to the user terminal, or may send it to another linked service server (e.g., a web server). For example, the analysis server may send the result obtained by performing a task of determining a fire to a fire-related server. In this case, the analysis server may include a separate control server. In this case, the control server may be used to control commands / responses between each device associated with the analysis server and the server. Additionally, the analysis server may request desired information from the web server based on the task to be performed by the user device and the task information that can be performed. When the analysis server requests a desired service from the web server, the web server sends it to the analysis server, and the analysis server may send the data to the user terminal. In this case, the control server of the content streaming system may be used to control commands / responses between devices in the streaming system.
[0368] Industrial applicability
[0369] Embodiments according to the present disclosure may be used for encoding / decoding images.
Claims
1. An image decoding method, the image decoding method comprising the following steps: Deriving the size of the processing unit of the image based on the available resources of the decoder and the peak memory; Deriving the splitting method of the image based on the derived size of the processing unit; Deriving the splitting information of the image based on the splitting method; And Connecting and reconstructing the processing unit of the image based on the derived splitting information of the image, wherein the splitting information of the image includes the position information of the processing unit.
2. The image decoding method according to claim 1, wherein, The connection and reconstruction of the processing unit of the image includes: determining the position of the processing unit in the image based on the splitting information of the image.
3. The image decoding method according to claim 2, wherein, The position of the processing unit is determined based on the position index of the processing unit.
4. The image decoding method according to claim 3, wherein, The position index includes a width index and a height index.
5. The image decoding method according to claim 3, wherein, The position index is derived based on the boundary position information about rows and the boundary position information about columns.
6. The image decoding method according to claim 1, wherein, The position information of the processing unit includes the boundary position information about rows and the boundary position information about columns.
7. The image decoding method according to claim 1, wherein, The size of the processing unit is derived based on the width splitting number and the height splitting number of the image.
8. The image decoding method according to claim 7, wherein, The width splitting number and the height splitting number of the image are obtained from the bitstream.
9. The image decoding method according to claim 1, wherein, The size of the processing unit is additionally derived based on the size of the maximum processing unit included in the profile, layer, and level information.
10. An image encoding method, the image encoding method comprising the following steps: Determining the size of the processing unit of the image; Determining the splitting method of the image based on the determined size of the processing unit; And Generating the splitting information of the image based on the splitting method, wherein the splitting information of the image includes the position information of the processing unit.
11. The image encoding method according to claim 10, wherein, The position information of the processing unit includes the position index of the processing unit.
12. The image encoding method according to claim 11, wherein, The position index includes a width index and a height index.
13. A bitstream sending method, the bitstream sending method comprising the following steps: Sending the bitstream generated by the image encoding method, wherein the image encoding method includes the following steps: Determining the size of the processing unit of the image; Determining the splitting method of the image based on the determined size of the processing unit; and Generating the splitting information of the image based on the splitting method, and the splitting information of the image includes the position information of the processing unit.
14. A computer-readable recording medium having recorded thereon a bitstream generated by the image encoding method according to claim 10.