Method and apparatus for video encoding and decoding on basis of optimized post-processing
The integration of neural network post-processing filters in video decoding and encoding methods addresses the challenge of high compression ratios and image quality degradation, enhancing efficiency and quality in digital video encoding and decoding processes.
Patent Information
- Application Number
- PCT/KR2025/004611
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-31
- Filing Date
- 2025-04-04
- Publication Date
- 2025-10-16
AI Technical Summary
Existing digital video encoding and decoding technologies face challenges in achieving high compression ratios while minimizing image quality degradation, particularly in terms of computational resources and time consumption, which can interfere with the production, recording, and distribution of high-quality digital video.
Implementing a video decoding method that includes receiving a bitstream with supplemental enhancement information, decoding the image, and performing post-processing using neural network filters based on this information for tasks such as chroma upsampling and resolution resampling, and encoding methods that determine post-processing purposes and generate auxiliary enhancement information for improved image quality and efficiency.
Enhances encoding and decoding efficiency, improves video quality, reduces computational and hardware requirements, and optimizes post-processing for various media formats and display devices, thereby addressing the limitations of existing technologies.
Smart Images

Figure KR2025004611_16102025_PF_FP_ABST
Abstract
Description
Video encoding / decoding method and device based on optimized post-processing
[0001] The present invention relates to the field of encoding and decoding of digital video, and relates to a method for encoding and decoding digital video, a method for recording such data, and components, devices, and systems for realizing such a method.
[0002] The present invention may be in the same technical field as at least one of the digital video compression technology standards known by the names of standards such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG, or may be in the technical field for improving the inherent efficiency of the standards, or may be in the technical field for improving or replacing the standards.
[0003] Digital video encoding and decoding are widely used in various digital video applications. For example, digital television broadcasting, video transmission over communication networks, video calls / video conversations / video chats, recording and providing video content using optical media including video compact discs (VCDs), digital versatile discs (DVDs), and Blu-Rays, all processes for producing, editing, collecting, and distributing video content, and devices such as video recording devices and camcorders for various purposes, including personal, commercial, industrial, and security purposes, all depend on video encoding and decoding technologies.
[0004] Accordingly, implementations that may be referred to as digital video encoders and decoders may form part of a wide range of devices related to the generation, recording, and provision of digital video, including digital televisions, digital broadcasting systems, wireless broadcasting systems, computers in the form of notebooks / desktops / tablets, e-book readers, digital cameras, digital recording devices, digital multimedia playback devices, video game devices / terminals / consoles, mobile phones (including smartphones) with multimedia playback capabilities, equipment for video conferencing, and other devices.
[0005] The above digital video encoders and decoders can be implemented by a digital video compression standard that is widely used and understood by those skilled in the art. The digital video compression standard may include at least one of compression standards known by a standard name such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.
[0006] Video encoders and decoders can be implemented to more efficiently encode or decode digital video information while complying with the above standards, or by improving or modifying the above standards. Attempts to modify the above standards can also lead to the development of new standards. A well-known example is the so-called enhanced compression model (ECM), an attempt to improve and replace the existing H.266 / VVC standard, currently being developed by the Joint Video Experts Team (JVET), a joint international standardization group of ISO, IEC, and ITU-T.
[0007] Digital video is known to require a significant amount of information to describe its content in its uncompressed state. Therefore, recording or transmitting this information in its original form can be inefficient. Therefore, digital video is compressed using various methods prior to recording or transmission. These compression methods include lossy and lossless encoding. Lossy encoding sacrifices some image quality to achieve high compression performance, while lossless encoding sacrifices some compression performance to prevent image quality degradation. Regardless of the encoding method, there is a need to implement a technology that achieves a high compression ratio while minimizing image quality degradation to meet the growing demand for high-quality digital video within the constraints of limited memory storage capacity and communication transmission bandwidth.
[0008] As described above, the encoding process for compression requires various operations, such as spatial segmentation of digital video, segmentation and / or processing in color channels, removal of spatial redundancy, removal of temporal redundancy, tracking of motion vectors within the video, encoding of differential images, quantization, coefficient scan, run-length coding, entropy coding, and loop filtering. These encoding operations generally consume computing resources and take a certain amount of time to complete. Similarly, the decoding operation for the encoding operation also requires certain computing resources and a certain amount of time. The main goal of video encoding and decoding technology is to ensure that the above resource consumption and time consumption do not interfere with the production, recording, distribution, and viewing of digital videos.
[0009] Accordingly, the present invention provides a new technology that can contribute to at least one of the following: improvement of encoding efficiency, improvement of decoding efficiency, improvement of video quality, reduction of computational amount, reduction of software size, reduction of hardware size, and improvement of other performances related to encoding and decoding in the technical tasks in the field of video encoding and decoding.
[0010] According to an embodiment of the present invention for solving the above-described technical problem, a video decoding method includes the steps of receiving a bitstream including an encoded image and supplemental enhancement information about the encoded image, extracting the encoded image and the supplemental enhancement information from the bitstream, decoding the encoded image to obtain a decoded image, and performing post-processing on the decoded image based on the supplemental enhancement information, wherein the supplemental enhancement information may include information about at least one of a post-processing method and a post-processing purpose for the encoded image.
[0011] The encoded image and the auxiliary enhancement information may be characterized in that they are combined by a high-level syntax within the bit string.
[0012] The above auxiliary improvement information may include information related to at least one of general visual quality improvement, chroma upsampling, resolution resampling, picture rate upsampling, bit depth upsampling, and colorization.
[0013] The above post-processing is performed by a neural network post-processing filter, and the auxiliary improvement information may include information describing at least one neural network included in the neural network post-processing filter.
[0014] The above post-processing may be sequentially executed by a plurality of neural network-based post-processing filters, and the auxiliary improvement information may be characterized by including descriptive information about at least one of the plurality of neural networks.
[0015] The encoded image may be an image processed from a first media format into an image by a preprocessing technique, the postprocessing includes a process of restoring the decoded image to the first media format, and the auxiliary improvement information may be characterized in that it includes information related to at least one postprocessing method related to restoring the decoded image to the first media format.
[0016] The method may be such that the first media format is a multi-view image, and the auxiliary enhancement information may include at least one of information for generating a three-dimensional view from the encoded image, information for merging images for multiple viewpoints, and information for imparting a three-dimensional effect.
[0017] The above method may be characterized in that the first medium format is a holographic image, and the encoded image includes at least one of real information, imaginary information, amplitude information, and phase information expressing the holographic image.
[0018] The method may be characterized in that the first media format is neural network information, and the encoded image includes information visualizing at least one of the parameters of the neural network constituting the neural network information and the feature values of the neural network.
[0019] The above auxiliary improvement information may be characterized in that it includes at least one bit string information generated by the encoder of the first media format in the same format.
[0020] The above first media format may be a media format included in at least one standard technical specification among MIV (MPEG Immersive Video), INVR (Implicit Neural Visual Representation), NNC (Neural Network Compression), VCM (Video Coding for Machines), and PCC (Point Cloud Coding).
[0021] The above post-processing purpose may include one of a first purpose for viewing the decoded image or a second purpose for a vision task based on the decoded image, and the auxiliary improvement information may be characterized in that it includes instruction information indicating one of the first purpose and the second purpose.
[0022] The method may further include, when the instruction information indicates the second purpose, information indicating at least one of object recognition, object detection, object segmentation, information retrieval, and captioning by the computer vision.
[0023] The above auxiliary improvement information may be characterized by including information indicating different post-processing methods for at least two segmentation units included in the encoded image.
[0024] A video encoding method according to an embodiment of the present invention for solving the above-described technical problem may include a step of determining at least one of a post-processing method and a post-processing purpose for an input image, a step of generating auxiliary enhancement information for the input image based on at least one of the method and the purpose, a step of encoding the input image to obtain an encoded image, and a step of generating a bit string including the encoded image and the auxiliary enhancement information.
[0025] The above method may further include a step of performing preprocessing on the input image based on the above method and purpose.
[0026] According to an embodiment of the present invention for solving the above-described technical problem, a video decoding device includes a receiving unit that receives a bitstream including an encoded image and supplementary enhancement information about the encoded image, a high-level syntax parsing unit that extracts the encoded image and the supplementary enhancement information from the bitstream, a decoding unit that decodes the encoded image to obtain a decoded image, and a postprocessing unit that performs postprocessing on the decoded image based on the supplementary enhancement information, wherein the supplementary enhancement information may include information about at least one of a postprocessing method and a postprocessing purpose for the encoded image.
[0027] The above post-processing unit is configured based on a neural network post-processing filter, and the auxiliary improvement information may include information describing at least one neural network included in the neural network post-processing filter.
[0028] The encoded image may be an image processed from a first media format into an image using a preprocessing technique, the post-processing unit may be configured to restore the decoded image to the first media format, and the auxiliary improvement information may be characterized by including information related to at least one post-processing method related to restoring the decoded image to the first media format.
[0029] The above post-processing purpose may include one of a first purpose for viewing the decoded image or a second purpose for a vision task based on the decoded image, and the auxiliary improvement information may be characterized in that it includes instruction information indicating one of the first purpose and the second purpose.
[0030] According to an embodiment of the present invention for solving the above-described technical problem, a method for decoding an image includes the steps of receiving a bitstream including an encoded image and supplemental enhancement information regarding the encoded image, extracting the encoded image and the supplemental enhancement information from the bitstream, and decoding the encoded image to obtain a decoded image, wherein the supplemental enhancement information may include post-processing information according to physical characteristics of a display device.
[0031] The physical characteristics of the display device may be characterized by including display model information including transmissive pixel technology and emissive pixel technology.
[0032] The above method may be characterized in that the auxiliary improvement information is included and expressed in the form of a bit mask, and a plurality of post-processing information or post-processing types are configured to be applied simultaneously.
[0033] The above auxiliary improvement information may be characterized by including information indicating at least one type of post-processing optimization among object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, privacy protection optimization, and luma range adaptation optimization.
[0034] The above method may further include a step of performing post-processing based on the auxiliary improvement information.
[0035] The step of performing post-processing based on the above auxiliary improvement information may include a step of selecting and applying different post-processing methods according to the physical characteristics of the display device included in the above auxiliary improvement information.
[0036] The method may be characterized in that at least one post-processing method is grouped into a post-processing group, and the auxiliary improvement information includes information about the post-processing group.
[0037] The above method may be characterized in that the post-processing group is hierarchically configured, and the auxiliary improvement information includes information defining a relationship between an upper post-processing group and a lower post-processing group according to the hierarchical configuration.
[0038] The encoded image may be an image processed from a first media format into an image by a preprocessing technique, the postprocessing includes a process of restoring the decoded image to the first media format, and the auxiliary improvement information may be characterized in that it includes information related to at least one postprocessing method related to restoring the decoded image to the first media format.
[0039] The above post-processing may be performed by a neural network post-processing filter, and the auxiliary improvement information may be characterized by including information describing at least one neural network included in the neural network post-processing filter.
[0040] The above auxiliary improvement information may be characterized by including information indicating different post-processing methods for at least two geometric segmentation areas defined within a single encoding unit.
[0041] According to an embodiment of the present invention for solving the above-described technical problem, a method for encoding an image includes a step of generating post-processing information for an input image, a step of generating auxiliary enhancement information for the input image based on the post-processing information, a step of encoding the input image to obtain an encoded image, and a step of generating a bit string including the encoded image and the auxiliary enhancement information, wherein the auxiliary enhancement information may include post-processing information according to a physical characteristic of a display device.
[0042] The physical characteristics of the display device may be characterized by including display model information including transmissive pixel technology and emissive pixel technology.
[0043] The above method may be characterized in that the auxiliary improvement information is included and expressed in the form of a bit mask, and a plurality of post-processing information or post-processing types are configured to be applied simultaneously.
[0044] The above auxiliary improvement information may be characterized by including information indicating at least one type of post-processing optimization among object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, privacy protection optimization, and luma range adaptation optimization.
[0045] The method may be characterized in that at least one post-processing method is grouped into a post-processing group, and the auxiliary improvement information includes information about the post-processing group.
[0046] The above method may further include a step of performing preprocessing on the input image.
[0047] The step of performing preprocessing on the input image may include a step of preprocessing the input image in the first media format and processing it into an image to be encoded, and the auxiliary improvement information may include information related to at least one postprocessing method related to restoring the decoded image to the first media format.
[0048] The above auxiliary improvement information may be characterized by including information about a neural network post-filter and including information describing at least one neural network included in the neural network post-filter.
[0049] The above auxiliary improvement information may be characterized by including information indicating different post-processing methods for at least two geometric segmentation areas defined within a single encoding unit.
[0050] According to an embodiment of the present invention for solving the above-described technical problem, a video decoding device includes a receiving unit that receives a bitstream including an encoded video and supplemental enhancement information about the encoded video, a parsing unit that extracts the encoded video and the supplemental enhancement information from the bitstream, a decoding unit that decodes the encoded video to obtain a decoded video, and a post-processing unit that reads at least one piece of post-processing information from the supplemental enhancement information, wherein the supplemental enhancement information includes post-processing information according to a physical characteristic of a display device.
[0051] According to the present invention, in the technical task of the video encoding and decoding field, it is possible to obtain advantageous effects in at least one of improvement in encoding efficiency, improvement in decoding efficiency, improvement in video quality, reduction in computational amount, reduction in software size, reduction in hardware size, and improvement in other performances related to encoding and decoding.
[0052] Figure 1 is a conceptual diagram of a video communication system according to one embodiment of the present invention;
[0053] FIG. 2 is a conceptual diagram of the arrangement of an encoder and decoder in a real-time video streaming environment according to one embodiment of the present invention.
[0054] Figure 3 is a functional unit conceptual diagram of a video decoder according to one embodiment of the present invention;
[0055] Figure 4 is a functional unit conceptual diagram of a video encoder according to one embodiment of the present invention;
[0056] Figure 5 is a conceptual diagram of a frame type according to one embodiment of the present invention;
[0057] Figure 6 is a conceptual diagram showing the structure of a video encoder using VVC.
[0058] Figure 7 is a conceptual diagram of preprocessing and postprocessing of a video encoding / decoding process according to one embodiment of the present invention.
[0059] FIG. 8 is an exemplary diagram showing a video-mediated application information processing method according to one embodiment of the present invention;
[0060] Figure 9 is an exemplary diagram showing an example of using the MIV standard according to one embodiment of the present invention;
[0061] Figure 10 is an exemplary diagram showing the use of neural network information according to one embodiment of the present invention;
[0062] FIG. 11 is an exemplary diagram of a method for optimizing a post-processing filter and a post-processing method according to one embodiment of the present invention.
[0063] FIG. 12 is an exemplary diagram of another method for optimizing a post-processing filter and a post-processing method according to one embodiment of the present invention;
[0064] FIG. 13 is an exemplary diagram showing an example of application of input information for multiple post-processors according to one embodiment of the present invention; and
[0065] Figure 14 is an exemplary diagram showing an example of application of different post-processing methods according to one embodiment of the present invention.
[0066] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention.
[0067] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, the first component could be referred to as the "second component," and similarly, the second component could also be referred to as the "first component." The term "and / or" includes any combination of multiple related items described herein or any of multiple related items described herein, and is non-exclusive unless otherwise indicated. The listing of items in this application is merely an exemplary description to facilitate the spirit and possible implementation methods of the present invention, and therefore is not intended to limit the scope of embodiments of the present invention.
[0068] As used herein, "A or B" can mean "only A," "only B," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B and C."
[0069] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0070] In this specification, "at least one of A and B" may mean "only A", "only B" or "both A and B". Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted identically to "at least one of A and B".
[0071] Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”
[0072] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0073] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0074] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning within the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0075] In describing the invention in this application, embodiments may be described or illustrated in terms of unit blocks that perform the described function or functions. The blocks may be expressed in this application as one or more devices, units, modules, parts, etc. The blocks may be implemented in hardware by one or more logic gates, integrated circuits, processors, controllers, memories, electronic components, or information processing hardware implementation methods, but not limited thereto. Alternatively, the blocks may be implemented in software by application software, operating system software, firmware, or information processing software implementation methods, but not limited thereto. A single block may be implemented by being separated into multiple blocks that perform the same function, or conversely, a single block may be implemented to perform the functions of multiple blocks simultaneously. The blocks may also be implemented by being physically separated or combined according to any criteria. The blocks may also be implemented to operate in an environment where their physical locations are not specified and are separated from each other by a communication network, the Internet, a cloud service, or a communication method, but not limited thereto. All of the above implementation methods are within the realm of various embodiments that can be taken by a person skilled in the field of information and communication technology to implement the same technical idea, and therefore, any detailed implementation method should be interpreted as being included within the technical idea scope of the invention of the present application.
[0076] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the attached drawings. To facilitate a comprehensive understanding of the present invention, the same reference numerals will be used for identical components in the drawings, and redundant descriptions of identical components will be omitted. Furthermore, the multiple embodiments are not mutually exclusive, and it is assumed that some embodiments may be combined with one or more other embodiments to form new embodiments.
[0077]
[0078] digital video codec
[0079] Figure 1 is a conceptual diagram of a video communication system according to one embodiment of the present invention. The video communication system (100) may be configured to include at least two terminals (110, 120) connected to each other via a network (105).
[0080] In one embodiment of the present invention, the above-described FIG. 1 may refer to a block diagram for configuring a unidirectional video communication network. A first terminal (110) among the terminals may encode video data in order to transmit (111) the video data via a network (105). A second terminal (120) among the terminals may be configured to receive (121) the encoded video data via a network and decode and display the same.
[0081] In another embodiment of the present invention, the above-described FIG. 1 may refer to a block diagram for configuring a two-way video communication network. For the two-way video communication, each terminal (110, 120) may be configured to encode video data acquired by itself for video transmission (112, 122) to each other terminal via the network. Each terminal may also be configured to receive (113, 123) video data transmitted by another terminal via the network, decode the same, and display the decoded video data.
[0082] The terminals (110, 120) shown in Fig. 1 may be exemplified as devices such as server computers, personal computers, portable computers, and smartphones, depending on the embodiment, but are not limited thereto. The present invention is applicable to all environments for establishing a one-way or two-way video communication network, and it should be understood that the network (105) may be established by any means for transporting encoded video data between the terminals (110, 120).
[0083] In one embodiment of the present invention, the network (105) may refer to a wired or wireless communication network. Depending on the embodiment, the network may be configured to communicate information using any communication standard, which may include packet-based communication. The packet communication may be understood to include packets, for example, known as TCP or UDP.
[0084] However, in another embodiment of the present invention, the network (105) may be understood to include a process of information transmission using a recording medium. In this case, the configuration of the network is not limited to a communication medium, and should be understood to include a process of temporarily storing and physically transporting information on a hard disk, a solid state disk (SSD), a flash memory, a compact disc (CD), a digital versatile disc (DVD), a Blu-ray disc, and other mechanical, electronic, or optical recording media.
[0085] Any other means of information communication or transport, regardless of the method employed, can be considered within the scope of the present invention as long as it has a structure that supports the transmission and decoding of video data in an encoded state. Therefore, in addition to the examples listed above, any means of information communication or transport, whether known in the past or newly available, can fall within the scope of application of the present invention.
[0086] FIG. 2 is a conceptual diagram illustrating the arrangement of an encoder and decoder in a real-time video streaming environment according to one embodiment of the present invention. The streaming system (200) illustrated in FIG. 2 can be applied to video data communication networks, including, for example, digital broadcasting, video telephony, and video conferencing. However, it should be noted that technical structures identical or similar to the streaming system can be equally applied even when information is transmitted via a recording medium, as described above.
[0087] According to one embodiment of the present invention, the streaming system may include a video source (210) that generates a video stream. The video source may include a digital video acquisition means (212), which may be, for example, a digital camera or other device, for acquiring uncompressed raw video. The raw video stream (215) may have a large capacity and may therefore be compressed by a video encoder (217) coupled or connected to the video source.
[0088] The above encoder (217) may be configured as a means including hardware, software, or a combination of the two, configured to implement an image encoding method and / or an implementation method thereof according to one embodiment of the present invention.
[0089] Through the encoder (217), an encoded bitstream (219) having a reduced capacity compared to the original video stream can be output. The bitstream (219) can be provided in real time via a relay device, which may be referred to as a streaming server (220), and / or can be stored in a recording medium (225) of the streaming server (220) for subsequent use.
[0090] The streaming system (200) may include at least one streaming client (230, 240) that connects to the streaming server (220) to receive the encoded bit string (229) in real time or obtain it later. The streaming client may include a video decoder (232) that obtains the encoded bit string (229) (which may also be regarded as a copy of the bit string (219) received by the streaming server), decodes the bit string (229), and outputs the resulting video data as video data in a form that can be displayed by a display (235) or other visual, auditory, or other sensory display means.
[0091] As described above, the functions for encoding and decoding video data are collectively called a coder-and-decoder, or video codec.
[0092] FIG. 3 is a conceptual diagram of a functional unit of a video decoder according to an embodiment of the present invention. As shown in FIG. 3, a receiving unit (310) can receive at least one encoded video data to be decoded by a decoder (305). In an embodiment of the present invention, the encoded video data may be independent for each reception, and the decoding procedure of each independent video data may be independent from the decoding procedure of other video data. The encoded video data may be received by the receiving unit (310) through a hardware or software connection (315) to a device storing the same, and as described above, the storing device may be a type of streaming server located at the other end of a communication network, or may mean a physical recording medium, but is not limited thereto.
[0093] The above-described receiving unit (310) can receive the encoded video data together with other data accompanying it, such as encoded audio data or other auxiliary data, and each of the data can be separated from the video data and provided to an appropriate processing function unit (312) other than the video decoder.
[0094] When the video data is provided through a communication network, a buffer memory (320) may be coupled between the receiving unit (310) and the decoder (305) to minimize delay and disconnection according to the network environment. The buffer memory (320) may refer to a computer-readable recording medium that temporarily stores the received video data and stably supplies it to a parser (330) corresponding to the input terminal of the decoder (305). However, if the bandwidth of the communication network is sufficient, if video data is read from a recording medium in a local location that is not physically separated, or if the possibility of communication delay is not predicted in other environments, the buffer memory may be unnecessary.
[0095] The video decoder (305) may include the parser (330) as its input terminal to interpret the encoded video data. The parser may perform a function of separating (parsing) a plurality of pieces of information stored in the form of a bit string in the encoded video data according to a predetermined rule, and, if necessary, performing an entropy decoding (335) of entropy-coded video data, thereby performing a function of reconstructing symbols (338), which are paragraphs of video encoding information. The symbols (338) may include all information for controlling the operation of the decoder (305), and / or may further include information for controlling a device that is attached to and operable with the decoder (305), such as a display device. Control information for controlling the above display device may include information in a format called supplementary enhancement information (SEI) or video usability information (VUI).
[0096] As described above, the parser (330) may be configured to perform entropy decoding (335) of the encoded video data. The entropy encoding method of the encoded video data may vary depending on the encoding standard, and decoding may be performed accordingly. Representative examples of the entropy encoding standard may include variable length coding, Huffman coding, and arithmetic coding, and each of the encoding methods may be a context-adaptive or context-sensitive method depending on the standard, and may also be based on principles widely known to those skilled in the art.
[0097] The parser (330) may be configured to extract at least one picture from the encoded video data. The definition of the picture may vary depending on the encoding standard, and depending on the standard, one or more of the examples listed below may correspond simultaneously and overlappingly. The picture may be grouped, defined, and / or divided into encoding / decoding units such as, for example, groups of pictures (GOPs), pictures / frames, tiles, slices, macroblocks, blocks, subblocks, transform units (TUs), and prediction units (PUs).
[0098] The parser (330) may be configured to extract encoding information, such as transform coefficients, quantization parameters (QPs), and / or motion vectors, from the encoded video data. The parser (330) may be configured to perform entropy decoding (335) and parsing operations on the video data received from the buffer memory, and to selectively decode symbols (338) representing the encoding information. In addition, the parser (330) may be configured to selectively supply a specific symbol (338) to a specific decoding function unit within the decoder (305), such as an inverse quantization and inverse transform unit (340), an intra prediction unit (350), an inter prediction unit (355), or a loop filter unit (360). Control of such information supply can be determined by the information sequence contained in the encoded video, and may vary depending on the encoding standard, and is not limited within the scope of the embodiments of the present invention, and is not described in detail in this conceptual diagram.
[0099] The decoder (305) may be comprised of a number of conceptual functional units that receive and process the encoded information provided by the parser (330). It should be readily apparent that these conceptual functional units may be combined or further subdivided, depending on implementation needs. For example, they may be further separated for ease of implementation, or integrated into one for operational efficiency. In any case, each functional unit may be configured to perform close interaction with each other. However, despite the possibility of such integration or separation, the following description will be given as a combination of conceptual functional units to illustrate the video data decoding procedure applied as an embodiment of the present invention.
[0100] The decoder may include an inverse quantization and inverse transformation unit (340). The inverse quantization and inverse transformation unit (340) may be configured to receive encoding information including a method to be used for numerical transformation (transform), a block size, quantization coefficients for recovering quantized information, and distinction information of a quantization matrix that simplifies and represents the quantized coefficients from the parser (330), and may be configured to output block values (341) that can be input to an aggregator (370) as a result of processing the encoding information.
[0101] In one embodiment of the present invention, the output values of the inverse quantization and inverse transformation unit (340) may include intra-prediction encoded block values. The intra-predicted block values may refer to values that can be decoded using prediction information within a picture currently being decoded, for example, the current frame, but not using prediction information from a previously decoded picture, for example, the previous frame.
[0102] The prediction information within the current picture may be provided by the intra prediction unit (350). According to an embodiment of the present invention, the intra prediction unit (350) generates a block value of the same form as the block being decoded as the prediction information by using picture information of a spatially adjacent area derived from a picture currently being decoded and of which decoding has been partially completed. The picture information may be provided (381) from a buffer for the current picture, a so-called line buffer (380). The merging unit (370), according to an embodiment, may be configured to merge the prediction information (351) generated by the intra prediction unit (350) with the block values (341) provided by the inverse quantization and inverse transformation unit (340).
[0103] In another embodiment, the output values of the inverse quantization and inverse transformation unit (340) may include block values subjected to motion compensation as inter-prediction encoded block values, and in some cases, block values subjected to motion compensation. In this case, the inter-prediction unit (355) may extract and use sample information (386) used for motion-based prediction from a reference picture buffer (385). The information (356) derived by performing motion compensation on the sample information based on the symbols (338) included in the block values as the output values may be configured to be merged with the block values (341) provided by the inverse quantization and inverse transformation unit (340) by the merger unit (370). In this case, the block values (341) may be referred to as so-called differential or residual values.
[0104] The position information within the memory used by the inter prediction unit (355) to extract the sample information from the reference picture may be determined by a motion vector provided to the inter prediction unit (355) and composed of a combination of symbols (338) for indicating, for example, X, Y, and other specific points of the reference picture. The inter prediction unit (355) may also include a function for interpolating and using the sample values when a so-called 'subsampling'-capable motion vector is provided, and may further include a function for predicting and reinforcing the value of the motion vector.
[0105] The output values (371) of the above merging unit (370) may be provided to the loop filter unit (360) and processed by various loop filtering methods. The loop filter unit (360) may be configured to receive not only the block unit output (371) of the merging unit (370) but also the symbol (338) provided from the parser (330) and control its operation. The output of the loop filter unit (360) may be output to an external display means such as the display device through an output connection (390), but may be stored (361) in a line buffer (380) for use in prediction for interpreting a subsequent intra or inter coded block value, and may also be stored in a reference picture buffer (385) through this.
[0106] Certain pictures, such as frames, after their decoding is completed, can be utilized as reference pictures for performing predictive decoding in a subsequent decoding process. A picture (or frame) can be gradually accumulated in a line buffer (380) and decoded. When the decoding of a frame is completed, the contents of the line buffer (380) are transferred (383) to the reference picture buffer (385), and a new line buffer (380) can be allocated for decoding the new frame.
[0107] The above video decoder (305) may be configured to perform a decoding operation according to a predetermined video compression technique that may be documented by various international standards or commercial standards. The standards may include, for example, international standard recommendations such as H.264, H.265, and H.266 defined by the International Telecommunication Union Standardization Sub-Division (ITU-T). Those skilled in the art will understand that each of the above recommendations is equivalent to an international standard jointly defined by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC). The encoded video data may comply with a specific bitstream syntax defined by the relevant standards, as defined and required by the video compression standard documents and standard documents, and specifically by the profiles and levels specified within such documents. In addition, the complexity of the encoded video data may be limited to a certain level to comply with the profiles and levels. For example, a profile or level may be configured to limit a maximum picture size, a maximum decoding speed, and a maximum reference picture size. These limitations may, in some embodiments, also be further limited via metadata signals for a hypothetical reference decoder (HRD) and HRD buffer management included in the encoded video data.
[0108] According to one embodiment of the present invention, the receiver (310) may receive additional redundant data along with the encoded video. The additional data may be considered as part of the encoded video data. The additional data may include information that may be used by the decoder (305) to properly decode the data or to more accurately reconstruct an image that approximates the original image. The additional data may be provided in the form of, for example, layers for temporal, spatial, or signal-to-noise ratio (SNR) enhancement, redundant slices, redundant pictures, and forward error correction codes.
[0109] Figure 4 is a functional unit conceptual diagram of a video encoder according to one embodiment of the present invention. The encoder (405) may be configured to receive original video information (402) from a video source (401) and perform encoding.
[0110] The original video information (402) may have any suitable bit depth, for example, 8 bits, 10 bits, 12 bits, etc. In addition, the original video information (402) may have any suitable color space, for example, R / G / B, Y / U / V, Y / Cb / Cr, etc. In addition, the original video information (402) may have any suitable sampling structure corresponding to the color space, for example, Y / Cb / Cr 4:2:0, Y / Cb / Cr 4:4:4, etc. The original video information (402) having such a predetermined format may be provided to the encoder in the form of a digital video stream.
[0111] In a one-way video communication network, the original video information (402) can be obtained from a recording medium storing a previously prepared video source. In a two-way video communication network, the original video information (402) can be obtained from an image acquisition device, such as a camera, that generates at least one video transmission stream included in the two-way video communication.
[0112] The video data including the above original video information (402) may be configured as a plurality of pictures configured to simulate motion by playing them in chronological order. In addition to pictures, the pictures may also be expressed in concepts such as frames. The pictures may include one or more samples depending on the type of sampling structure, color space, etc. being used. Those skilled in the art will understand that the samples are closely related terms to pixels in digital images. The operation of the encoder will be described below with reference to such samples.
[0113] According to one embodiment of the present invention, the encoder (405) may be configured to encode and compress pictures (and / or grouped or segmented information) constituting the original video information (402) in real time (or according to other temporal requirements required according to the implementation method) into the form of encoded video information.
[0114] In the encoder (405), the control unit (450) may be a functional unit configured to control an appropriate encoding speed. The control unit (450) may be configured to control other functional units and be functionally coupled to the following functional units as described below. The parameters set by the control unit (450) may include parameters related to bitrate control, such as picture skip, quantizer, and variable values for applying image quality optimization techniques, and may also include values such as picture size, group of pictures (GOP) structure, and maximum search range of motion vectors. A person skilled in the art will be able to understand various other functions that the control unit (450) may have, and such other functions may be added or removed according to the design of a video encoder optimized for individual system design.
[0115] According to an embodiment of the present invention, the encoder (405) may be configured to operate in a structure such as a "coding loop" well known to those skilled in the art. To simplify the description by way of example, the coding loop may be configured with an internal encoder (so-called "source coder") (410) responsible for receiving a picture to be encoded and generating symbols based on at least one reference picture that has been encoded in the past, and a local decoder (420) configured to be connected to the internal encoder. The local decoder (420) may be configured to perform an operation to reproduce sample data to be generated by a decoder (490) located at an actual remote location that will receive encoded video information from the encoder (405) by receiving an output of the internal encoder (410).
[0116] Video data composed of sample data reconstructed by the internal decoder (420) may be configured to be input to the reference picture buffer of the encoder (405). As described above, the internal decoder (420) is implemented to reproduce the result output by the encoder (405) and to be decoded by a remote decoder, so the video data recorded in the reference picture buffer may also be identical in bit unit to the information of the reference picture buffer of the remote decoder. That is, the prediction function unit that may be included in the encoder (405) may read the same values as the sample values of the previous frame that the decoder will later refer to in the decoding process from the reference picture buffer of the encoder (405).
[0117] As described above, the principle of achieving matching of the reference picture buffers between the encoder (405) and the decoder (490) by means of the internal decoder (420) on the encoder (405) side is well known to those skilled in the art, and a method of responding to an environment in which such an environment is not guaranteed (e.g., information loss due to communication failure, etc.) can also follow what is known to those skilled in the art.
[0118] An embodiment of the operation method of the internal decoder (420) has been described in detail above with reference to FIG. 3. The decoder of FIG. 3 may be regarded as the aforementioned "remote" decoder (490). The internal decoder (420) may be implemented excluding lossless encoding and decoding sections such as the parser (330) or entropy decoding (335). This is because the internal encoder (405) is implemented to simply reproduce the operation of a decoder located at a remote location, and thus may directly decode symbols without requiring a process of compressing and then decompressing symbols. Accordingly, the functional units preceding the parser and entropy decoder as shown in FIG. 3 may not be provided or may be implemented at least partially.
[0119] As described above, according to a preferred embodiment of the present invention, any decoder function (excluding a parser and an entropy decoder) present in the decoder can naturally exist as a substantially identical function in the corresponding encoder (405).
[0120] The operation of the encoding function unit that may be included in the above encoder (405) can be considered as the reverse operation of the decoder function unit. Therefore, the embodiment can be explained by performing the operation of the decoder function unit in reverse. For example, a quantization and transform function unit corresponding to the inverse quantization and inverse transform unit may be provided, and an inter prediction encoding unit corresponding to the inter prediction unit may be provided. In addition, some additional explanations will be added.
[0121] The internal encoder (410) may be configured to perform encoding on input picture information, for example, an input frame, by a prediction encoding method executed by a prediction encoding unit (440) that operates by referencing at least one reference picture information, for example, at least one temporally previous encoded picture (or frame) from a reference picture buffer (430) from video data designated as a reference frame. In this case, the encoder (405) may be configured to encode a differential between blocks of samples constituting the input picture and blocks of samples constituting the reference picture.
[0122] The internal decoder (420) can decode video data that can be designated as the reference picture from symbols generated by the internal encoder (410). As described above, since the video data is identical to the decoding operation performed by the remote decoder, the video data used as the reference picture may be provided to the encoder (405) in a form that has undergone lossy compression and has been partially damaged, and this operation may be intended to ensure operational consistency with the decoder.
[0123] The prediction encoding unit (440) may be configured to perform a prediction search operation within the encoder (405). The prediction search operation may refer to an operation corresponding to the inter prediction or intra prediction described in the description of the decoder. For picture information that is input and scheduled to be newly encoded, the prediction unit may access the reference picture buffer (430) to retrieve information such as a motion vector, a block shape, and metadata that may include the same, which are information indicating a point of a reference picture that can function as prediction reference information suitable for the new picture information, and a sample block to be actually referenced. The above prediction encoding unit (440) may operate on the basis of the so-called "sample block by pixel block" criteria in order to obtain appropriate prediction reference information. According to one embodiment of the present invention, at least one prediction reference information designating at least one reference picture information stored in the reference picture buffer (430) may be designated for the input picture, as determined based on the search results obtained by the prediction encoding unit (440).
[0124] In one embodiment of the present invention, the control unit (450) may be configured to manage the overall encoding operation of the internal encoder (410), including setting parameters used to encode video data.
[0125] All outputs of the above-described functional units may be subjected to entropy encoding (460) in order to be finally output. The entropy encoding (460) may include various entropy coding techniques, such as variable length coding, Huffman coding, and arithmetic coding, for the symbols generated by the various functional units as described above, and each encoding method may be a context-adaptive or context-sensitive method according to the standard, or may be based on principles widely known to those skilled in the art. Such entropy encoding (460) can typically achieve lossless compression, and thus can be configured to convert at least one symbol generated by the functional units into encoded video data.
[0126] The above control unit (450) may, when controlling the operation of the encoder (405), apply to each picture (or frame) the type of encoding in which a specific picture is encoded during the encoding period. Depending on the type, the method by which the picture is encoded may be affected. Depending on the embodiment, the type may include what is categorized as the following "frame type."
[0127] Fig. 5 is a conceptual diagram of a frame type according to one embodiment of the present invention. The following description will be made with reference to Fig. 5.
[0128] An intra (“I”) picture (510) may refer to a picture that can be encoded and decoded using only its own information without referring to other picture information in video data through predictive encoding. The “I” picture may be designated by names such as a key frame, an independent / instantaneous decoder refresh (IDR) frame, and a clean random-access (CRA) frame, depending on the video encoding standard, and the “I” pictures designated by the various names as described above may have various modifications and application methods as permitted by each standard and may be partially different from each other. In addition to those listed above, various application methods for implementing the “I” picture may be by various methods that are already known to those skilled in the art or may be newly provided.
[0129] A prediction ("P") picture (520) may refer to a picture that can be encoded and decoded through intra or inter prediction based on at least one prediction information and / or a motion vector that designates at least one reference picture to predict sample values of a block constituting the picture. The "P" picture may be configured to refer to only one reference frame, or may be configured to refer to one or more reference frames, according to a video encoding standard. When referring to more than one reference frame, sample information and / or associated metadata derived from multiple reference pictures may be used to reconstruct a single block. However, in common cases, a picture designated as a "P" picture may be understood as a picture that performs reference only to a temporally preceding picture.
[0130] A bidirectional prediction ("B") picture (530) may refer to a picture that can be encoded and decoded through intra or inter prediction based on at least one piece of prediction information and / or a motion vector that designates at least two reference pictures in order to predict sample values of blocks constituting the picture. In a common case, a picture designated as the "B" picture is distinguished from a picture designated as the "P" picture, and may be understood as a picture that performs a reference without being limited to a temporally preceding picture.
[0131] Video data may be spatially divided into a plurality of sample blocks during the encoding and decoding process, and encoding may be performed in units of the blocks. The block units may include, but are not limited to, sizes such as 4x4, 8x8, 4x8, or 16x16 in units of horizontal / vertical pixels, as is widely known. The block may be encoded using a predictive encoding method with reference to any other (already encoded) blocks, as permitted and / or restricted by the type specified for each picture in which the block is included. For example, the blocks of the "I" picture (510) may be encoded without using a predictive encoding method, or with reference to blocks that have already been encoded within the same partial picture. That is, only the so-called intra prediction method may be used. In contrast, the "P" picture (520) may further reference a reference picture encoded in at least one previous time unit, and thus, inter prediction may also be used for encoding along with intra prediction. In the case of a "B" picture (530), reference can be made not only to a previously encoded picture in the encoding order but also to a later reference picture in terms of time unit. However, it is widely known that there may be blocks within a "P" picture or a "B" picture that are encoded without relying on predictive encoding.
[0132] The above video encoder (405) may be configured to perform encoding operations according to a predetermined video compression technique that may be documented by various international standards or commercial standards. Examples of the above standards may include all of those described in the above decoder.
[0133] According to one embodiment of the present invention, the transmitter (470) may buffer the encoded video data generated by the entropy encoding in order to provide / transmit the video data (ultimately to a remote decoder (490)) to a device storing the encoded video data via a hardware or software connection (495). According to an embodiment, when providing / transmitting the encoded video data from the video encoder (405), the transmitter (470) may receive and merge other data accompanying the encoded video data, for example, encoded audio data or other auxiliary data, from a separate source (480).
[0134] According to one embodiment of the present invention, the transmitter (470) may be configured to transmit additional data along with the encoded video. The additional data may be considered part of the encoded video data. The additional data may include information that can be used by a decoder to properly decode the data or to more accurately reconstruct an image that approximates the original image. Examples of the additional data may include all of the examples previously presented with respect to the receiver (310) of the decoder.
[0135] The present invention can be implemented by a digital video compression standard that is widely used and understood by those skilled in the art as described above. The digital video compression standard may include at least one of compression standards known by the standard name such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.
[0136] Fig. 6 is a conceptual diagram illustrating the structure of a video encoder according to another embodiment of the present invention. What is depicted in Fig. 6 may be a rough structure of a video encoder widely known as a standard code such as ITU-T H.266 and ISO / IEC 23090-3, and also known as MPEG-I Part 3 or versatile video coding (VVC).
[0137] According to FIG. 6, a video encoder (605) may be configured to receive uncompressed and unencoded original video data (601) as input and output an encoded bit stream (602). The video data (601) may be supplied directly to a luma mapping unit (610a) when intra-encoded, or may be supplied to a luma mapping unit (610b) via an inter-prediction unit (620) including motion vector extraction. In the intra-encoded case, the mapped luma signal may be supplied to an output merger (606) by selecting (608) at least one of an intra-prediction encoded signal via an intra-prediction unit (625) or an inter-prediction encoded signal output from the luma mapping unit (610b) via the inter-prediction unit (620). The result of the above output merger can be applied to a chroma scaling unit (615). (The operation of the luminance signal mapping unit (610) and the operation of the chroma scaling unit (615) are collectively referred to as a luma mapping / chroma scaling (LMCS) process.) The reduced chroma signal can be provided to a transform unit (630), and the transform unit (630) can perform an adaptive color transform, particularly on the chroma signal. The coefficients derived as a result of the transform are applied to a quantization unit (640) and quantized. As a result, lossy compression is achieved, and the result of the lossy compression can be output as a bit string (602) through a multi-hypothesis CABAC (650), which is a lossless compression method.
[0138] Meanwhile, the result of the lossy compression may actually enter the decoding process by going through the processes of inverse quantization (645), inverse transform (635), and luminance signal expansion (617) to generate a coding loop. The result of the luminance signal expansion may be supplied to the internal merger (607) together with the result of selecting (608) at least one of the previously generated intra prediction encoding signal or inter prediction encoding signal. The result of the internal merger may go through inverse luma mapping (617), and then may go through processing such as a deblocking filter (660), sample adaptive offset (SAO) (670), and an adaptive loop filter (ALF) to reproduce the image quality improvement process in the decoder. The result of reproducing the operation in the decoder as described above is applied to the reference picture buffer (690) and can be reused for prediction encoding by the inter prediction unit (620).
[0139] The present invention can also be utilized by or incorporated into an enhanced compression model (ECM), which is an implementation of a next-generation video codec currently being developed by the Joint Video Experts Team (JVET), an international standardization expert organization. The enhanced compression model can include an enhanced intra prediction coding method, an enhanced inter prediction coding method, an enhanced transform and transform coefficient coding method, an enhanced adaptive loop filtering method, a bilateral filtering method, a new sample adaptive offset (SAO) method for improving image quality, an extended entropy coding method, and an improved gradual decoding refresh (GDR) technique.
[0140]
[0141] Composition of the present invention
[0142] The present invention discloses a video encoding / decoding method and device for optimized post-processing that can be used in the video encoding and decoding field including the above-described embodiments.
[0143] FIG. 7 is a conceptual diagram of preprocessing and postprocessing of a video encoding / decoding process according to an embodiment of the present invention. In general, postprocessing (740) may have the purpose or usage or mode of improving (741) the image quality of a decoded frame after decoding (730) of the image. For the postprocessing, a suitable method may be applied depending on the display or application field after video decoding. For example, a neural network-based postprocessing filter (NNPF), which recently provides high efficiency, may be used. The postprocessing may have the purpose or usage or mode of improving image quality, increasing spatial / temporal resolution (i.e., upsampling), increasing bit depth, etc.
[0144] In addition to the post-processing aimed at improving image quality as described above, the present invention also considers post-processing to optimize the use, purpose, and mode of the decoded image. For example, the video encoding / decoding process may be an application information processing method that uses compression by video data as a medium, for example, implementation of immersive video (e.g., which may mean something like a technology called MPEG Immersive Video (MIV)) (742), implementation of point cloud (e.g., which may mean something like Point Cloud Compression (PCC)) (743), improvement of efficiency of computer vision mission (744), and other application methods that use video data as a medium that are not limited to the examples described above, etc., as a medium, and the like, may be the purpose of the post-processing.
[0145] For post-processing, there exists a technology that autonomously performs post-processing by analyzing the contents of decoded images or encoded information. However, since it is sometimes difficult to obtain optimized post-processing information using this method, a method of transmitting essential information required for post-processing in the form of a message during the encoding stage is also being used. For example, so-called supplemental enhancement information (SEI) (705) can be used. The SEI (705) is generated in the encoding (720) stage and / or the pre-processing (710) stage of encoding and can be transmitted and used together with the image (700), preferably as part of a fused bit stream (707), in the decoding (730) and / or post-processing (740) stage.
[0146] The SEI (705) used in the present invention should be understood as a general term for the high-level bitstream syntax (HLS) defined in various video encoding / decoding technical specifications and information representation methods equivalent thereto. For example, the SEI message format of the H.264 / AVC, H.265 / HEVC standards, the VSEI (versatile SEI) format defined to correspond to the H.266 / VVC standard, or other similar SEI message formats and / or additional information representation methods that are known in the past or may be newly provided should all be considered to be encompassed by the meaning of the SEI (705) indicated in the present invention. The present invention does not place any restrictions on the specific implementation of such information transmission formats, and any form of information transmission format suitable for information transmission for post-processing optimization can be utilized to implement the technical idea of the present invention.
[0147] According to one embodiment of the present invention, information about the physical characteristics of the display device included in the SEI (705) and / or related post-processing information may be used for the purpose of instructing the decoding device and / or the post-processor associated therewith to perform actual post-processing, but this is not necessarily the case. That is, the information may be used simply for the purpose of providing general information to the decoder side, or may be configured so that the decoder reads it in the process of processing the bit string, but the corresponding operation of the decoder and / or post-processor is not accompanied. However, if the information is included in the process of generating the auxiliary improvement information as a bit string and such auxiliary improvement information is configured so that it can be appropriately parsed by the decoder side, it should be considered to satisfy the configuration of the present invention.
[0148] According to one embodiment of the present invention, the processing of the SEI (705) included in the fused bit stream (707) can be performed in various ways. In one embodiment, the SEI (705) can be implemented to be generated in the encoding (720) step, and, if necessary, with reference to information in the preprocessing (710) step, so as to form a single fused bit stream (707) together with information about the image (700). In another embodiment, the SEI (705) can be implemented to be generated in the preprocessing (710) step, and then formed into a fused bit stream (707) by an information merging means including a high-level syntax or multiplexing in parallel with the image (700) information generated in the encoding (720) step. Likewise, in one embodiment, the SEI (705) may be input as a unit of a merged bit string (707) in the decryption (730) step, acquired and processed, and then, if necessary, all or part of it may be transferred to the post-processing (740) step. In another embodiment, the SEI (705) may be separated from the image (700) information by parsing or demultiplexing of high-level syntax, and all or part of it may be directly input to the post-processing (740) step.
[0149] According to one embodiment of the present invention, if preprocessing optimized for a specific display type has already been performed in the preprocessing (710) step, related information may be transmitted through the SEI (705), and the information related to the preprocessing may be considered for processing in the decoding (730) and / or postprocessing (740) steps. Specifically, if the decoded image is already optimized for a specific display type, the postprocessor may recognize the information and omit or adjust redundant or unnecessary postprocessing. In addition, detailed information regarding what type of optimization was performed in the preprocessing step may also be included in the SEI (705), and the information may be used for the purpose of determining or controlling a postprocessing method.
[0150] According to one embodiment of the present invention, the SEI (705) may include post-processing information for various purposes such as general visual quality improvement, chroma upsampling, resolution resampling, picture rate upsampling, bit depth upsampling, and colorization. In addition, when the NNPF is used, information describing the neural network of the NNPF may be included in the SEI (705). According to another embodiment of the present invention, when a general post-processing method that does not use a neural network is used, information indicating the post-processing method may be included in the SEI (705). It will be readily understood that various post-processing information may be included in the SEI (705) in addition to the examples described above.
[0151] According to one embodiment of the present invention, the SEI (705) may include information indicating a post-processing optimization type. This type information may include object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, privacy protection optimization, luminance range adaptation optimization, etc.
[0152] According to one embodiment of the present invention, the SEI (705) may include post-processing information and optimization information expressed in bit mask form, and may be configured to simultaneously transmit, apply, and / or activate multiple pieces of information or types. Additionally, additional parameter information for each optimization method and optimization type may be included in the SEI (705).
[0153] According to one embodiment of the present invention, the SEI (705) may include post-processing information according to display characteristics. Such display characteristics may be included in the SEI (705) in a form that designates the physical characteristics of the display device, and may include, for example, display model information such as transmissive pixel technology and / or emissive pixel technology. In this case, the transmissive pixel technology is used in displays such as LCD and refers to a method of expressing color by passing backlight, and the emissive pixel technology may refer to a method of expressing color by using each pixel independently emitting light in displays such as OLED and QLED. In addition to the pixel technologies described above, various physical characteristics of displays may be considered, and may include, for example, stereoscopic displays, holographic displays, volumetric displays, and other display characteristics that are not limited in this specification. Since these different display technologies differ in color expression range, contrast ratio, response speed, energy consumption pattern, etc., post-processing tailored to the display characteristics may be required to optimize the final display result of the image, and the SEI (705) can be configured to provide information that allows such post-processing to be performed appropriately.
[0154] According to another embodiment of the present invention, the post-processing information regarding the physical characteristics of the display device included in the SEI (705) may be used as a signal to indicate or provide information that the image has already been optimized for a specific display type during the pre-processing process before the image is encoded. As described above, the purpose of the post-processing (740) in the present invention is not limited to improving the image quality (741) of general video (moving image), and may correspond to a separate application information processing method using video data as a medium.
[0155] FIG. 8 is an exemplary diagram illustrating a video-mediated application information processing method according to one embodiment of the present invention. The image in the present invention is not limited to a typical two-dimensional color image, and may include various forms of image information, such as those illustrated in FIG. 8. That is, the image in the present invention may refer to an image (810) that is subject to encoding / decoding by a video codec, and further, the image (810) may further include auxiliary information.
[0156] According to one embodiment of the present invention, an image in the present invention, such as the image (810), may include a video atlas. The video atlas refers to a video frame composed of a plurality of distinct image regions acquired from the same or different sensors. Specifically, the video atlas can be understood as a two-dimensional image that packs image data from multiple viewpoints into a single image for efficiently representing three-dimensional image data, such as multi-view images or immersive videos.
[0157] According to another embodiment of the present invention, an image in the present invention, for example, the image (810) may include tristimulus color components along with alpha, occupancy information, and / or depth components. Here, occupancy information refers to data indicating whether valid information exists in a specific pixel or region within the image, and this can be used importantly to indicate the validity of each region, especially in a form such as a patch-based atlas. The alpha component refers to transparency information of the image, and can be utilized in various image synthesis and processing processes.
[0158] According to another embodiment of the present invention, an image in the present invention, for example, the image (810) may be treated as an image including auxiliary components for representing the structure of the image, such as stereoscopic 3D, multi-view, depth, alpha, or an object instance ID map.
[0159] According to one embodiment of the present invention, an image in the present invention, for example, the image (810), in addition to image data, may additionally include entity information (820) indicating depth or object information indicating the distance between a camera and an object, a patch-based atlas (830) representing a multi-view image, a feature or feature map (840) extracted from a neural network parameter and / or image information from a neural network, a point cloud (850), hologram implementation data (860) composed of information such as real, imaginary, and phase, and information on the usage and / or usage purpose of post-processing including a computer vision mission including an image recognition task, and such post-processing information may be generated in the pre-processing step and / or the encoding step and may be transmitted by being included in the SEI. According to another embodiment of the present invention, when a general post-processing method that does not use a neural network is used, parameter information for the post-processing method may be included in place of information (840) related to parameters of the neural network, etc.
[0160] In a more specific example, the post-processing information may include information regarding the intended use of the image. For example, some images may be intended for a specific use, such as a flat (2D) display, a hologram, a metaverse (virtual reality), or a computer vision task. In another example, some images may be intended for human viewing, while others may be intended for computer processing.
[0161] As another specific embodiment, the post-processing information may include information on an additional encoding / decoding method. For example, a certain image may be processed as a kind of pre-processing by another application information processing method before being encoded as a video, and may be re-processed as a kind of post-processing by another application information processing method after being decoded as a video. For example, an image may be processed from data and / or images of an original shape using standard technologies such as MIV (MPEG Immersive Video), INVR (Implicit Neural Visual Representation), NNC (Neural Network Compression), VCM (Video Coding for Machines), or PCC (Point Cloud Coding) before encoding, and thus, this original shape may need to be restored or utilized in the post-processing process after decoding the image. The above processing may include, for example, re-expressing an original 3D-based image as a patch-based atlas, or re-expressing 3D-based data in a point cloud format, or extracting and processing depth information of an image, or re-expressing neural network parameters, features obtained by a neural network, or feature maps as an image, or applying a non-neural network post-processing filter based on predetermined parameters, or separating multiple objects from the original given image data and configuring them into multiple layers and / or entities, and may represent various processing methods other than those described above.
[0162] According to one embodiment of the present invention, when information of another compression standard is required for post-processing, information defining whether information used in the other compression standard is used may be included. For example, when it is determined that information derived from the standard technology needs to be transmitted through the SEI in conjunction with processing by another standard technology including MIV, INVR, NNC, VCM, or PCC, information such as syntax defined by the standard technology may be included in the SEI information in the same form or in a form derived therefrom. Referring again to FIG. 7, the processing may correspond to the pre-processing (710) step, and the syntax information (e.g., 703) generated in the pre-processing (710) step may be included in the SEI (705) in the same form or after re-processing, and thus transmitted through a bit string (707). As described above, the syntax information included in the SEI (705) may be used in the post-processing (740) step corresponding to the processing.
[0163] As another specific embodiment, the post-processing information may include information regarding the display and / or processing method of the image involved in the post-processing. For example, the post-processing may include information indicating that the purpose of the post-processing is 3D viewpoint generation (image generation), image merging, or stereoscopic effect. As another example, the post-processing may include information indicating that the purpose of the post-processing is a vision task such as object recognition / detection / segmentation, information retrieval, or captioning. As another example, the post-processing may include information indicating that the purpose of the post-processing is a special effect such as blurring or noise addition. As another example, the post-processing may include information related to the configuration of a post-processing method, a post-processing filter, or an NNPF that optimizes the purpose of the post-processing. As another example, the post-processing may include information indicating that the image to be subjected to the post-processing is an image having a specific purpose, including a face, fingerprint, gene, game, or screen content.
[0164] According to one embodiment of the present invention, at least one post-processing method can be grouped into a post-processing group according to its characteristics and purpose. The post-processing group can be defined as a set of at least one post-processing method, and each group can be defined to share a common purpose or characteristic. For example, the post-processing group can be classified into a neural network-based post-processing group, a post-processing group for image quality improvement, a post-processing group for computer vision optimization, and a post-processing group for display characteristics optimization. Each post-processing group includes different algorithms and parameter sets, and can be configured as a collection of methods optimized for a specific post-processing purpose.
[0165] The above post-processing groups can be hierarchically structured, and the relationships between upper and lower groups, or between individual groups that are performed in parallel / serially, can be defined. For example, the image quality enhancement group can be divided into a spatial image quality enhancement subgroup and a temporal image quality enhancement subgroup, and each subgroup can be further structured to include more specific methods. Through this hierarchical grouping, the decoder can be configured to select and apply a chain structure of appropriate post-processing methods depending on the situation. In particular, the inter-application, order, and dependency between various post-processing methods can be defined by the structure of the post-processing group and / or the structure of the hierarchical post-processing group.
[0166] As in the examples described above, information necessary for post-processing an image according to its corresponding usage and purpose may be transmitted by being included in the SEI. The information may be generated by an encoder and / or a corresponding preprocessor, preferably included in the SEI and transmitted through a bitstream, and configured to be processed by a decoder and / or a post-processor. According to one embodiment, information for determining one or more post-processing filters or post-processing methods may be included in the SEI, or information for determining one or more post-processing filters or post-processing groups may be included in the SEI, depending on the type of the image, the state in which the image is expressed, the purpose for which the image was acquired, information on whether other standards were used in encoding / decoding, etc. As described above, various post-processing filters or post-processing methods may be included in the group, and the implementation method of the present invention is not limited thereto.
[0167] According to one embodiment of the present invention, the post-processing information may include information indicating that the original image was initially created for a specific display type. This may imply that the image was not simply optimized through post-processing, but rather was created from the image acquisition or generation stage with consideration given to the characteristics of a specific display technology.
[0168] FIG. 9 is an exemplary diagram showing an example of using the MIV standard according to one embodiment of the present invention. Referring to FIG. 9, a multi-view image (910) can be expressed as a patch-based atlas (920) by a preprocessing encoder (915) according to the MIV standard, and the corresponding information can be encoded by a video encoder (925) following a standard such as VVC (Versatile Video Coding). At this time, an SEI (932) can be attached to the encoded video bitstream (930), and the SEI can include at least one processing information (917) derived in the preprocessing encoder (915) step of expressing the multi-view image as a patch-based atlas according to the MIV standard. When the decoder (935) decodes the above image and then performs post-processing, since the information used in the MIV is available through SEI, it can be reused in the decoding process (935) according to the VVC standard and used for post-processing of the image, or it can be configured to be referenced (937) in restoring the decoded image (940) back into a multi-view image (950) by a post-processing decoder (945) according to the MIV standard.
[0169] According to one embodiment of the present invention, the post-processing that can be executed within the decryption (935) process or the post-processing decoder (945) may be configured to perform different types of post-processing depending on the generated patch type. In this case, the degree of the different types of post-processing methods may be indicated by the SEI (932).
[0170] According to one embodiment of the present invention, the patch-based atlas (920) is a special type of image data processed by the MIV standard, and may have a data structure prepared for efficiently expressing multi-view image data. According to one embodiment, the atlas (920) may include occupancy information. The occupancy information may be used to distinguish between an area where actual valid pixel data exists and a padding or empty area within the atlas, and may be expressed in the form of a binary map indicating the validity of each pixel within the atlas.
[0171] According to one embodiment of the present invention, the SEI (932) may include type information related to the atlas (920), and the type information may include information indicating that the corresponding image is an atlas image. The type information supports the decoder (935) to recognize that the image is an atlas and apply appropriate decoding and / or post-processing based on it. In addition, according to one embodiment, the SEI (932) may further include data regarding the presence or absence of the occupancy information and the characteristics of the information.
[0172] According to one embodiment of the present invention, the SEI (932) may include information indicating the presence and characteristics of auxiliary components such as an alpha channel, a depth map, and an object instance ID. The auxiliary components may include information that enables the decoding process (935) and / or the post-processing decoder (945) operating based on the MIV standard to effectively restore and post-process a multi-view image (950).
[0173] According to one embodiment of the present invention, when generating a patch in the MIV standard, images from multiple viewpoints may be compared to each other to select dissimilar regions, and the target for patch generation may be selected by considering the dispersion, outline, complexity, etc. of each region or pixel within a color or depth image. In some embodiments, an image from one viewpoint may be synthesized using images from multiple viewpoints, and then the generated image may be compared with the original image to select the target. In some embodiments, a neural network may be used to determine whether to generate a patch. In this case, the reference or neural network information for the patch generation may be transmitted and received by being included in the bit string (930) and / or the SEI (932). In some embodiments, a color image or a depth image may be corrected using the neural network, and then a determination may be made as to whether to generate a patch using the corrected image. In some embodiments, a depth image may be corrected based on a color image, and then a determination may be made as to whether to generate a patch using the corrected depth image. In some embodiments, a color image may be corrected based on a depth image, and then a determination may be made as to whether to generate a patch using the corrected color image.
[0174] According to one embodiment of the present invention, when selecting a patch generation target by considering an outline in the step of determining whether to generate a patch, an outline detector filter including a Sobel and / or a Canny filter may be applied to each pixel in the depth image to derive the intensity of the outline, and if the outline intensity is greater than a certain threshold, it may be determined as an outline and a patch may be generated. At this time, depending on the embodiment, it may be configured to use two or more thresholds. Depending on the embodiment, the final outline may be determined by performing an additional correction operation on the outline detected by the outline detector. Depending on the embodiment, it may be configured to use the corrected depth image after correcting the depth image. Depending on the embodiment, whether or not a depth image has an outline may be determined based on a neural network. Depending on the embodiment, whether or not a depth image has an outline may be determined by comparing the outline intensity with surrounding pixel values. Depending on the embodiment, whether or not a depth image has an outline may be determined by considering both a color and a depth image.
[0175] Fig. 10 is an exemplary diagram illustrating the use of neural network information according to one embodiment of the present invention. Referring to Fig. 10, according to one embodiment of the present invention, a feature (1020) may be extracted from an original image (1010) using a neural network (1015) for object detection, and the feature may be encoded by an image encoder (1025) following a standard such as VVC. At this time, an SEI (1032) may be attached to the encoded image bitstream (1030), and the SEI (1032) may include extracted information (1017) related to the neural network and / or feature. Accordingly, when the feature is decoded as an image (1040) by the decoder (1035) and post-processing is performed by the neural network (1045), the neural network or feature extraction information derived from the SEI (1032) may be referenced (1037) to perform tasks such as object recognition (1050).
[0176] According to one embodiment of the present invention, when the compression unit to which the post-processing filter or post-processing method is applied is different, information on whether post-processing is performed for each frame, slice, sub-picture, tile, or any arbitrary region can be transmitted and received, and post-processing information for each frame, slice, sub-picture, tile, or any arbitrary region can be transmitted and received, preferably through the SEI, and information on whether the same post-processing filter or post-processing method is used for each arbitrary image segmentation unit such as the same frame, slice, sub-picture, tile, block, and coding unit or whether different post-processing filters or post-processing methods are used can be transmitted and received, and information on whether post-processing is performed for each region of interest, image patch, or any arbitrarily defined region can be transmitted and received, and information indicating horizontal / vertical coordinates and horizontal / vertical sizes for each region can be transmitted and received, and when pixels outside the range (boundary) to which the post-processing filter or post-processing method is applied are needed, information on whether padding is performed or a padding method, etc. can be transmitted and received.
[0177] According to one embodiment of the present invention, when a post-processing filter or a post-processing method is used for the purpose of image generation, information indicating a reference / target viewpoint position, a reference / target temporal position, or an image generation method may be transmitted and received. The image generation may refer to a case where an additional image must be generated from a decoded image in a post-processing process, and may include, but is not limited to, a process of interpolating and generating an image of a viewpoint that has not been input when reconstructing a multi-viewpoint image according to the MIV standard or the like. In the case where the post-processing has the purpose of generating an image as described above, information indicating camera parameters, positions, or arrangement methods required for image generation may be preferably transmitted and received through the SEI, information indicating position information for applying post-processing such as decoding, image generation, and post-processing may be transmitted and received, information indicating a method for filling in unfilled pixel values in the generated image may be transmitted and received, or some of the above information may be configured to be transmitted and received by referencing or deriving syntax information of a bit string generated by another technology, for example, another standard technology used in the pre-processing process.
[0178] According to one embodiment of the present invention, a post-processing filter or post-processing method may be determined depending on the type of image, such as a general natural image, a computer-generated image, an atlas, a feature map, a point cloud, or a hologram. For example, a natural image is an image of the real world captured by a camera, and general post-processing for noise removal and image quality improvement may be appropriate. For another example, a computer-generated image is an image generated through 3D rendering or graphics processing, and may include so-called screen content images, and post-processing such as aliasing removal or texture enhancement may be appropriate. For another example, an atlas may refer to a multi-view image expressed patch-based in the MIV technique, as described above in FIG. 9, and post-processing for patch boundary blurring or removal of inconsistencies between patches may be appropriate. For another example, a feature map is an image that visualizes feature information of an object extracted from a neural network, as described above in FIG. 10, and specialized post-processing for object recognition or classification may be appropriate. For another example, a point cloud is data representing points in three-dimensional space, and post-processing for point density adjustment or surface reconstruction may be appropriate. For another example, a hologram may be composed of multiple component images, including real, imaginary, and phase images, and independent post-processing appropriate for each image composition may be appropriate.
[0179] Information specific to the image type described above can be included and transmitted in the SEI, and the decoder can be configured to select and apply an optimal post-processing method based on this information. Furthermore, specific methods may vary within a post-processing group depending on the image type. For example, within a neural network-based post-processing group, different neural network structures and parameters may be applied to natural images and feature maps.
[0180] According to another embodiment of the present invention, even within the same image, different regions may have different characteristics, and in this case, different post-processing methods may be applied to each region. For example, natural and computer-generated image portions may coexist within a single image, and by applying optimized post-processing to each region, the quality of the overall image can be improved. To this end, the SEI may include information on the characteristics of each region and the post-processing method to be applied to each region. Furthermore, optimized post-processing that considers both display characteristics and image type is also possible. For example, when a natural image is displayed on an emissive display and when it is displayed on a transmissive display, post-processing that takes into account different color expressions and luminance characteristics is required. Post-processing optimization that takes these complex conditions into account can be achieved by comprehensively utilizing the various pieces of information contained in the SEI.
[0181] According to one embodiment of the present invention, a post-processing filter or post-processing method may be determined according to an image generation method, a viewpoint position, a temporal position, an area position, an area characteristic, or a restored image quality.
[0182] According to one embodiment of the present invention, when a post-processing filter or post-processing method is used for a vision task, information indicating the vision task may be transmitted and received. The information may indicate the type of vision task, including object recognition, object detection, object segmentation, information retrieval, and captioning. In addition, information indicating the location of post-processing application, such as decoding, vision task, and post-processing, may be transmitted and received, and the same or different post-processing filters or post-processing methods may be used depending on the vision task. In addition, the post-processing filter or post-processing method may be determined based on the vision task, pre-processing characteristics, region location, region characteristics, or restored image quality. In addition, information indicating the location and application method of post-processing according to a region of interest (ROI) within an image may be transmitted and received, and a post-processing filter or post-processing method may be determined based on the location and application method of post-processing. At least a portion of the information exemplified above may correspond to information derived from other technologies that handle the corresponding information, for example, information derived from a preprocessing process performed prior to the image encoding process, or information derived from such information. In one embodiment, the information derived from the other technology may be included in the SEI in its original form or in the form of information derived therefrom and transmitted from the image encoder to the image decoder and / or from the preprocessor to the postprocessor.
[0183] According to one embodiment of the present invention, when a post-processing filter or a post-processing method is used for a hologram, the post-processing filter or the post-processing method may be determined according to the type of the hologram real image, imaginary image, or phase image, or according to the quality of the restored image. In addition, information indicating the position information of the post-processing application, such as decoding, hologram generation, and post-processing, may be transmitted and received. According to another embodiment of the present invention, at least a part of the information necessary for hologram generation may be referenced from or derived from bitstream syntax information generated by a technology related to hologram processing, such as another standard technology, and may be configured to be transmitted and received, preferably by including it in the SEI.
[0184] FIG. 11 is an exemplary diagram illustrating a method for optimizing a post-processing filter and a post-processing method according to an embodiment of the present invention. Referring to FIG. 11, in an embodiment of the present invention, when a post-processing filter or a post-processing method optimized for a vision mission is different, a different post-processing filter or a post-processing method may be configured to be used according to vision mission objective information, post-processing objective information, or post-processing information. According to a preferred embodiment of the present invention, information specifying the post-processing filter or the post-processing method may be inserted in the form of an SEI into a bit stream (1120) transmitted by an encoder (1110) according to a video compression standard (e.g., VVC), and a bit stream including the SEI may be received by a corresponding decoder (1130), and then transmitted to a post-processor (1140) along with an image decoded by the decoder (1130). The post-processor (1140) may be configured to analyze post-processing information (1142) and select and execute an appropriate post-processing method (1145). Depending on the application result of the post-processing method (1145), corresponding desirable vision task operations (1150), such as object recognition, object detection, and object segmentation, may be individually executed. Of course, it is self-evident that various post-processing methods and vision tasks as post-processing purposes other than those described above may exist, and all of these are included within the technical spirit of the present invention.
[0185] FIG. 12 is an exemplary diagram illustrating another method for optimizing a post-processing filter and a post-processing method according to an embodiment of the present invention. Referring to FIG. 12, in an embodiment of the present invention, when the post-processing filter or post-processing method optimized for a plurality of image components constituting one image is different, for example, when the post-processing filter or post-processing method optimized for real images, imaginary images, etc. constituting a holographic image is different, it may be configured to use different post-processing filters or post-processing methods depending on the image type or post-processing information. According to a preferred embodiment of the present invention, the information specifying the post-processing filter or post-processing method may be inserted in the form of an SEI into a bit stream (1220) transmitted by an encoder (1210) according to a video compression standard (e.g., VVC), and a bit stream including the SEI may be received by a corresponding decoder (1230), and then transmitted to a post-processor (1240) together with an image decoded by the decoder (1230). The post-processor (1240) may be configured to analyze (1242) the post-processing information and select and execute an appropriate post-processing method for each partial image information constituting the hologram. For example, if the hologram image is composed of a real image and an imaginary image, a first post-processing (1245) that is preferable for the real image and a second post-processing (1246) that is preferable for the imaginary image may be selectively executed according to the type of the image. The result of applying the post-processing (1245, 1246) may be utilized in the process of generating an image (1250). For example, in the case of a hologram composed of the real image and the imaginary image, the real image and the imaginary image may be merged using a corresponding hologram technology and utilized to generate a hologram image (1250). Of course, it is obvious that a method of implementing a composite image using various image components and post-processing methods other than those described above may be used.Furthermore, the above description of FIG. 12 does not exclude the possibility of implementing additional and detailed image configurations such as the third and fourth ones and corresponding post-processing methods.
[0186] According to one embodiment of the present invention, the SEI may include various types of information depending on the physical characteristics of the display on which the image is to be displayed. In addition to the display for the holographic image described above, display models such as a transmissive pixel display and an emissive pixel technology may be considered, and also, taking into account the display characteristics of a stereoscopic display, a volumetric display, and other displays not limited thereto, other post-processing filters or post-processing methods may be configured to be used depending on the provision of appropriate post-processing information. Accordingly, the first post-processing (1245), the second post-processing (1246), and / or the image generation (1250) processes illustrated in FIG. 12 should be understood to comprehensively include post-processing and image generation methods corresponding to image configurations corresponding to each display type.
[0187] FIG. 13 is an exemplary diagram showing an application example of input information for a plurality of post-processors according to one embodiment of the present invention. According to one embodiment of the present invention, when one or more post-processors are used for post-processing, information for each post-processor may be configured to be provided. Referring to FIG. 13, a decoder (1310) may be configured to extract information included in an SEI accompanying an image bitstream after decoding an image, and / or information of a decoding process derived from the image bitstream, such as information including quantization parameter information, motion information, mode information, or in-loop filtering information, and provide each piece of information to an appropriate post-processor (1320, 1330). For example, first post-processing information (1325) may be provided to a first post-processor (1320), and second post-processing information (1335) may be provided to a second post-processor (1330). According to one embodiment of the present invention, at least one of the plurality of post-processors (1320, 1330) may be a neural network-based post-processor. In this case, for each of the plurality of post-processors (1320, 1330), usage or usage purpose information corresponding to at least one neural network configuration and / or neural network group may be transmitted and received, and parameter information of each neural network or neural network application location information may be provided. According to one embodiment of the present invention, at least one of the plurality of post-processors (1320, 1330) may be a post-processor that is not based on a neural network. In this case, for each of the plurality of post-processors (1320, 1330), usage or usage purpose information corresponding to at least one post-processing method may be transmitted and received by an information transmission means including the SEI, and further, parameter information or post-processing application location information for each post-processing method may be provided.In addition, when the above post-processing group is used, for each of the plurality of post-processors (1320, 1330), usage or usage purpose information corresponding to at least one post-processing group composed of at least one post-processing method can be transmitted and received by an information transmission means including the SEI, and parameter information or post-processing application location information for each post-processing group can be provided.
[0188] According to one embodiment of the present invention, when a general post-processing method and a neural network-based post-processing method are mixed and used in the plurality of post-processors (1320, 1330), information on whether a neural network is used and / or the post-processing method may be transmitted and received. According to an embodiment, when a neural network-based post-processing method is used in at least one post-processor, neural network information may be additionally transmitted and received according to the corresponding information. According to an embodiment, a post-processing group may be defined depending on whether a neural network is used or regardless of whether a neural network is used. According to an embodiment, in the case of a post-processing group using a neural network, neural network information may be additionally transmitted and received for each post-processing method. According to an embodiment, post-processing information or application location information for each post-processing group or each post-processing method may be transmitted and received.
[0189] FIG. 14 is an exemplary diagram showing an application example of different post-processing methods according to one embodiment of the present invention. According to one embodiment of the present invention, the same decoded image can be used by distinguishing it into a video for viewing and a video for non-mission based on post-processing information. Referring to FIG. 14, a decoder (1430) decodes an image, and then extracts information included in an SEI accompanying an image bitstream (1420), and / or information of a decoding process derived from the image bitstream (1420), such as information including quantization parameter information, motion information, mode information, or filtering information within a loop, and according to an analysis (1440) of post-processing information depending on whether the purpose of use of the image is viewing by a human or execution of a vision task by a computer device, supplies necessary information to a corresponding appropriate post-processor (1445, 1446), and operates the corresponding post-processor. Preferably, the distinction for the purpose of use can be made by information derived from the process of preprocessing performed in the encoder (1410) that encodes the image (1420) to be decoded by the decoder (1430) and / or in the preprocessor (1405).
[0190] According to one embodiment of the present invention, information regarding the physical characteristics of a display device and / or post-processing information can be utilized in various ways. The decoder (1430) may receive and read the information, and then perform post-processing on its own based on the information, transmit the information to the display device and / or post-processing device to perform corresponding processing, simply recognize and / or display the information and do not perform any special processing, or decide not to apply additional processing to an already optimized image based on the information.
[0191]
[0192] Although the present invention has been described with reference to drawings and embodiments, as already mentioned above, it does not mean that the scope of protection of the present invention is limited to the drawings or embodiments presented above, and it will be understood that a person skilled in the relevant technical field can modify and change the present invention in various ways within a scope that does not depart from the spirit and scope of the present invention described in the claims of the present invention patent.
Claims
1. In the video decryption method, A step of receiving a bitstream including an encoded image and supplemental enhancement information regarding the encoded image; A step of extracting the encoded image and the auxiliary enhancement information from the bit string; and A step of decoding the encoded image to obtain a decoded image; including, An image decoding method, characterized in that the above auxiliary improvement information includes post-processing information according to the physical characteristics of the display device.
2. In paragraph 1, A method for decoding an image, characterized in that the physical characteristics of the display device include display model information including transmissive pixel technology and emissive pixel technology.
3. In paragraph 1, An image decoding method characterized in that the above auxiliary improvement information is included and expressed in the form of a bit mask, and a plurality of post-processing information or post-processing types are configured to be applied simultaneously.
4. In paragraph 1, A video decoding method, characterized in that the auxiliary improvement information includes information indicating at least one type of post-processing optimization among object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, privacy protection optimization, luma range adaptation optimization, video atlas optimization, occupancy processing optimization, and multi-view synthesis optimization.
5. In paragraph 1, An image decoding method, further comprising a step of performing post-processing based on the above auxiliary improvement information.
6. In paragraph 5, The step of performing post-processing based on the above auxiliary improvement information is: An image decoding method, comprising: a step of selecting and applying different post-processing methods according to the physical characteristics of the display device included in the above auxiliary improvement information.
7. In paragraph 1, At least one post-processing method is grouped into a post-processing group, An image decoding method, characterized in that the auxiliary improvement information includes information about the post-processing group.
8. In paragraph 1, The above encoded image is an image processed from a first media format into an image using a preprocessing technique, The above post-processing includes a process of restoring the decrypted image to the first media format, An image decoding method, characterized in that the auxiliary improvement information includes information related to at least one post-processing method related to restoring the decoded image into the first media format.
9. In paragraph 8, The first media format includes a video atlas, The above video atlas comprises video frames composed of multiple distinct image regions acquired from the same or different sensors, A video encoding method, characterized in that the auxiliary improvement information includes occupancy information indicating the validity of each pixel in the video atlas.
10. In paragraph 1, An image decoding method, characterized in that the auxiliary improvement information includes information indicating different post-processing methods for at least two geometric segmentation areas defined within a single encoding unit.
11. In the video encoding method, A step of generating post-processing information for an input image; A step of generating auxiliary improvement information for the input image based on the post-processing information; A step of encoding the input image to obtain an encoded image; and A step of generating a bit string including the encoded image and the auxiliary enhancement information; An image encoding method, characterized in that the auxiliary improvement information includes post-processing information according to the physical characteristics of the display device.
12. In paragraph 11, A method for encoding an image, characterized in that the physical characteristics of the display device include display model information including transmissive pixel technology and emissive pixel technology.
13. In paragraph 11, An image encoding method characterized in that the above auxiliary improvement information is included and expressed in the form of a bit mask, and a plurality of post-processing information or post-processing types are configured to be applied simultaneously.
14. In paragraph 11, A video encoding method, characterized in that the auxiliary improvement information includes information indicating at least one type of post-processing optimization among object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, privacy protection optimization, luma range adaptation optimization, video atlas optimization, occupancy processing optimization, and multi-view synthesis optimization.
15. In paragraph 11, At least one post-processing method is grouped into a post-processing group, A method for encoding an image, characterized in that the auxiliary improvement information includes information about the post-processing group.
16. In paragraph 11, An image encoding method, further comprising: a step of performing preprocessing on the input image.
17. In paragraph 16, The step of performing preprocessing on the input image includes a step of preprocessing the input image in the first media format and processing it into an image to be encoded; A video encoding method, characterized in that the auxiliary improvement information includes information related to at least one post-processing method related to restoring the decoded video into the first media format.
18. In paragraph 17, The first media format includes a video atlas, The above video atlas comprises video frames composed of multiple distinct image regions acquired from the same or different sensors, A video encoding method, characterized in that the auxiliary improvement information includes occupancy information indicating the validity of each pixel in the video atlas.
19. In paragraph 11, An image encoding method, characterized in that the auxiliary improvement information includes information indicating different post-processing methods for at least two geometric segmentation areas defined within a single encoding unit.
20. In the video decryption device, A receiving unit that receives a bitstream including an encoded image and supplemental enhancement information regarding the encoded image; A parsing unit that extracts the encoded image and the auxiliary improvement information from the bit string; A decoding unit that decodes the encoded image to obtain a decoded image; and A post-processing unit that reads at least one post-processing information from the above auxiliary improvement information; An image decoding device, characterized in that the above auxiliary improvement information includes post-processing information according to the physical characteristics of the display device.
Citation Information
Patent Citations
Method and system for generating region nesting messages for video pictures
JP6816166B2
Apparatus and method for transmitting and receiving 3DTV broadcasting
KR1020160062716A
Support of multi-mode extraction for multi-layer video codecs
KR102054040B1
Manufacturing method of whole wheat bread using solomon's seal tea
KR102490839B1
KR20220106101A
Cited By
Compression methods, devices, and storage media for object maps
CN122412377A
Compression methods, devices, and storage media for object maps
CN122412377B