Improved decoder-side intra prediction mode estimation method and device for video encoding and decoding
The improved merged DIMD method addresses the challenge of high compression ratio and complexity in video encoding and decoding by combining neighboring block histograms for intra prediction, resulting in enhanced efficiency and quality.
Patent Information
- Application Number
- PCT/KR2025/004505
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-03
- Filing Date
- 2025-04-03
- Publication Date
- 2025-10-09
AI Technical Summary
Existing video encoding and decoding technologies face challenges in achieving a high compression ratio while minimizing image quality degradation and reducing computational complexity, particularly in decoder-side intra mode derivation (DIMD) methods.
An improved merged DIMD method that combines histogram information from neighboring blocks to derive intra prediction modes, enhancing encoding and decoding efficiency by generating a merged histogram for current blocks.
The method achieves a desirable compromise between compression ratio and complexity, improving encoding and decoding efficiency, reducing computational resources, and enhancing video quality.
Smart Images

Figure KR2025004505_09102025_PF_FP_ABST
Abstract
Description
Improved decoder-side intra prediction mode estimation method and device for video encoding and decoding
[0001] The present invention relates to video compression technology, and more particularly, to an improved application method and utilization of decoder-side intra mode derivation (DIMD) technology, which is one of the intra prediction application methods that contributes to improving compression performance in a video encoder and decoder.
[0002] The present invention may be in the same technical field as at least one of the digital video compression technology standards known by the names of standards such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG, or may be in the technical field for improving the inherent efficiency of the standards, or may be in the technical field for improving or replacing the standards.
[0003] Digital video encoding and decoding are widely used in various digital video applications. For example, digital television broadcasting, video transmission through communication networks, video calls / video conversations / video chats, recording and providing video content using optical media including video compact discs (VCDs) / digital versatile discs (DVDs) / Blu-Rays, all processes for producing, editing, collecting, and distributing video content, and devices such as video recording devices and camcorders for shooting and recording video for various reasons including personal, commercial, industrial, and security purposes, all depend on video encoding and decoding technologies.
[0004] Accordingly, implementations that may be referred to as digital video encoders and decoders may form part of a wide range of devices related to the generation, recording, and provision of digital video, including digital televisions, digital broadcasting systems, wireless broadcasting systems, computers in the form of notebooks / desktops / tablets, e-book readers, digital cameras, digital recording devices, digital multimedia playback devices, video game devices / terminals / consoles, mobile phones (including smartphones) with multimedia playback capabilities, equipment for video conferencing, and other devices.
[0005] The above digital video encoders and decoders can be implemented by a digital video compression standard that is widely used and understood by those skilled in the art. The digital video compression standard may include at least one of compression standards known by a standard name such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.
[0006] Video encoders and decoders can be implemented to more efficiently encode or decode digital video information while complying with the above standards, or by improving or modifying the above standards. Attempts to modify the above standards can also lead to the development of new standards. A well-known example is the so-called enhanced compression model (ECM), an attempt to improve and replace the existing H.266 / VVC standard, currently being developed by the Joint Video Experts Team (JVET), a joint international standardization group of ISO, IEC, and ITU-T.
[0007] Among the various detailed technologies applied to conventional video encoders and decoders, there is a technology collectively called decoder-side intra mode derivation (DIMD). DIMD proposes a method of calculating the intra prediction direction of the current block from the values of previously decoded images adjacent to the block currently being encoded / decoded, in order to reduce the amount of information transmitted from the encoder to the decoder regarding the intra prediction mode during the process of increasing the compression ratio of an image through intra prediction, i.e., intra prediction. Therefore, even without the encoder providing a separate signal, a decoder in DIMD mode can autonomously derive intra prediction direction information from the preceding decoding process, resulting in improved compression efficiency. Furthermore, a technology has been proposed for implementing a merged DIMD (DIMD Merge), which utilizes DIMD to quote the results of DIMD processing from a previously encoded / decoded block and reuse them in subsequent blocks.
[0008] Digital video is known to require a significant amount of information to describe its content in its uncompressed state. Therefore, recording or transmitting this information in its original form can be inefficient. Therefore, digital video is compressed using various methods prior to recording or transmission. These compression methods include lossy and lossless encoding. Lossy encoding sacrifices some image quality to achieve high compression performance, while lossless encoding sacrifices some compression performance to prevent image quality degradation. Regardless of the encoding method, there is a need to implement a technology that achieves a high compression ratio while minimizing image quality degradation to meet the growing demand for high-quality digital video within the constraints of limited memory storage capacity and communication transmission bandwidth.
[0009] As described above, the encoding process for compression requires various operations, such as spatial segmentation of digital video, segmentation and / or processing in color channels, removal of spatial redundancy, removal of temporal redundancy, tracking of motion vectors within the video, encoding of differential images, quantization, coefficient scan, run-length coding, entropy coding, and loop filtering. These encoding operations generally consume computing resources and take a certain amount of time to complete. Similarly, the decoding operation for the encoding operation also requires certain computing resources and a certain amount of time. The main goal of video encoding and decoding technology is to ensure that the above resource consumption and time consumption do not interfere with the production, recording, distribution, and viewing of digital videos.
[0010] Accordingly, the present invention provides a new technology that can contribute to at least one of the following: improvement of encoding efficiency, improvement of decoding efficiency, improvement of video quality, reduction of computational amount, reduction of software size, reduction of hardware size, and improvement of other performances related to encoding and decoding in the technical tasks in the field of video encoding and decoding.
[0011] Digital video is known to require a significant amount of information to describe its content in its uncompressed state. Therefore, recording or transmitting this information in its original form can be inefficient. Therefore, digital video is compressed using various methods prior to recording or transmission. These compression methods include lossy and lossless encoding. Lossy encoding sacrifices some image quality to achieve high compression performance, while lossless encoding sacrifices some compression performance to prevent image quality degradation. Regardless of the encoding method, there is a need to implement a technology that achieves a high compression ratio while minimizing image quality degradation to meet the growing demand for high-quality digital video within the constraints of limited memory storage capacity and communication transmission bandwidth.
[0012] As described above, the encoding process for compression requires various operations, such as spatial segmentation of digital video, segmentation and / or processing in color channels, removal of spatial redundancy, removal of temporal redundancy, tracking of motion vectors within the video, encoding of differential images, quantization, coefficient scan, run-length coding, entropy coding, and loop filtering. These encoding operations generally consume computing resources and take a certain amount of time to complete. Similarly, the decoding operation for the encoding operation also requires certain computing resources and a certain amount of time. The main goal of video encoding and decoding technology is to ensure that the above resource consumption and time consumption do not interfere with the production, recording, distribution, and viewing of digital videos.
[0013] Accordingly, the present invention provides a new technology that can contribute to at least one or more of the following: improvement of encoding efficiency, improvement of decoding efficiency, improvement of video quality, reduction of computational amount, reduction of software size, reduction of hardware size, and improvement of other performances related to encoding and decoding in the technical problems in the field of video encoding and decoding as described above. In particular, the present invention proposes an improved merged DIMD method to overcome the limitations of compression ratio and complexity of the conventional merged DIMD.
[0014] An image decoding method according to an embodiment of the present invention for solving the above-described technical problem comprises the steps of: obtaining histogram information for at least one neighboring block among first candidate group blocks adjacent to a current block and second candidate group blocks located close to but not adjacent to a current block; generating a merged histogram for a current block by combining histogram information of the neighboring blocks; deriving at least one intra prediction mode from the merged histogram; and intra-prediction decoding the current block using the derived intra prediction mode.
[0015] According to an embodiment of the present invention for solving the above-described technical problem, a video encoding method comprises the steps of: obtaining histogram information for at least one neighboring block among first candidate group blocks adjacent to a current block and second candidate group blocks located close to but not adjacent to a current block; generating a merged histogram for a current block by combining histogram information of the neighboring blocks; deriving at least one intra prediction mode from the merged histogram; performing intra prediction encoding on the current block using the derived intra prediction mode; and including in a bitstream information indicating that the current block is encoded using merged DIMD (Decoder-side Intra Mode Derivation).
[0016] According to an embodiment of the present invention for solving the above-described technical problem, a video decoding device may include a processor, a memory connected to the processor, a histogram information acquisition unit for acquiring histogram information for at least one surrounding block among first candidate group blocks adjacent to a current block and second candidate group blocks located close to but not adjacent to the current block, a merged histogram generation unit for generating a merged histogram for a current block by combining histogram information of the surrounding blocks, a prediction mode derivation unit for deriving at least one intra prediction mode from the merged histogram, and an intra prediction decoding unit for intra prediction-decoding the current block using the derived intra prediction mode.
[0017] According to an embodiment of the present invention for solving the above-described technical problem, a video encoding device may include a processor, a memory connected to the processor, a histogram acquisition unit for acquiring histogram information for at least one neighboring block among first candidate group blocks adjacent to a current block and second candidate group blocks located close to but not adjacent to the current block, a merged histogram generation unit for generating a merged histogram for a current block by combining histogram information of the neighboring blocks, a prediction mode derivation unit for deriving at least one intra prediction mode from the merged histogram, an intra prediction encoding unit for intra prediction encoding the current block using the derived intra prediction mode, and a bitstream generation unit for including information indicating that the current block is encoded using merged DIMD (Decoder-side Intra Mode Derivation) in a bitstream.
[0018] According to the present invention, at least one effect of improving encoding efficiency, improving decoding efficiency, improving video quality, reducing computational amount, reducing software size, reducing hardware size, and improving other performances related to encoding and decoding can be derived in video encoding and decoding.
[0019] According to the present invention, a desirable compromise between compression ratio and complexity can be achieved when executing a merged DIMD.
[0020] Figure 1 is a conceptual diagram of a video communication system according to one embodiment of the present invention;
[0021] FIG. 2 is a conceptual diagram of the arrangement of an encoder and decoder in a real-time video streaming environment according to one embodiment of the present invention.
[0022] Figure 3 is a functional unit conceptual diagram of a video decoder according to one embodiment of the present invention;
[0023] Figure 5 is a conceptual diagram of a frame type according to one embodiment of the present invention;
[0024] Figure 6 is a conceptual diagram showing the structure of a video encoder according to the H.266 / VVC standard.
[0025] Figure 7 is a conceptual diagram for selection of a set of adjacent pixels according to one embodiment of the present invention;
[0026] Figure 8 is a conceptual diagram of gradient calculation for each pixel in a template according to one embodiment of the present invention.
[0027] Figure 9 is a conceptual diagram showing an example in which limited slope analysis is performed according to one embodiment of the present invention.
[0028] Figure 10 is a conceptual diagram showing a prediction fusion algorithm of DIMD according to one embodiment of the present invention.
[0029] Figure 11 is a conceptual diagram showing a prediction fusion algorithm of DIMD according to another embodiment of the present invention.
[0030] Figure 12 is a conceptual diagram of a merged DIMD according to one embodiment of the present invention, and
[0031] FIG. 13 is an exemplary diagram of a DIMD merge candidate block according to one embodiment of the present invention.
[0032] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention.
[0033] Although terms such as “first,” “second,” etc. may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, the first component could be referred to as the second component, and similarly, the second component could also be referred to as the first component. The term “and / or” includes any combination of multiple related listed items or any of multiple related listed items, and is non-exclusive unless otherwise indicated. The listing of items in this specification is merely an exemplary description to easily explain the spirit and possible implementation methods of the invention herein, and therefore is not intended to limit the scope of embodiments of the present invention.
[0034] As used herein, "A or B" can mean "only A," "only B," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B and C."
[0035] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0036] In this specification, "at least one of A and B" may mean "only A", "only B" or "both A and B". Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted identically to "at least one of A and B".
[0037] Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”
[0038] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0039] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this specification, it should be understood that the terms "comprises" or "has" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0040] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0041] In describing the invention herein, embodiments may be described or illustrated in terms of unit blocks that perform the described function or functions. The blocks may be expressed herein as one or more devices, units, modules, parts, etc. The blocks may be implemented in hardware by one or more logic gates, integrated circuits, processors, controllers, memories, electronic components, or information processing hardware implementation methods, but not limited thereto. Alternatively, the blocks may be implemented in software by application software, operating system software, firmware, or information processing software implementation methods, but not limited thereto. A single block may be implemented by being separated into multiple blocks that perform the same function, or conversely, a single block may be implemented to perform the functions of multiple blocks simultaneously. The blocks may also be implemented by being physically separated or combined according to any criteria. The blocks may be implemented to operate in an environment where their physical locations are not specified and are separated from each other by a communication network, the Internet, a cloud service, or a communication method, but not limited thereto. All of the above implementation methods are within the scope of various embodiments that can be taken by a person skilled in the field of information and communication technology to implement the same technical idea, and therefore, any detailed implementation method should be interpreted as being included within the scope of the technical idea of the invention in this specification.
[0042] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the attached drawings. To facilitate a comprehensive understanding of the present invention, the same reference numerals will be used for identical components in the drawings, and redundant descriptions of identical components will be omitted. Furthermore, the multiple embodiments are not mutually exclusive, and it is assumed that some embodiments may be combined with one or more other embodiments to form new embodiments.
[0043]
[0044] digital video codec
[0045] Figure 1 is a conceptual diagram of a video communication system according to one embodiment of the present invention. The video communication system (100) may be configured to include at least two terminals (110, 120) connected to each other via a network (105).
[0046] In one embodiment of the present invention, the above-described FIG. 1 may refer to a block diagram for configuring a unidirectional video communication network. A first terminal (110) among the terminals may encode video data in order to transmit (111) the video data via a network (105). A second terminal (120) among the terminals may be configured to receive (121) the encoded video data via a network and decode and display the same.
[0047] In another embodiment of the present invention, the above-described FIG. 1 may refer to a block diagram for configuring a two-way video communication network. For the two-way video communication, each terminal (110, 120) may be configured to encode video data acquired by itself for video transmission (112, 122) to each other terminal via the network. Each terminal may also be configured to receive (113, 123) video data transmitted by another terminal via the network, decode the same, and display the decoded video data.
[0048] The terminals (110, 120) shown in Fig. 1 may be exemplified as devices such as server computers, personal computers, portable computers, and smartphones, depending on the embodiment, but are not limited thereto, and may be any computing device generally used. For example, according to an embodiment of the present invention, each terminal (110, 120) may mean a desktop computer, a laptop computer, a tablet PC, a mobile phone, a smart phone, a personal digital assistant (PDA), a workstation, an electronic calculator, a server computer, a cloud computer, a virtual computer, a quantum computer, or any other electronic, electrical, or quantum computing device implemented in a movable or non-movable form, and in particular, it may be interpreted as any device that is designed to operate as a terminal device according to an embodiment of the present invention among such devices, is capable of operating as a terminal device according to an embodiment of the present invention, or is capable of installing and / or executing a computer program that enables the terminal device to operate as an embodiment of the present invention or perform a method corresponding to such an operation.
[0049] Each of the above terminals (110, 120) may be implemented by a plurality of functional units that are interconnected in various forms, such as a bus, a circuit, or a relationship between a routine and a subroutine, and configured to exchange information within each of the above terminals (110, 120). In addition, through the interconnection, the terminals may be configured to include a processor (130) having a calculation function and a memory (140) connected to the processor for the purpose of executing or supporting the operation of a functional unit that primarily requires calculation among the above functional units.
[0050] The processor (130) described in this specification may mean one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding.
[0051] Even if the processor (130) is expressed singularly for the sake of ease of understanding, a person of ordinary skill in the art will recognize that the processor (130) may include multiple processing elements and / or multiple types of processing elements. For example, a device according to one embodiment of the present invention may include multiple processors or one processor and one controller as the processor (130). In addition, the processor (130) may be implemented using various processing configurations, such as a parallel processor or a multi-core processor.
[0052] The processor (130) may be configured to execute an operating system (OS) and one or more software programs running on the OS. In addition, the processor may access, store, manipulate, process, and generate data in response to the execution of the software. The software program may include a computer program, code, instructions, or a combination of one or more of these, and may configure a processing device to operate as desired or independently or collectively command a processing device. The software may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave to be interpreted by the processor (130) or to provide instructions or data to the processor. The software may be distributed among multiple computer systems connected to the network (105), such as the terminals (110, 120), and stored or executed in a distributed manner.
[0053] The software may also be implemented in the form of program commands that can be executed through various computer means and recorded or stored in the memory (140). The memory (140) may be a computer-readable recording medium, and program commands, data files, data structures, etc. may be recorded singly or in combination in the computer-readable recording medium. The program commands stored in the memory (140) may be based on a command system specifically designed and configured for the embodiment of the present invention, or may follow a command system known and available to those skilled in the art of computer software, for example, a command system exemplified by the assembly language, C, C++, Java, Python, etc. It should be understood that the command system and the program commands therefrom include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by the device and / or the processor (130) according to an embodiment of the present invention using an interpreter, etc.
[0054] The computer-readable recording medium constituting the device according to one embodiment of the present invention, including the memory (140) described in this specification, may include a temporary or volatile recording medium that is maintained only while the processor is operating, such as a processor cache, a RAM, a flash memory, or a relatively non-volatile or long-term recording medium, such as a magnetic media such as a hard disk, a floppy disk, and a magnetic tape, an optical media such as a CD-ROM, a DVD, a magneto-optical media such as a floptical disk, or a solid state memory, or may include a read-only recording medium, such as a ROM arranged on hardware, and furthermore, the hardware itself configured to perform an operation equivalent to a series of program commands by a hard-wired structure by circuit wiring, and each step for performing the operation for implementing the embodiment of the present invention can be viewed as being recorded by the connection and arrangement of the hardware components, so that the connection and arrangement method It is obvious to those skilled in the art that this can be seen as equivalent to the above memory (140).
[0055] The embodiments described above with respect to the processor (130) and the memory (140) are not mutually exclusive, and may be selected or combined and implemented as needed. For example, one hardware device may be configured to operate as a module composed of one or more of the software to perform the operations of an embodiment of the present invention, and vice versa. As another example, in the present specification, all or part of the operations assigned to a certain functional unit may be implemented by one or more of the software stored in a device according to an embodiment of the present invention (preferably, in any one of the recording media belonging to the category of the memory) and configured to be executed by the processor, and in this case, such a functional unit may be referred to as a functional unit "included" in the processor.
[0056] The present invention is applicable to all environments for forming a one-way or two-way video communication network, and it should be understood that the network (105) can be formed by any means for transporting encoded video data between the terminals (110, 120).
[0057] In one embodiment of the present invention, the network (105) may refer to a wired or wireless communication network. In this case, depending on the embodiment, the network may be configured to communicate information using any communication standard, and the communication standard may include packet-based communication. The packet communication may be understood to mean including packets known as TCP or UDP, for example. The wired communication method of the network (105) may be a method of connecting to an external communication network by means of a telephone line of the RJ-11 standard, an Ethernet cable belonging to various categories of the RJ-45 standard, other coaxial cables, metal cables, optical cables, and various other wired media. The wireless communication method of the above network (105) may include, depending on the embodiment, a short-range wireless communication method including Bluetooth, Wi-Fi, Zigbee, and NFC (near field communication), or may include a long-range wireless communication method that may be referred to as a name of a wireless communication technology collectively called a generation name of a communication standard such as Wibro, WiMax, Global Systems for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long-Term Evolution (LTE), New Radio (NR), and other 2G, 3G, 4G, 5G, and 6G, and a name of an international standard wireless communication standard such as IMT-2000, IMT-Adcanced, IMT-2020, and IMT-2030. Of course, even if any conventional or newly developed wired or wireless communication means, method, standard, and protocol are applied to the implementation of the network (105), there is no problem in achieving the purpose of the present invention as long as it is a means configured to perform transmission and reception in an information and communication device such as the terminal (110, 120).It is also obvious that one network (105) can be configured in a mixed manner by one or more wired and / or wireless standards.
[0058] However, in another embodiment of the present invention, the network (105) may be understood to include a process of transmitting information using a computer-readable recording medium. In this case, the configuration of the network is not limited to a communication medium, and should be understood to include a process of temporarily storing and physically transporting information in a computer-readable memory and / or recording medium. The computer-readable recording medium used for transmitting information may be understood to mean a relatively non-volatile or long-term recording medium such as a magnetic media such as a hard disk, a floppy disk, and a magnetic tape, an optical media such as a CD-ROM or DVD, a magneto-optical media such as a floptical disk, or a solid state memory, which is mainly used as a means of transporting data between computing devices.
[0059] Any other means of information communication or transport, regardless of the method employed, can be considered within the scope of the present invention as long as it has a structure that supports the transmission and decoding of video data in an encoded state. Therefore, in addition to the examples listed above, any means of information communication or transport, whether known in the past or newly available, can fall within the scope of application of the present invention.
[0060] FIG. 2 is a conceptual diagram illustrating the arrangement of an encoder and decoder in a real-time video streaming environment according to one embodiment of the present invention. The streaming system (200) illustrated in FIG. 2 can be applied to video data communication networks, including, for example, digital broadcasting, video telephony, and video conferencing. However, it should be noted that technical structures identical or similar to the streaming system can be equally applied even when information is transmitted via a recording medium, as described above.
[0061] According to one embodiment of the present invention, the streaming system may include a video source (210) that generates a video stream. The video source may include a digital video acquisition means (212), which may be, for example, a digital camera or other device, for acquiring uncompressed raw video. The raw video stream (215) may have a large capacity and may therefore be compressed by a video encoder (217) coupled or connected to the video source.
[0062] The above encoder (217) may be configured as a means including hardware, software, or a combination of the two, configured to implement an image encoding method and / or an implementation method thereof according to one embodiment of the present invention.
[0063] Through the encoder (217), an encoded bitstream (219) having a reduced capacity compared to the original video stream can be output. The bitstream (219) can be provided in real time via a relay device, which may be referred to as a streaming server (220), and / or can be stored in a recording medium (225) of the streaming server (220) for subsequent use.
[0064] The streaming system (200) may include at least one streaming client (230, 240) that connects to the streaming server (220) to receive the encoded bit string (229) in real time or obtain it later. The streaming client may include a video decoder (232) that obtains the encoded bit string (229) (which may also be considered as a copy of the bit string (219) received by the streaming server), decodes the bit string (229), and outputs the resulting video data as video data in a form that can be displayed by a display (235) or other visual, auditory, or other sensory display means.
[0065] As described above, the functions for encoding and decoding video data are collectively called a coder-and-decoder, or video codec.
[0066] FIG. 3 is a conceptual diagram of a functional unit of a video decoder according to an embodiment of the present invention. As shown in FIG. 3, a receiving unit (310) can receive at least one encoded video data to be decoded by a decoder (305). In an embodiment of the present invention, the encoded video data may be independent for each reception, and the decoding procedure of each independent video data may be independent from the decoding procedure of other video data. The encoded video data may be received by the receiving unit (310) through a hardware or software connection (315) to a device storing the same, and as described above, the storing device may be a type of streaming server located at the other end of a communication network, or may mean a physical recording medium, but is not limited thereto.
[0067] The above-described receiving unit (310) can receive the encoded video data together with other data accompanying it, such as encoded audio data or other auxiliary data, and each of the data can be separated from the video data and provided to an appropriate processing function unit (312) other than the video decoder.
[0068] When the video data is provided through a communication network, a buffer memory (320) may be coupled between the receiving unit (310) and the decoder (305) to minimize delay and disconnection according to the network environment. The buffer memory (320) may refer to a computer-readable recording medium that temporarily stores the received video data and stably supplies it to a parser (330) corresponding to the input terminal of the decoder (305). However, if the bandwidth of the communication network is sufficient, if video data is read from a recording medium in a local location that is not physically separated, or if the possibility of communication delay is not predicted in other environments, the buffer memory may be unnecessary.
[0069] The video decoder (305) may include the parser (330) as its input terminal to interpret the encoded video data. The parser may perform a function of separating (parsing) a plurality of pieces of information stored in the form of a bit string in the encoded video data according to a predetermined rule, and, if necessary, performing an entropy decoding (335) of entropy-coded video data, thereby performing a function of reconstructing symbols (338), which are paragraphs of video encoding information. The symbols (338) may include all information for controlling the operation of the decoder (305), and / or may further include information for controlling a device that is attached to and operable with the decoder (305), such as a display device. Control information for controlling the above display device may include information in a format called supplementary enhancement information (SEI) or video usability information (VUI).
[0070] As described above, the parser (330) may be configured to perform entropy decoding (335) of the encoded video data. The entropy encoding method of the encoded video data may vary depending on the encoding standard, and decoding may be performed accordingly. Representative examples of the entropy encoding standard may include variable length coding, Huffman coding, and arithmetic coding, and each of the encoding methods may be a context-adaptive or context-sensitive method depending on the standard, and may also be based on principles widely known to those skilled in the art.
[0071] The parser (330) may be configured to extract at least one partial image from the encoded video data. The definition of the partial image may vary depending on the encoding standard, and depending on the standard, one or more of the examples listed below may correspond simultaneously and overlappingly. The partial image may be defined in units such as, for example, a group of pictures (GOPs), pictures / frames, tiles, slices, macroblocks, blocks, subblocks, transform units (TUs), and prediction units (PUs).
[0072] The parser (330) may be configured to extract encoding information, such as transform coefficients, quantization parameters (QPs), and / or motion vectors, from the encoded video data. The parser (330) may be configured to perform entropy decoding (335) and parsing operations on the video data received from the buffer memory, and to selectively decode symbols (338) representing the encoding information. In addition, the parser (330) may be configured to selectively supply a specific symbol (338) to a specific decoding function unit within the decoder (305), such as an inverse quantization and inverse transform unit (340), an intra prediction unit (350), an inter prediction unit (355), or a loop filter unit (360). Control of such information supply can be determined by the information sequence contained in the encoded video, and may vary depending on the encoding standard, and is not limited within the scope of the embodiments of the present invention, and is not described in detail in this conceptual diagram.
[0073] The decoder (305) may be comprised of a number of conceptual functional units that receive and process the encoded information provided by the parser (330). It should be readily apparent that these conceptual functional units may be combined or further subdivided, depending on implementation needs. For example, they may be further separated for ease of implementation, or integrated into one for operational efficiency. In any case, each functional unit may be configured to perform close interaction with each other. However, despite the possibility of such integration or separation, the following description will be given as a combination of conceptual functional units to illustrate the video data decoding procedure applied as an embodiment of the present invention.
[0074] The decoder may include an inverse quantization and inverse transformation unit (340). The inverse quantization and inverse transformation unit (340) may be configured to receive encoding information including a method to be used for numerical transformation (transform), a block size, quantization coefficients for recovering quantized information, and distinction information of a quantization matrix that simplifies and represents the quantized coefficients from the parser (330), and may be configured to output block values (341) that can be input to an aggregator (370) as a result of processing the encoding information.
[0075] In one embodiment of the present invention, the output values of the inverse quantization and inverse transformation unit (340) may include intra-prediction encoded block values. The intra-predicted block values may refer to values that can be decoded without using prediction information from a previously decoded partial image, for example, a previous frame, but using prediction information within a partial image currently being decoded, for example, a current frame.
[0076] Prediction information within the current partial image may be provided by the intra prediction unit (350). According to an embodiment of the present invention, the intra prediction unit (350) generates a block value of the same form as the block being decoded as prediction information by using image information of a spatially adjacent area derived from a partial image currently being decoded and of which decoding has been partially completed. The partial image information may be provided (381) from a current image buffer, a so-called line buffer (380). The merging unit (370), according to an embodiment, may be configured to merge the prediction information (351) generated by the intra prediction unit (350) with the block values (341) provided by the inverse quantization and inverse transformation unit (340).
[0077] In another embodiment, the output values of the inverse quantization and inverse transformation unit (340) may include block values subjected to motion compensation as inter-prediction encoded block values, and in some cases, block values subjected to motion compensation. In this case, the inter-prediction unit (355) may extract and use sample information (386) used for motion-based prediction from a reference picture buffer (385). The information (356) derived by performing motion compensation on the sample information based on the symbols (338) included in the block values as the output values may be configured to be merged with the block values (341) provided by the inverse quantization and inverse transformation unit (340) by the merger unit (370). In this case, the block values (341) may be referred to as so-called differential or residual values.
[0078] The position information within the memory used by the inter prediction unit (355) to extract the sample information from the reference image may be determined by a motion vector provided to the inter prediction unit (355) which is composed of a combination of symbols (338) for representing, for example, X, Y, and other specific points of the reference image. The inter prediction unit (355) may also include a function for interpolating and using the sample values when a so-called 'subsampling' capable motion vector is provided, and may further include a function for predicting and reinforcing the value of the motion vector.
[0079] The output values (371) of the above merging unit (370) may be provided to the loop filter unit (360) and processed by various loop filtering methods. The loop filter unit (360) may be configured to receive not only the block unit output (371) of the merging unit (370) but also the symbol (338) provided from the parser (330) to control its operation. The output of the loop filter unit (360) may be output to an external display means such as the display device through an output connection (390), but may be stored (361) in a line buffer (380) for use in prediction for interpreting a subsequent intra or inter coded block value, and may also be stored in a reference image buffer (385) through this.
[0080] Certain partial images, such as frames, after their decoding is completed, can be utilized as reference images for performing predictive decoding in a subsequent decoding process. One partial image, such as a frame, can be gradually accumulated in a line buffer (380) and decoded, and when one frame is decoded, the contents of the line buffer (380) are transferred (383) to the reference image buffer (385), and a new line buffer (380) can be allocated for decoding the new frame.
[0081] The above video decoder (305) may be configured to perform a decoding operation according to a predetermined video compression technique that may be documented by various international standards or commercial standards. The standards may include, for example, international standard recommendations such as H.264, H.265, and H.266 defined by the International Telecommunication Union Standardization Sub-Division (ITU-T). Those skilled in the art will understand that each of the above recommendations is equivalent to an international standard jointly defined by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC). The encoded video data may comply with a specific bitstream syntax defined by the relevant standards, as defined and required by the video compression standard documents and standard documents, and specifically by the profiles and levels specified within such documents. In addition, the complexity of the encoded video data may be limited to a certain level to comply with the profiles and levels. For example, a profile or level may be configured to limit a maximum picture size, a maximum decoding speed, and a maximum reference picture size. These limitations may, in some embodiments, also be further constrained via metadata signals for a hypothetical reference decoder (HRD) and HRD buffer management included in the encoded video data.
[0082] According to one embodiment of the present invention, the receiver (310) may receive additional redundant data along with the encoded video. The additional data may be considered as part of the encoded video data. The additional data may include information that may be used by the decoder (305) to properly decode the data or to more accurately reconstruct an image that approximates the original image. The additional data may be provided in the form of, for example, layers for temporal, spatial, or signal-to-noise ratio (SNR) enhancement, redundant slices, redundant images, and forward error correction codes.
[0083] FIG. 4 is a functional unit conceptual diagram of a video encoder according to one embodiment of the present invention. The encoder (405) may be configured to receive original video information (402) from a video source (401) and perform encoding.
[0084] The original video information (402) may have any suitable bit depth, for example, 8 bits, 10 bits, 12 bits, etc. In addition, the original video information (402) may have any suitable color space, for example, R / G / B, Y / U / V, Y / Cb / Cr, etc. In addition, the original video source may have any suitable sampling structure corresponding to the color space, for example, Y / Cb / Cr 4:2:0, Y / Cb / Cr 4:4:4, etc. The original video source having such a predetermined format may be provided to the encoder in the form of a digital video stream.
[0085] In a one-way video communication network, the original video information (402) can be obtained from a recording medium storing a previously prepared video source. In a two-way video communication network, the original video information (402) can be obtained from an image acquisition device, such as a camera, that generates at least one video transmission stream included in the two-way video communication.
[0086] The video data including the above original video information (402) may be configured as a plurality of partial images configured to simulate motion by being played back in time sequence. The partial images may be expressed as concepts such as pictures or frames, for example. The partial images may include one or more samples depending on the type of sampling structure, color space, etc. being used. Those skilled in the art will understand that the terms "samples" and "pixels" in digital images are closely related. The operation of the encoder will be described below with reference to such samples.
[0087] According to one embodiment of the present invention, the encoder (405) may be configured to encode and compress (partial) images constituting the original video information (402) in real time (or according to other temporal requirements required according to the implementation method) into the form of encoded video information.
[0088] In the encoder (405), the control unit (450) may be a functional unit configured to control an appropriate encoding speed. The control unit (450) may be configured to control other functional units and be functionally coupled to the following functional units as described below. The parameters set by the control unit (450) may include parameters related to bitrate control, such as skip of the image, quantizer, variable values for applying a picture quality optimization technique, and may also include values such as the size of the image, the structure of a group of pictures (GOP), and the maximum search range of a motion vector. A person skilled in the art will be able to understand various other functions that the control unit (450) may have, and such other functions may be added or removed according to the design of a video encoder optimized for an individual system design.
[0089] According to an embodiment of the present invention, the encoder (405) may be configured to operate in a structure such as a "coding loop" well known to those skilled in the art. To simplify the description by way of example, the coding loop may be configured with an internal encoder (so-called "source coder") (410) responsible for receiving an image to be encoded and generating symbols based on at least one reference image that has been encoded in the past, and a local decoder (420) configured to be connected to the internal encoder. The local decoder (420) may be configured to perform an operation to reproduce sample data to be generated by a decoder (490) located at an actual remote location that will receive encoded video information from the encoder (405) by receiving an output of the internal encoder (410).
[0090] Video data composed of sample data reconstructed by the internal decoder (420) may be configured to be input to the reference picture buffer of the encoder (405). As described above, the internal decoder (420) is implemented to reproduce the result output by the encoder (405) and to be decoded by a remote decoder, so the video data recorded in the reference picture buffer may also be identical in bit units to the information of the reference picture buffer of the remote decoder. That is, the prediction function unit that may be included in the encoder (405) may read the same values as the sample values of the previous frame that the decoder will later refer to in the decoding process from the reference picture buffer of the encoder (405).
[0091] As described above, the principle of achieving matching of the reference image buffer between the encoder (405) and the decoder (490) by means of the internal decoder (420) on the encoder (405) side is well known to those skilled in the art, and a method of responding to an environment in which such an environment is not guaranteed (e.g., information loss due to communication failure, etc.) can also follow what is known to those skilled in the art.
[0092] An embodiment of the operation method of the internal decoder (420) has been described in detail above with reference to FIG. 3. The decoder of FIG. 3 may be regarded as the aforementioned "remote" decoder (490). The internal decoder (420) may be implemented excluding lossless encoding and decoding sections such as the parser (330) or entropy decoding (335). This is because the internal encoder (405) is implemented to simply reproduce the operation of a decoder located at a remote location, and thus may directly decode symbols without requiring a process of compressing and then decompressing symbols. Accordingly, the functional units preceding the parser and entropy decoder as shown in FIG. 3 may not be provided or may be implemented at least partially.
[0093] As described above, according to a preferred embodiment of the present invention, any decoder function (excluding a parser and an entropy decoder) present in the decoder can naturally exist as a substantially identical function in the corresponding encoder (405).
[0094] The operation of the encoding function unit that may be included in the above encoder (405) can be considered as the inverse of the decoder function unit. Therefore, the embodiment can be explained by performing the operation of the decoder function unit in reverse. For example, a quantization and transform function unit corresponding to the inverse quantization and inverse transform unit may be provided, and an inter prediction encoding unit corresponding to the inter prediction unit may be provided. In addition, some additional explanations will be added.
[0095] The internal encoder (410) may be configured to perform encoding on input image information, for example, an input frame, by a predictive encoding method executed by a predictive encoding unit (440) that operates by referencing at least one temporally previous encoded partial image, for example, frames, from a reference image buffer (430) from at least one reference image information, for example, video data designated as a reference frame. In this case, the encoder (405) may be configured to encode a differential between blocks of samples constituting the input image and blocks of samples constituting the reference image.
[0096] The internal decoder (420) can decode video data that can be designated as the reference picture from symbols generated by the internal encoder (410). As described above, since the video data is subject to the same decoding operation as that performed by a remote decoder, the video data used as the reference picture may be provided to the encoder (405) in a form that has undergone lossy compression and has suffered some damage, and this operation may be intended to ensure operational consistency with the decoder.
[0097] The prediction encoding unit (440) may be configured to perform a prediction search operation within the encoder (405). The prediction search operation may refer to an operation corresponding to the inter prediction or intra prediction described in the description of the decoder. For image information that is input and scheduled to be newly encoded, the prediction unit may access the reference image buffer (430) to retrieve information such as a motion vector, a block shape, and metadata that may include the same, which are information indicating points of a reference image that can function as prediction reference information suitable for the new image information, and a sample block to be actually referenced. The prediction encoding unit (440) may operate on the basis of a so-called "sample block by pixel block" to obtain appropriate prediction reference information. According to one embodiment of the present invention, at least one prediction reference information may be designated for the input image, which designates at least one reference image information stored in the reference image buffer (430), as determined based on the search results obtained by the prediction encoding unit (440).
[0098] In one embodiment of the present invention, the control unit (450) may be configured to manage the overall encoding operation of the internal encoder (410), including setting parameters used to encode video data.
[0099] All outputs of the above-described functional units may be subjected to entropy encoding (460) in order to be finally output. The entropy encoding (460) may include various entropy coding techniques, such as variable length coding, Huffman coding, and arithmetic coding, for the symbols generated by the various functional units as described above, and each encoding method may be a context-adaptive or context-sensitive method according to the standard, or may be based on principles widely known to those skilled in the art. Such entropy encoding (460) can typically achieve lossless compression, and thus can be configured to convert at least one symbol generated by the functional units into encoded video data.
[0100] The above control unit (450) may, when controlling the operation of the encoder (405), apply the type of encoding of a specific partial image during the encoding period to each partial image, such as a picture or frame. Depending on the type, the method by which the partial image is encoded may be affected. Depending on the embodiment, the type may include what is categorized as the following "frame type."
[0101] Fig. 5 is a conceptual diagram of a frame type according to one embodiment of the present invention. The following description will be made with reference to Fig. 5.
[0102] An intra (“I”) picture (510) may refer to a picture that can be encoded and decoded using only its own information without referring to other picture information in the video data through predictive encoding. The “I” picture may be designated by names such as a key frame, an independent / instantaneous decoder referh (IDR) frame, and a clean random-access (CRA) frame, depending on the video encoding standard, and the “I” pictures designated by the various names as described above may have various modifications and application methods as permitted by each standard and may be partially different from each other. In addition to those listed above, various application methods for implementing the “I” picture may be by various methods that are already known to those skilled in the art or may be newly provided.
[0103] A prediction ("P") picture (520) may refer to a picture that can be encoded and decoded through intra or inter prediction based on at least one prediction information and / or a motion vector that designates at least one reference picture to predict sample values of a block constituting the picture. The "P" picture may be configured to refer to only one reference frame, or may be configured to refer to one or more reference frames, according to a video encoding standard. When referring to more than one reference frame, sample information and / or associated metadata derived from multiple reference pictures may be used to reconstruct a single block. However, in common cases, a picture designated as a "P" picture may be understood as a picture that performs reference only to a temporally preceding picture.
[0104] A bidirectional prediction ("B") picture (530) may refer to a picture that can be encoded and decoded through intra or inter prediction based on at least one prediction information and / or motion vector that designates at least two reference pictures to predict sample values of blocks constituting the picture. In a common case, a picture designated as the "B" picture is distinct from a picture designated as the "P" picture, and may be understood as a picture that performs a reference without being limited to a temporally preceding picture.
[0105] Video data may be spatially divided into a plurality of sample blocks during the encoding and decoding process, and encoding may be performed in units of the blocks. The block units may include, but are not limited to, sizes such as 4x4, 8x8, 4x8, or 16x16 in units of horizontal / vertical pixels, as is widely known. The block may be encoded using a predictive encoding method with reference to any other (already encoded) blocks, as permitted and / or restricted by the type specified for each partial picture including the block. For example, the blocks of the "I" picture (510) may be encoded without using a predictive encoding method, or with reference to blocks already encoded within the same partial picture. That is, only the so-called intra prediction method may be used. In contrast, the "P" picture (520) may further reference a reference picture encoded in at least one previous time unit, and thus, inter prediction may also be used for encoding along with intra prediction. In the case of a "B" picture (530), reference can be made not only to a previously encoded picture in the encoding order but also to a later reference picture in terms of time unit. However, it is widely known that there may be blocks within a "P" picture or a "B" picture that are encoded without relying on predictive encoding.
[0106] The above video encoder (405) may be configured to perform encoding operations according to a predetermined video compression technique that may be documented by various international standards or commercial standards. Examples of the above standards may include all of those described in the above decoder.
[0107] According to one embodiment of the present invention, the transmitter (470) may buffer the encoded video data generated by the entropy encoding in order to provide / transmit the video data (ultimately to a remote decoder (490)) to a device storing the encoded video data via a hardware or software connection (495). According to an embodiment, when providing / transmitting the encoded video data from the video encoder (405), the transmitter (470) may receive and merge other data accompanying the encoded video data, for example, encoded audio data or other auxiliary data, from a separate source (480).
[0108] According to one embodiment of the present invention, the transmitter (470) may be configured to transmit additional data along with the encoded video. The additional data may be considered part of the encoded video data. The additional data may include information that can be used by a decoder to properly decode the data or to more accurately reconstruct an image that approximates the original image. Examples of the additional data may include all of the examples previously presented with respect to the receiver (310) of the decoder.
[0109] The present invention can be implemented by a digital video compression standard that is widely used and understood by those skilled in the art as described above. The digital video compression standard may include at least one of compression standards known by the standard name such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.
[0110] Fig. 6 is a conceptual diagram illustrating the structure of a video encoder according to the H.266 / VVC standard. What is depicted in Fig. 6 corresponds to the rough structure of a video encoder widely known by standard codes such as ITU-T H.266 and ISO / IEC 23090-3, and also by the name MPEG-I Part 3 or the general name versatile video coding (VVC).
[0111] According to FIG. 6, a video encoder (605) may be configured to receive uncompressed and unencoded original video data (601) as input and output an encoded bit stream (602). The video data (601) may be supplied directly to a luma mapping unit (610a) when intra-encoded, or may be supplied to a luma mapping unit (610b) via an inter-prediction unit (620) including motion vector extraction. In the intra-encoded case, the mapped luma signal may be supplied to an output merger (606) by selecting (608) at least one of an intra-prediction encoded signal via an intra-prediction unit (625) or an inter-prediction encoded signal output from the luma mapping unit (610b) via the inter-prediction unit (620). The result of the above output merger can be applied to a chroma scaling unit (615). (The operation of the luminance signal mapping unit (610) and the operation of the chroma scaling unit (615) are collectively referred to as a luma mapping / chroma scalaing (LMCS) process.) The reduced chroma signal can be provided to a transform unit (630), and the transform unit (630) can perform an adaptive color transform, particularly on the chroma signal. The coefficients derived as a result of the transform are applied to a quantization unit (640) and quantized. As a result, lossy compression is achieved, and the result of the lossy compression can be output as a bit string (602) through a multi-hypothesis CABAC (650), which is a lossless compression method.
[0112] Meanwhile, the result of the lossy compression may actually enter the decoding process by going through the processes of inverse quantization (645), inverse transform (635), and luminance signal expansion (617) to generate a coding loop. The result of the luminance signal expansion may be supplied to the internal merger (607) together with the result of selecting (608) at least one of the previously generated intra prediction encoding signal or inter prediction encoding signal. The result of the internal merger may go through inverse luma mapping (618), and then may go through processing such as a deblocking filter (660), sample adaptive offset (SAO) (670), and an adaptive loop filter (ALF) to reproduce the image quality improvement process in the decoder. The result of reproducing the operation in the decoder as described above is applied to the reference image buffer (690) and can be reused for prediction encoding by the inter prediction unit (620).
[0113] The present invention can also be used by or in combination with the enhanced compression model (ECM), which is an implementation of a next-generation video codec being developed by the joint video experts team (JVET), an international standardization expert group, to improve H.266 / VVC. According to the standardization progress document of the above JVET, document number ISO / IEC JTC 1 / SC 29 / WG 5 N 190 (also document number JVET AC2025-v1), the improved compression model may include an improved intra prediction coding method, an improved inter prediction coding method, an improved transform and transform coefficient coding method, an improved adaptive loop filtering method, a bilateral filtering method, a new sample adaptive offset (SAO) method for image quality improvement, an extended entropy coding method, and an improved gradual decoding refresh (GDR) technique, and such techniques may be included in the present invention as detailed techniques for implementing the present invention.
[0114]
[0115] General method of applying decoder-side intra prediction mode estimation (DIMD)
[0116] The present invention provides a method for applying an improved decoder-side intra mode derivation (DIMD) technique that can be used in the video encoding and decoding field including the above-described embodiments, and a device to which such a method is applied. As described above, DIMD proposes a method for calculating the intra prediction direction of a current block from values of a previously decoded picture adjacent to a block currently being encoded / decoded (which, depending on the embodiment, can be understood as replacing it with a similar division unit such as a CTU as described above) in order to reduce the amount of information transmitted from an encoder to a decoder regarding an intra prediction mode in the process of increasing the compression rate of an image through intra prediction, i.e., intra prediction. Accordingly, even if the encoder does not provide a separate signal, a decoder in DIMD mode can autonomously derive directional information of intra prediction from a preceding decoding process, thereby resulting in an improvement in compression efficiency.
[0117] histogram of gradients
[0118] In order to estimate the intra prediction mode for a block, the DIMD method may be configured to use a histogram of gradients (HoG). The histogram may be understood as a graph representing the intensity of each prediction direction of a block allowed by the video encoding / decoding standard to which the DIMD method is applied. In general video encoding / decoding standards, multiple directional prediction modes are defined. For example, H.265 / HEVC defines 33 directional prediction modes, and H.266 / VVC defines 65 directional prediction modes. Since each directional mode represents a specific angle, the intensity may be defined corresponding to each directional mode. Those skilled in the art will understand that the histogram may be configured adaptively even when a different number or type of directional prediction mode is used in a conventional or newly provided video encoding / decoding standard. Furthermore, those skilled in the art will appreciate that, in generating the histogram, the specific number of directional prediction modes defined in the video encoding / decoding standard as described above may be arbitrarily further subdivided or partially omitted. For example, in generating a histogram, the 65 directional prediction modes of H.266 / VVC may be further subdivided to define more than 66 angles, or may be partially omitted to define less than 64 angles.
[0119] In order to generate the histogram, a set of adjacent pixels on which gradient analysis is to be performed may first be selected. The set of adjacent pixels may be derived from pixels that have already been decoded and reconstructed. Fig. 7 is a conceptual diagram for selecting a set of adjacent pixels according to an embodiment of the present invention. Referring to Fig. 7, in a frame (700) of an image being decoded, an area (750) that has already been decoded and reconstructed and an area (760) that has not yet been decoded may be distinguished based on a current block (710) that is currently being decoded. At this time, as illustrated in Fig. 7, a set of pixels surrounding the current block (710) to the left by TW pixels (717) and to the top by TH pixels (718) may be selected as a template (715) within the reconstructed area (750). According to an embodiment of the present invention, the TW and the TH may be set to 3. According to another embodiment of the present invention, the TH and the TW may be set to the same value or different values. In one embodiment of the present invention, either the TW or the TH may be undefined or defined as 0. In other words, the template (715) may be defined only in the upward direction by specifying only one or more THs, or may be defined only in the left direction by specifying only one or more TWs.
[0120] Next, a gradient analysis can be performed on pixels included in the template (715). Through this, the gradient direction of the image of the pixels within the template (715) can be determined, which is likely to be identical to the gradient direction of the image of the current block. Therefore, according to one embodiment of the present invention, horizontal and vertical Sobel filters defined as in mathematical expression 1 can be used to estimate the gradient direction of the image of the current block by measuring the gradient with respect to the template (715).
[0121]
[0122] In order to apply the Sobel filter matrices to each pixel of the template (715), a window composed of eight directly adjacent pixels centered on each pixel and surrounding at least one pixel of the template (715) may be formed to apply the filter. Through this, the Sobel filter matrix M of the mathematical expression 1 x By applying the horizontal slope value G x , M y By applying the vertical slope value G y can be obtained.
[0123] FIG. 8 is a conceptual diagram of gradient calculation for each pixel in a template according to an embodiment of the present invention. The example illustrated in FIG. 8 is for a template (810) having a width of 3 pixels (817) for the current block (710) illustrated in FIG. 7. As shown in FIG. 8 (a), pixels (820) corresponding to the center line of the template (810) are used as the target pixels for gradient analysis, and a window (830) of 3 pixels wide and 3 pixels high is formed for each reference pixel (835) to apply the Sobel filter. By applying the filter, the amplitude ("Ampl") and angle ("Angle") of the gradient can be calculated using the gradients Gx and Gy for each pixel (820) on the center line of the template (810) as in the following mathematical expression 2.
[0124]
[0125] Next, individual histogram values according to the direction of the gradient can be obtained. The histogram values can be formed by the angle-specific index of the intra prediction mode allowed by the video encoder and decoder to which the DIMD method is applied, as described above. According to each angle value, the histogram value of the corresponding internal angle mode can be configured to increase by Ampl. When all the pixels (820) on the center line of the template (810) are processed, the histogram values can include the accumulated value of the gradient amplitude for each internal angle mode. Fig. 8 (b) shows an example of a histogram calculated after applying the above operation to all pixel positions of the template. Referring to Fig. 8 (b), it can be seen that the histogram values can be expressed as the amplitude (Ampl) for each angle (Angle).
[0126] According to one embodiment of the present invention, other calculation methods than those described above may be used to generate the histogram values. For example, the angle (Angle) may be calculated based on an angle value of a direction prediction mode of a previously determined neighboring block instead of being calculated based on the direction of the gradient. As another example, the amplitude (Ampl) may be determined and / or adjusted based on a value that is proportionally or exponentially determined based on the size and / or number of samples included in the template or samples included in a neighboring block included in the template.
[0127] Intra prediction encoding methods that refer to pixels above the current block may cause complexity issues related to so-called line buffering in the decoder. Therefore, in one implementation, for a block located at the upper boundary of a single large encoding / decoding unit (e.g., a coding tree unit (CTU) as indicated in the VVC codec standard), gradient analysis may be configured not to be performed on pixels located at the upper part of the template. Fig. 9 is a conceptual diagram illustrating an example in which limited gradient analysis is performed according to one embodiment of the present invention.
[0128] According to one implementation method, the mode exhibiting the highest peak in the histogram may be configured to be selected as the intra prediction mode for the current block. Furthermore, if the maximum value of the histogram is 0 (meaning that slope analysis cannot be performed or that the image of the area constituting the template is flat without slope), the mean (DC) mode may be selected as the intra prediction mode for the current block.
[0129] According to one implementation method, intra prediction for the current block can be performed using a mode that exhibits at least one upper peak in the histogram. For this purpose, a prediction fusion technique based on weighted summation or weighted average can be utilized.
[0130] According to one implementation method, intra prediction for the current block can be performed using prediction results from a mode showing at least one upper peak in the histogram and a planar mode. For this purpose, prediction fusion techniques based on weighted summation or weighted averaging can be utilized, as described above.
[0131] According to one implementation method, when the block size is less than a certain size, for example, less than a size of 4x4 pixels, the frequency of application of the Sobel filter can be made more rare. Fig. 9 is a conceptual diagram for application of a Sobel filter with a reduced frequency according to one embodiment of the present invention. Referring to Fig. 9 (a), for a small block (910), a window (931, 932) can be configured to be calculated by using one pixel (935) from the left and one pixel (936) from the top, respectively. In addition to reducing the number of operations for calculating the gradient, such an implementation method can also contribute to simplifying the selection of the two best modes from the histogram, as illustrated in Fig. 9 (b).
[0132] According to one implementation, the template can extend in the upper right direction and the lower left direction. If there are available pixels, in a block of W pixels in width and H pixels in height, the template can extend in the upper right direction by up to W pixels and in the lower left direction by up to H pixels. According to another implementation, the template can be excluded or disabled in a specific direction. For example, the template can be defined only in the upper direction or only in the left direction. This decision to disable the template can be made when the block satisfies an edge condition, or can be made based on another condition / flag / signal regardless.
[0133] According to one implementation method, information of other blocks spatially and / or temporally adjacent to the current block may be referenced to derive a histogram for the current block. For example, the method may be configured to derive a histogram to be applied to predictive encoding / decoding of the current block by applying intra-template matching (IntraTMP) based on information of the current block or its neighboring blocks, or by identifying at least one block vector, a motion vector, and other reference vectors from the current block or its neighboring blocks, selecting a specific block region spatially and / or temporally adjacent based on the template matching result and / or as indicated by the vector, and performing histogram derivation according to the present invention for the block region.
[0134] prediction fusion
[0135] A prediction fusion algorithm may refer to an algorithm that fuses at least one prediction value to derive a single prediction value. According to an embodiment of the present invention, at least one prediction value may be generated by a spike in the histogram, and the prediction fusion algorithm may be used to combine them.
[0136] Fig. 10 is a conceptual diagram showing a prediction fusion algorithm of DIMD according to one embodiment of the present invention. According to one embodiment of the present invention, three angles corresponding to the three highest spikes of the histogram can be detected as M1 (1011), M2 (1012), and M3 (1013). Next, by applying (1021, 1022, 1023) pixel predictions by the three angles (1011, 1012, 1013) to reference pixels (1020), respectively, intra prediction pixel information Pred1 (1031), Pred2 (1032), and Pred3 (1033) are obtained, and then pixel information obtained by fusing (1030) each of the intra prediction pixel information can be calculated as the final prediction value (1050) of the block. For the above fusion (1030), a weighted average of three predictor variables can be calculated (1041, 1042, 1043) with respect to the spike amplitudes of the histogram. In a similar manner, at least two or more different numbers of spikes can be selected from the histogram and merged together.
[0137] Fig. 11 is a conceptual diagram illustrating a prediction fusion algorithm of DIMD according to another embodiment of the present invention. According to one embodiment, two angles corresponding to the two highest spikes of the histogram can be detected as M1 (1111) and M2 (1112). Next, by applying pixel predictions (1121, 1122) based on the two angles (1111, 1112) to reference pixels (1120), Pred1 (1131) and Pred2 (1132) are obtained, and Pred3 (1123) can also be obtained by conventional planar prediction. Pixel information obtained by fusing (1130) each of the intra prediction pixel information can be calculated as the final prediction value of the block. For this purpose, weights (1141, 1142, 1143) for the fusion (1130) can be applied. The weighted prediction value (1143) for Pred3 (1133) can be fixed to 1 / 3, and may be set to 21 / 64 for the convenience of bit-based operation. Each of the weighted prediction values (1141, 1142) for Pred1 (1131) and Pred2 (1132) can be determined in proportion to the spike amplitude size in the histogram within a range where their total is 2 / 3 (or 43 / 64). According to one embodiment of the present invention, three, four, five, or more spikes can be selected from the histogram and merged with the prediction result (1143) of the planar prediction mode by a method similar to or combining those shown in FIGS. 11 and 12.
[0138] According to one embodiment of the present invention, in a block of W pixels wide and H pixels high, if the upper or left histogram size is twice as large as the others, the weights for the predictions derived from each spike may be modified. In this case, the weights may vary depending on the pixel location within the block from which the predictions are derived.
[0139] For example, if the histogram on the top is twice that on the left, the weights can be modified as in Equation 3 below.
[0140]
[0141] As another example, if the left histogram is twice as large as the top, the weights can be modified as in Equation 4 below.
[0142]
[0143] In the above mathematical equations 3 and 4, wDimd i can mean the unmodified weights, and Δi is a predefined arbitrary number, for example, 10.
[0144]
[0145] General method of applying merged DIMD
[0146] A technology for realizing a merged DIMD (DIMD Merge) that quotes the result of DIMD processing from a previously encoded / decoded block and reuses it in a subsequent block is also proposed.
[0147] FIG. 12 is a conceptual diagram of a merged DIMD according to an embodiment of the present invention. Referring to FIG. 12, the merged DIMD may be configured to execute a DIMD process for the current block (1210) using DIMD information (1225, 1235) extracted from neighboring blocks (1220, 1230) of the current block (1210). In particular, a new merged histogram (1215) for the current block (1210) may be calculated based on the histograms (1225, 1235) of the neighboring blocks (1220, 1230). For example, a new merged HoG (Merged Histogram of Gradients; MHoG) for the current block may be calculated based on the HoG of the neighboring block. Also, for example, a new Merged Histogram of Occurrences (MHoC) for the current block can be calculated based on the HoC of the neighboring block. According to one embodiment of the present invention, in order for the merged DIMD to be used, at least one block (1220, 1230) adjacent to the block (1210) currently being encoded / decoded must be encoded by DIMD or merged DIMD. If one adjacent block (1220, 1230) encoded by DIMD or merged DIMD is available, the histogram (1225, 1235) of the corresponding block can be directly quoted and used as the merged histogram (1215) for the current block. If two or more adjacent blocks (1220, 1230) are available, the respective histograms (1225, 1235) can be combined to derive the merged histogram (1215). According to one embodiment of the present invention, each histogram (1225, 1235) can be merged into the merged histogram (1215) by taking the average of its amplitude.
[0148]
[0149] Occurrence-based intracoding (OBIC)
[0150] Occurrence-based intra coding (OBIC) is a technology that operates similarly to the DIMD method, and in the present specification, it can be considered as an embodiment of implementing the DIMD method or an embodiment applied from the DIMD method. According to the OBIC, a Histogram of Occurences (HoC) can be used as the histogram, which is used in conjunction with, compatible with, and / or as a concept that replaces the HoG. The HoC can mean a histogram that collects the frequency of occurrence of intra prediction modes of samples of neighboring blocks in proportion to the sample size of the neighboring blocks. Most HoC and the HoG can be used in conjunction with, compatible with, and / or replace each other in that all or most of the intra prediction modes collected by the HoC include the directional prediction modes collected by the HoG, and for example, when generating the histogram, each amplitude value can be considered to be generated based on the frequency of occurrence that is proportional to the size of the sample. When a histogram including HoC is generated as described above, fusion of intra prediction results as illustrated in Fig. 10 or Fig. 11 based on this histogram can be performed by the same or similar method.
[0151] Accordingly, it is self-evident that all embodiments described in the present invention can be applied to both DIMD and OBIC, as long as they do not conflict with the technical spirit of the present invention. In particular, with respect to the description of the histogram described in the present invention, it will be understood that the disclosure of the present invention regarding its operation and manipulation can be equally applied when using a gradient histogram (HoG) and an occurrence histogram (HoC).
[0152]
[0153] Advanced application of merged DIMD
[0154] According to one embodiment of the present invention, a more sophisticated calculation may be used to calculate the merged histogram. In a method for generating a merged histogram referencing N surrounding blocks, the referenced blocks may be blocks encoded in at least one directional intra prediction mode. For example, assume any block i among the N referenced blocks. A histogram H, which may be a single histogram or a merged histogram, is calculated from the block i. i In , when the total number of directional intra prediction modes defined by the encoding / decoding standard is M, and a number m defined as 0 or more and less than M represents any intra prediction mode defined by the standard, H i (m) is the histogram H for the above prediction mode m i can be regarded as the amplitude (magnitude) in .
[0155] According to one embodiment of the present invention, the referenced block i may be a block that is spatially or temporally associated with the current block, for example, a block within a spatially non-adjacent but temporally identical restored region, or may include a spatially adjacent but temporally separated co-located block. According to an embodiment, the location of the block i may be determined based on a spatial offset from the current block, a block vector, a motion vector, a reference vector, or information derived from the current block or its neighboring blocks.
[0156] According to one embodiment of the present invention, if the block i is encoded with DIMD or merged DIMD, the histogram of the DIMD method or the merged histogram of the merged DIMD method for the block i is the histogram H. i can be considered as.
[0157] According to one embodiment of the present invention, when the block i is encoded in a directional intra prediction mode other than DIMD or merged DIMD, and the sign of the intra prediction mode used for the encoding is n which is greater than or equal to 0 and less than M, the histogram H for the block i i is H i (n) has the amplitude of the first value only for the other H i () can be set to have an amplitude of the second value. At this time, according to an embodiment, the first value may be a value determined in relation to at least one of the horizontal and vertical pixel sample sizes of the block i. In addition, according to an embodiment, the second value may be a value for initialization, and according to an embodiment, may be 0 or a value determined in relation to the first value. In addition, the second value may be a single value, or H i It may also mean a series of values that vary across (0, 1, ,…, M-1). For example, the second value may be composed of values that are linearly or non-linearly extrapolated in order of proximity to the directionality of the directional intra prediction mode indicated by the first value.
[0158] According to one embodiment of the present invention, when the block i is composite-encoded with one or more directional intra prediction modes (for example, it may mean that an encoding technique such as spatial geometry partitioning mode (SGPM) or template-based intra mode derivation (TIMD) is used), all of the one or more directional intra predictions are H ican be reflected in the calculation. For example, assuming that the signs of the intra prediction mode used in the composite encoding are n and o greater than or equal to 0 and less than M, the histogram H for the block i i is H i For (n), the first value, H i (o) has the amplitude of the second value, and other H i () may be set to have an amplitude of the third value. In this case, depending on the embodiment, the first value and the second value may be the same or different. Further, depending on the embodiment, the first value and the second value may be values determined in relation to at least one of the horizontal and vertical pixel sample sizes of the block i. Further, depending on the embodiment, the first value and the second value may be determined based on the method and / or structure of the composite encoding. For example, when the SGPM encoding technique is used, the first value and the second value may be determined in proportion to the area of the geometrically divided region in which the intra prediction modes n and o are used, respectively.
[0159] At least one H from the surrounding blocks as described above i Once the at least one H is determined or calculated, i A merged histogram H~ can be calculated from the merged histogram. According to one embodiment of the present invention, the merged histogram H~ can be calculated as shown in the following mathematical expression 5.
[0160]
[0161] In the above mathematical expression 5, N is a number greater than or equal to 1 and the histogram H i It can mean the total number of surrounding blocks i that are referenced. That is, the above mathematical expression 5 is the N histograms H that are referenced. iThe process of deriving a merged histogram H~ by taking the arithmetic mean of each individual amplitude included in can be represented. However, according to another embodiment of the present invention, the merged histogram H~ may be derived by any other averaging method (for example, it may include a geometric mean or a weighted mean) or any other algorithm that combines multiple matrices in addition to the method shown in the above mathematical expression 5.
[0162] The above-described derived merged histogram can be used in the merged DIMD mode like a normal histogram. According to one embodiment, from the merged histogram, intra prediction directions having at least one amplitude are selected in descending order, and one, two, three, four, five, or more are selected and used for intra prediction. According to one embodiment, the method of intra prediction may refer to the above-described prediction fusion method. According to one embodiment, in intra prediction using the above-described prediction fusion, the amplitude for each of the above-described directions may be used as a merging proportion of each intra prediction result to be merged.
[0163] According to one embodiment of the present invention, information about the use of the merged DIMD may be transmitted by being included in an encoded video bitstream. The information may be encoded as a signal of at least 1 bit, and / or encoded in a compressed form by an entropy encoding method of at least 1 bit. According to the entropy encoding method, the signal may ultimately occupy an information amount of less than 1 bit. According to one embodiment of the present invention, in order to encode the signal, the signal may be included in and / or added to context information input to context-adaptive arithmetic coding (CABAC) used in standards such as H.264, H.265, H.266, and similar specifications.
[0164] According to one embodiment of the present invention, in order to improve the encoding speed, information regarding the use of the merged DIMD may be signaled only for blocks larger than a certain size. Furthermore, according to one embodiment of the present invention, information regarding the use of the merged DIMD may be signaled only when at least one adjacent block (e.g., the adjacent block on the left or top) is encoded using DIMD or merged DIMD.
[0165] According to one embodiment of the present invention, in order to reduce complexity, in the prediction fusion used in the DIMD including the merged DIMD, an application technique such as position-dependent prediction combination (PDPC) in planar prediction may be used. In this case, the intra prediction block derived in the process of calculating the sum of absolute transformed difference (SATD) during the encoding process may be reused in the prediction fusion of the DIMD process.
[0166] According to one embodiment of the present invention, the blocks to which DIMD information is referenced for generating the merged DIMD need not necessarily be limited to adjacent blocks. According to one embodiment of the present invention, the referenced blocks may be composed of DIMD merge candidates.
[0167] FIG. 13 is an exemplary diagram of a DIMD merge candidate block according to one embodiment of the present invention.
[0168] According to one embodiment of the present invention, the DIMD merge candidate may include blocks of a first candidate group, and the first candidate group may refer to blocks adjacent to the current block (1310). According to an embodiment, the first candidate group may include any block belonging to and / or overlapping an area indicated by area 1321 of FIG. 13. According to an embodiment, the first candidate group may include at least one block located at the upper left, upper right, left, and lower left of the current block (1310). According to an embodiment, the first candidate group may be configured to include one block located at the upper left, two, three, four, five, six, or eight blocks located at the upper and upper right, and two, three, four, five, six, or eight blocks located at the left and lower left, based on the current block (1310). In some embodiments, the first candidate group may include at least one block spaced apart from the top of the current block (1310) by a vertical pixel length of the current block (1310). The first candidate group may include at least one block spaced apart from the left of the current block (1310) by a horizontal pixel length of the current block (1310).
[0169] According to one embodiment of the present invention, the DIMD merge candidate may include blocks of a second candidate group, and the second candidate group may refer to blocks that are not adjacent to but are located in close proximity to the current block (1310). According to an embodiment, the second candidate group may include any block that belongs to and / or overlaps the area indicated by area 1322 of FIG. 13. According to an embodiment, the second candidate group may include at least one block that is located in the upper left, upper right, left, and lower left directions of the current block (1310), but is spaced apart by a predetermined multiple in proportion to the horizontal and vertical lengths of the current block (1310).
[0170] According to one embodiment of the present invention, the blocks belonging to the second candidate group (1322) may include blocks indicated by a block vector, a motion vector, or a similar reference vector. The block vector represents a displacement from the current block to a specific block location within the current picture, and may be derived based on the current block or its neighboring blocks. The motion vector may point to a reference area of another picture that is expected to have similarity to the current block in a manner similar or the same as that used in inter prediction. In this way, blocks that are spatially distant from the current block but have structural or content similarity can be included in the second candidate group (1322) and used in the DIMD merging process.
[0171] According to one embodiment of the present invention, at least one block included in the first candidate group (1321) and the second candidate group (1322) may be excluded from the DIMD merge candidate if it has not yet been encoded according to the order of encoding processing.
[0172] According to one embodiment of the present invention, the DIMD merge candidate may include a merge histogram derived from a merged DIMD. In some embodiments, the merged histogram as the merge candidate may mean one determined by a method of forming a merged histogram that is the same as or different from the method of deriving the merged DIMD. In some embodiments, the merged histogram as the DIMD merge candidate may mean a merged histogram calculated from blocks included in the first candidate group (1321) and the second candidate group (1322).
[0173] According to one embodiment of the present invention, a redundancy check may be performed during the process of deriving the list of DIMD merge candidates. DIMD information (e.g., histogram information) newly included in the list of DIMD merge candidates may be selected so as not to overlap with information not already included in the list of DIMD merge candidates or obtained when DIMD is performed independently from the current block (1310).
[0174] According to one embodiment of the present invention, information on which DIMD information for a specific block to select and use from a list of DIMD merge candidates may be transmitted by being included in an encoded video bitstream. The information may be encoded as a signal of at least 1 bit, and / or encoded in a compressed form by an entropy encoding method of at least 1 bit. According to the entropy encoding method, the signal may ultimately occupy an information amount of less than 1 bit. According to one embodiment of the present invention, in order to encode the signal, the signal may be included in and / or added to context information input to context-adaptive arithmetic coding (CABAC) used in H.264, H.265, H.266, and similar standards / specifications.
[0175] According to one embodiment of the present invention, when the merged DIMD is available, the effect of each DIMD method derived from the merged DIMD candidate list can be evaluated and determined and used by the same procedure in the decoder and the encoder. Preferably, the evaluation can be evaluated by the Hadamard method, and / or the cost optimality can be evaluated through a rate-distortion optimization (RDO) loop. In one embodiment of the present invention, when the encoding method by the merged DIMD is evaluated to be inferior in performance to any other intra prediction encoding method, the candidate list of the merged DIMD may not be transmitted as a signal.
[0176]
[0177] Applied Examples
[0178] According to one embodiment of the present invention, the above-described DIMD-based encoding method can be implemented by combining and applying the various implementation methods described above.
[0179] According to one embodiment of the present invention, the DIMD-based encoding method described above may be performed by extracting at least one directional intra prediction encoding mode used in at least one non-DIMD-encoded block belonging to the first candidate group (1321) and / or the second candidate group (1322) described with reference to FIG. 13, and then assigning amplitudes to values of the modes, and combining the amplitudes derived from the blocks according to the above-described procedure, thereby deriving a histogram. Depending on the embodiment, this method may be used in parallel with or instead of the various types of histograms described herein.
[0180] According to one embodiment of the present invention, the DIMD-based encoding method described above may be configured to perform DIMD-based prediction, including derivation of a histogram and / or a merged histogram, as described above, based on reference blocks belonging to the first candidate group (1321) and / or the second candidate group (1322) described with reference to FIG. 13, and to perform DIMD-based predictive encoding for the current block (1310) based on the prediction result. At this time, at least one block vector, motion vector, or reference vector may be identified from blocks belonging to the current block (1310) and / or the first candidate group (1321) to designate the reference block. Depending on the embodiment, this method may be used in parallel with or instead of the various types of histograms described herein.
[0181] According to one embodiment of the present invention, the DIMD-based encoding method described above may be performed based on obtaining and / or generating a histogram from at least one DIMD and / or non-DIMD encoded block belonging to the first candidate group (1321) and / or the second candidate group (1322) described with reference to FIG. 13, and then merging the histograms to obtain and / or generate a histogram. Accompanying this encoding, the encoded video bitstream may include at least one index for indicating at least one block from a list of DIMD merge candidates that may include the first candidate group (1321) and / or the second candidate group (1322).
[0182] According to one embodiment of the present invention, the DIMD-based encoding method described above may be performed by extracting a histogram by the DIMD method post-hoc for at least one non-DIMD-encoded block belonging to the first candidate group (1321) and / or the second candidate group (1322) described with reference to FIG. 13, and then based on a merged histogram generated including the at least one extracted histogram. In this case, the histogram included in the merged histogram may be derived from at least one block included in the DIMD and / or non-DIMD method.
[0183] According to one embodiment of the present invention, when the DIMD-based encoding method described above generates a list of DIMD merging candidates from at least one DIMD and / or non-DIMD encoded block belonging to the first candidate group (1321) and / or the second candidate group (1322) described with reference to FIG. 13, the number of blocks included in the list of candidates or the type of blocks may be varied based on the horizontal and / or vertical sample size of the current block (1310). According to one embodiment, all or part of the first candidate group (1321) may be included, or only all or part of the second candidate group (1322) may be included, depending on the pixel sample size of the current block (1310).
[0184] According to one embodiment of the present invention, the above-described DIMD-based encoding method can take various fusion forms when performing the prediction fusion. For example, according to one embodiment, one or more intra prediction modes can be derived based on a histogram or a merged histogram, and the one or more intra prediction modes can be fused with a weight proportional to the amplitude size of each prediction mode in the histogram or the merged histogram. At this time, the number of intra prediction modes to be fused can be one, two, three, four, five, or more. At this time, the result by the planar prediction mode may or may not be fused. According to one embodiment of the present invention, the number of intra prediction modes to be fused can vary depending on the size of the block currently being encoded (decoded). For example, in the case of a large block, a larger number of intra prediction modes can be configured to be fused.
[0185] According to one embodiment of the present invention, the above-described DIMD-based encoding method may operate based on calculations based on blocks that are spatially adjacent to the current block (1310) as well as blocks that are temporally adjacent (e.g., may mean belonging to a frame that appeared immediately before or in the near past in display order or encoding order). For example, the calculation based on the Sobel filter may operate on blocks and / or pixels that are temporally adjacent and spatially co-located or adjacent. For example, information referenced for generating a merge histogram for the merged DIMD may be extracted from blocks and / or pixels that are temporally adjacent and spatially co-located or adjacent. For example, a DIMD merge candidate list generated for generating a merge histogram may be configured to include blocks of a third candidate group that include blocks that are temporally adjacent and spatially co-located or adjacent.
[0186] According to one embodiment of the present invention, the above-described DIMD-based encoding method may include a block vector-based DIMD technique that designates a reference block having similar characteristics to a current block through a block vector, a motion vector, or a similar reference vector, and utilizes information of the corresponding reference block. In one embodiment, the block vector may mean a vector indicating a spatial offset between a current block and a reference block within a current picture, and may be provided through a bit string, reuse a block vector used in a neighboring block of the current block, derive from intra / inter prediction information used in the current block or neighboring blocks, or derive through similarity analysis between the current block and another block within the picture. The reference block designated by the block vector may be regarded as a block included in the second candidate group and / or the third candidate group, and histogram information and / or intra prediction information obtained from the corresponding block may be utilized in the DIMD process of the current block. For example, the intra prediction mode of the current block may be determined by merging a histogram of the reference block and a histogram surrounding the current block.
[0187] The methods for implementing the various DIMDs of the present invention described above do not exclude the possibility that at least two implementation methods may be implemented in a completely and / or partially overlapping manner.
[0188]
[0189] Encoder and decoder
[0190] As with the characteristics of the DIMD technology described above, the procedure described through the encoding process in this specification can be implemented as a corresponding method in the decoding process, and if necessary, it can be implemented as the same procedure (for example, the process of predicting the value of a certain block can be performed in the same way in the encoding process and the decoding process), or implemented in reverse order (for example, a value transformed by a function can be inversely transformed by an inverse function), or implemented in a substantially corresponding other form, which is obvious to those skilled in the art. Therefore, it should be understood that all procedures described through the encoding process in this specification include the corresponding decoding method as the specification of the invention. In addition, as exemplified through the internal decoder (420) and the coding loop including the same in FIG. 4, the decoding process described above can be implemented in the same way within the encoder to predict the state of the decoder.
[0191] The encoding method according to the present invention described above can be implemented through an encoder as a device. The encoder as a device can be implemented in a form that maintains the conventional encoder structure exemplified through FIGS. 1 to 6 above or applies a predetermined change thereto. However, the form of implementation is not necessarily limited to that exemplified, and any form of encoder structure that can function as a video encoder should be considered an encoder established by the present invention as long as it implements the technical idea of the present invention.
[0192] In addition, the method for decoding the encoding result according to the present invention described above can be implemented through a decoder as a device. The decoder as a device can be implemented in a form that maintains the conventional decoder structure exemplified through FIGS. 1 to 6 above or applies a predetermined change thereto. However, the form of implementation is not necessarily limited to that exemplified, and any form of decoder structure that can function as a video decoder should be considered a decoder established according to the present invention as long as it implements the technical idea of the present invention.
[0193] A person skilled in the art will readily understand that a bit string encoded by the above-described method and device can be decoded by applying a method symmetrical and / or reverse to the encoding method. In one embodiment, when reading information for decoding from the encoded bit string, at least one variable length coded phrase included in the encoded bit string can be interpreted, and furthermore, in one embodiment, the variable length coding can be performed by an entropy coding method. The technical details and application method of implementing such a decoding procedure can be readily understood from the above-described encoding procedure.
[0194] The processor that may be included in the encoder and / or decoder described herein may mean one or more general purpose computers or special purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding.
[0195] Even if the processor is expressed singularly for ease of understanding, those skilled in the art will appreciate that the processor may include multiple processing elements and / or multiple types of processing elements. For example, a device according to one embodiment of the present invention may include multiple processors or one processor and one controller as the processor. Furthermore, the processor may be implemented using various processing configurations, such as a parallel processor or a multi-core processor.
[0196] The processor may be configured to execute an operating system (OS) and one or more software programs running on the operating system. Furthermore, the processor may access, store, manipulate, process, and generate data in response to the execution of the software.
[0197] The software may include a computer program, code, instructions, or a combination of one or more of these, and may be configured to control the processor to perform a desired operation and to issue commands to the processor, either independently or collectively. The software may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for interpretation by the processor or for providing commands or data to the processor. The software may also be distributed over networked computer systems, and stored or executed in a distributed manner.
[0198] The software may also be implemented in the form of program commands that can be executed through various computer means and recorded or stored in the memory. The memory may be a computer-readable recording medium, and program commands, data files, data structures, etc. may be recorded singly or in combination in the computer-readable recording medium. The program commands stored in the memory may be based on a command system specifically designed and configured for the embodiment of the present invention, or may follow a command system known and available to those skilled in the art of computer software, such as a command system exemplified by the assembly language, C, C++, Java, Python, etc. It should be understood that the command system and the program commands therefrom include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by the device and / or the processor according to an embodiment of the present invention using an interpreter or the like.
[0199] The computer-readable recording medium constituting the device according to one embodiment of the present invention, including the memory described in this specification, may include a temporary or volatile recording medium that is maintained only while the processor is operating, such as a processor cache, a RAM, a flash memory, or a relatively non-volatile or long-term recording medium, such as a magnetic media such as a hard disk, a floppy disk, and a magnetic tape, an optical media such as a CD-ROM, a DVD, a magneto-optical media such as a floptical disk, or a solid state memory, or may include a read-only recording medium, such as a ROM arranged on hardware, and further, the hardware itself configured to perform an operation equivalent to a series of program commands by a hard-wired structure by circuit wiring, and each step for performing the operation for implementing the embodiment of the present invention can be viewed as being recorded by the connection and arrangement of the hardware components, and thus the connection and arrangement method is the memory and It is obvious to a person skilled in the art that they can be considered equivalent.
[0200] The embodiments described above with respect to the processor and the memory are not mutually exclusive and may be selected or combined and implemented as needed. For example, a single hardware device may be configured to operate as a module composed of one or more of the software to perform the operations of an embodiment of the present invention, and vice versa. As another example, in the present specification, all or part of the operations assigned to a certain functional unit may be implemented by one or more of the software stored in a device according to an embodiment of the present invention (preferably, in any one of the recording media belonging to the category of the memory) and configured to be executed by the processor. In such a case, such a functional unit may be referred to as a functional unit "included" in the processor.
[0201]
[0202] Although the present invention has been described with reference to drawings and embodiments, as already mentioned above, it does not mean that the scope of protection of the present invention is limited to the drawings or embodiments presented above, and it will be understood that a person skilled in the relevant technical field can modify and change the present invention in various ways within a scope that does not depart from the spirit and scope of the present invention described in the claims of the present invention patent.
Claims
1. In the video decryption method, A step of obtaining histogram information for at least one surrounding block among first candidate group blocks adjacent to the current block and second candidate group blocks located close to but not adjacent to the current block; A step of generating a merged histogram for the current block by combining histogram information of the surrounding blocks; A step of deriving at least one intra prediction mode from the merged histogram; and An image decoding method comprising: a step of intra-prediction decoding the current block using the derived intra-prediction mode; 2. In the video encoding method, A step of obtaining histogram information for at least one surrounding block among first candidate group blocks adjacent to the current block and second candidate group blocks located close to but not adjacent to the current block; A step of generating a merged histogram for the current block by combining histogram information of the surrounding blocks; A step of deriving at least one intra prediction mode from the merged histogram; a step of intra-prediction encoding the current block using the derived intra-prediction mode; and A video encoding method comprising the step of including information indicating that the current block is encoded using a merged DIMD (Decoder-side Intra Mode Derivation) in a bitstream.
3. In the video decryption device, processor; Memory connected to the processor; A histogram information acquisition unit that acquires histogram information for at least one surrounding block among first candidate group blocks adjacent to the current block and second candidate group blocks located close to but not adjacent to the current block; A merged histogram generation unit that generates a merged histogram for the current block by combining histogram information of the surrounding blocks; A prediction mode derivation unit for deriving at least one intra prediction mode from the above merged histogram; and An image decoding device, comprising: an intra prediction decoding unit for intra prediction decoding the current block using the derived intra prediction mode.
4. In the video encoding device, processor; Memory connected to the processor; A histogram acquisition unit that acquires histogram information for at least one surrounding block among first candidate group blocks adjacent to the current block and second candidate group blocks located close to but not adjacent to the current block; A merged histogram generation unit that generates a merged histogram for the current block by combining histogram information of the surrounding blocks; A prediction mode derivation unit for deriving at least one intra prediction mode from the above merged histogram; An intra prediction encoding unit that intra-predicts and encodes the current block using the derived intra prediction mode; and An image encoding device, comprising a bitstream generation unit that includes information indicating that the current block is encoded using a merged DIMD (Decoder-side Intra Mode Derivation) into a bitstream.
Citation Information
Patent Citations
Intra-prediction direction refining method, intra-prediction direction refining device and intra-prediction direction refining program
JP2014225795A
Energy saving system and method for Cargo Hold within ventilation system
KR1020240045630A
Semiconductor device
KR1020250003146A
Method and device for exchanging secret keys based on reconfigurable and unclonable cryptographic component
KR1020250052002A