VIDEO ENCODING AND DECODING METHODS, AND DECODING EQUIPMENT

VN126616APending Publication Date: 2026-07-01UNIVERSITY INDUSTRY COOPERATION GROUP OF KYUNG HEE UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
VN · VN
Patent Type
Applications
Current Assignee / Owner
UNIVERSITY INDUSTRY COOPERATION GROUP OF KYUNG HEE UNIVERSITY
Filing Date
2024-09-26
Publication Date
2026-07-01

AI Technical Summary

Technical Problem

Conventional Local Illuminance Compensation (LIC) technology in video encoding has limited application methods, which restricts its ability to improve encoding efficiency effectively.

Method used

An improved LIC method that uses a cross filter with specific coefficients applied to each pixel, including a central pixel and adjacent pixels, to enhance prediction performance and minimize bitstream information capacity waste.

Benefits of technology

The improved LIC method enhances prediction performance, increases coding efficiency, improves video quality, reduces computational complexity, and optimizes hardware requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure VN1202603433_0
    Figure VN1202603433_0
Patent Text Reader

Abstract

The invention relates to video compression technology, and more specifically to a method of video encoding and decoding, and a decoding device. It relates to an improvement in Local Illuminance Compensation (LIC) technology, incorporated into predictive encoding technology that contributes to improved compression performance in video encoders and decoders. According to a sample implementation of the invention, a method for decoding an encoded video bitstream may include: acquiring at least one prediction vector for the current block being decoded from the bitstream; determining at least one reference block based on the prediction vector; determining at least one of the first patterns comprising at least one pixel adjacent to the current block with the current block as an anchor block, and the second pattern comprising at least one pixel adjacent to the reference block with the reference block as an anchor block; extracting parametric information related to the LIC by referencing at least one pixel information inferred from at least one of the first and second patterns;Determine at least one coefficient to be applied to the first filter based on parameter information; perform filtering on the reference block using the first filter to which the coefficients have been applied, to obtain the filtered result value; create prediction patterns for the current block based on the filtered result value; and reproduce the current block using the prediction patterns.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding / decoding method and device based on improved local illumination compensation cross-filter

[0001] The present invention relates to video compression technology, and more particularly, to improvement of local illuminance compensation (LIC) technology included in prediction coding technology that contributes to improvement of compression performance in video encoders and decoders.

[0002] The present invention may be in the same technical field as at least one of the digital video compression technology standards known by the names of standards such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG, or may be in the technical field for improving the inherent efficiency of the standards, or may be in the technical field for improving or replacing the standards.

[0003] Digital video encoding and decoding are widely used in various digital video applications. For example, digital television broadcasting, video transmission through communication networks, video calls / video conversations / video chats, recording and providing video content using optical media including video compact discs (VCDs) / digital versatile discs (DVDs) / Blu-Rays, all processes for producing, editing, collecting, and distributing video content, and devices such as video recording devices and camcorders for shooting and recording video for various reasons including personal, commercial, industrial, and security purposes, all depend on video encoding and decoding technologies.

[0004] Accordingly, implementations that may be referred to as digital video encoders and decoders may form part of a wide range of devices related to the generation, recording, and provision of digital video, including digital televisions, digital broadcasting systems, wireless broadcasting systems, computers in the form of notebooks / desktops / tablets, e-book readers, digital cameras, digital recording devices, digital multimedia playback devices, video game devices / terminals / consoles, mobile phones (including smartphones) with multimedia playback capabilities, equipment for video conferencing, and other devices.

[0005] The above digital video encoders and decoders can be implemented by a digital video compression standard that is widely used and understood by those skilled in the art. The digital video compression standard may include at least one of compression standards known by a standard name such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.

[0006] Video encoders and decoders can be implemented to more efficiently encode or decode digital video information while complying with the above standards, or by improving or modifying the above standards. Attempts to modify the above standards can also lead to the development of new standards. A well-known example is the so-called enhanced compression model (ECM), an attempt to improve and replace the existing H.266 / VVC standard, currently being developed by the Joint Video Experts Team (JVET), a joint international standardization group of ISO, IEC, and ITU-T.

[0007] Among the various detailed technologies applied to conventional video encoders and decoders, there is a technology collectively called local illuminance compensation (LIC). LIC technology refers to a technology that calculates a prediction residual by correcting illuminance (i.e., compensation) for the value of a prediction target block selected for predictive encoding in inter-frame predictive encoding (so-called inter encoding) and / or intra-frame predictive encoding (so-called intra encoding), and in a decoding process corresponding to each of the above encoding methods, and in particular, it can mean that the compensation is performed locally, for example, on a block-by-block basis.

[0008] Conventional LIC technology has limited applicability, and various improvements are expected to improve encoding efficiency. The present invention aims to address these technical challenges and provide improved encoding efficiency.

[0009] The present invention proposes an improved LIC implementation method to solve the above-described technical problems.

[0010] According to an embodiment of the present invention for solving the above-described technical problem, a method for decoding an encoded video bitstream may include the steps of: obtaining at least one prediction vector for a current block currently being decoded from the bitstream; determining at least one reference block based on the prediction vector; defining at least one of a first template including at least one pixel adjacent to the current block as a reference block, and a second template including at least one pixel adjacent to the reference block as a reference block; extracting parameter information related to local illumination compensation (LIC) with reference to at least one pixel information derived from at least one of the first template and the second template; determining at least one coefficient to be applied to a first filter based on the parameter information; performing filtering on the reference block using the first filter to which the coefficients are applied, thereby obtaining a filtering result value; generating a prediction sample for the current block based on the filtering result value; and restoring the current block using the prediction sample.

[0011] The first filter may be characterized as being a cross filter configured to receive at least one left pixel, at least one right pixel, at least one top pixel, at least one bottom pixel, and a center pixel as inputs, and having at least one coefficient that can be applied to each pixel.

[0012] The first filter may be characterized as a 3x3 sized cross filter configured to receive as input five pixels including a left pixel, a right pixel, an upper pixel, a lower pixel, and a center pixel, and having five coefficients that can be applied to each of the pixels.

[0013] The above first filter is,

[0014]

[0015] It may be characterized in that it operates by a mathematical formula expressed as, Pred is a filtering result value of the first filter, A represents a left pixel, B represents a right pixel, C represents an upper pixel, D represents a lower pixel, E represents an input value of a central pixel, and C1, C2, C3, C4, and C5 represent coefficients applied to each of the inputs.

[0016] The above first filter is,

[0017]

[0018] It operates by a mathematical formula expressed as, Pred is a filtering result value of the first filter, A represents a left pixel, B represents a right pixel, C represents an upper pixel, D represents a lower pixel, E represents an input value of a central pixel, F represents a nonlinear term based on E, and V mid is the median constant, V above div is a leveling division constant, and C1, C2, C3, C4, C5, and C6 may be characterized as representing coefficients applied to each of the inputs.

[0019] Above V mid represents the median value of the bit depth for the pixels of the current block ("bitDepth_mid"), and the V divDivision by can be characterized as being replaced by a right binary bit shift operation by the bit depth value (">>bitDepth").

[0020] The above first filter is,

[0021]

[0022] It operates by a mathematical formula expressed as, Pred is a filtering result value of the first filter, A represents a left pixel, B represents a right pixel, C represents an upper pixel, D represents a lower pixel, E represents an input value of a central pixel, F represents a nonlinear term based on E, and V mid is the median constant, V above div is the equalization division constant, V above const is a batch addition constant, and C1, C2, C3, C4, C5, C6, and C7 may be characterized as representing coefficients applied to each of the inputs.

[0023] Above V mid and above Vconst represents the median value of the bit depth for the pixels of the current block ("bitDepth_mid"), and the V div Division by can be characterized as being replaced by a right binary bit shift operation by the bit depth value (">>bitDepth").

[0024] The first filter may be characterized as an image processing filter configured to receive as input a group of pixels having at least one shape among a square, a diamond, and a circle, each having a width equal to a first pixel and a height equal to a second pixel, centered on a central pixel, and having at least one coefficient that can be applied to each pixel.

[0025] The step of determining at least one coefficient to be applied to the first filter may be characterized in that it is performed by a minimum mean squared error (MSE) minimization method based on a Wiener filter, which is calculated based on at least one of the restored pixel values ​​on the top and left of the current block.

[0026] According to an embodiment of the present invention for solving the above-described technical problem, a method for encoding a bit stream generated from a video comprises the steps of: determining at least one prediction vector for a current block currently being encoded; determining at least one reference block based on the prediction vector; defining at least one of a first template including at least one pixel adjacent to the current block as a reference block, and a second template including at least one pixel adjacent to the reference block as a reference block; extracting parameter information related to local illumination compensation (LIC) with reference to at least one pixel information derived from at least one of the first template and the second template; determining at least one coefficient to be applied to a first filter based on the parameter information; performing filtering on the reference block using the first filter to which the coefficients are applied, thereby obtaining a filtering result value; generating a prediction sample for the current block based on the filtering result value; encoding the current block using the prediction sample, and adding the encoding result to the bit stream. It may include a recording step.

[0027] The first filter may be characterized as being a cross filter configured to receive at least one left pixel, at least one right pixel, at least one top pixel, at least one bottom pixel, and a center pixel as inputs, and having at least one coefficient that can be applied to each pixel.

[0028] The first filter may be characterized as a 3x3 sized cross filter configured to receive as input five pixels including a left pixel, a right pixel, an upper pixel, a lower pixel, and a center pixel, and having five coefficients that can be applied to each of the pixels.

[0029] The above first filter is,

[0030]

[0031] It may be characterized in that it operates by a mathematical formula expressed as, Pred is a filtering result value of the first filter, A represents a left pixel, B represents a right pixel, C represents an upper pixel, D represents a lower pixel, E represents an input value of a central pixel, and C1, C2, C3, C4, and C5 represent coefficients applied to each of the inputs.

[0032] The above first filter is,

[0033]

[0034] It operates by a mathematical formula expressed as, Pred is a filtering result value of the first filter, A represents a left pixel, B represents a right pixel, C represents an upper pixel, D represents a lower pixel, E represents an input value of a central pixel, F represents a nonlinear term based on E, and V mid is the median constant, V above divis a leveling division constant, and C1, C2, C3, C4, C5, and C6 may be characterized as representing coefficients applied to each of the inputs.

[0035] Above V mid represents the median value of the bit depth for the pixels of the current block ("bitDepth_mid"), and the V div Division by can be characterized as being replaced by a right binary bit shift operation by the bit depth value (">>bitDepth").

[0036] The above first filter is,

[0037]

[0038] It operates by a mathematical formula expressed as, Pred is a filtering result value of the first filter, A represents a left pixel, B represents a right pixel, C represents an upper pixel, D represents a lower pixel, E represents an input value of a central pixel, F represents a nonlinear term based on E, and V mid is the median constant, V above div is the equalization division constant, V above const is a batch addition constant, and C1, C2, C3, C4, C5, C6, and C7 may be characterized as representing coefficients applied to each of the inputs.

[0039] Above V mid and above Vconst represents the median value of the bit depth for the pixels of the current block ("bitDepth_mid"), and the V div Division by can be characterized as being replaced by a right binary bit shift operation by the bit depth value (">>bitDepth").

[0040] The first filter may be characterized as an image processing filter configured to receive as input a group of pixels having at least one shape among a square, a diamond, and a circle, each having a width equal to a first pixel and a height equal to a second pixel, centered on a central pixel, and having at least one coefficient that can be applied to each pixel.

[0041] The step of determining at least one coefficient to be applied to the first filter may be characterized in that it is performed by a minimum mean squared error (MSE) minimization method based on a Wiener filter, which is calculated based on at least one of the restored pixel values ​​on the top and left of the current block.

[0042] According to an embodiment of the present invention for solving the above-described technical problem, a decoder device configured to decode a video bit stream encoded by a computing device includes a receiving unit for receiving a bit stream, an output unit for outputting a decoded video, a reference buffer for storing information of at least one decoded picture, a processor, and a memory for storing instructions executable by the processor, wherein the instructions obtain at least one prediction vector for a current block currently being decoded from the bit stream, determine at least one reference block based on the prediction vector, define at least one of a first template including at least one pixel adjacent to the current block as a reference block, and a second template including at least one pixel adjacent to the reference block as a reference block, and extract parameter information related to Local Illumination Compensation (LIC) with reference to at least one pixel information derived from at least one of the first template and the second template, and extract at least one parameter information applied to a first filter based on the parameter information. The method may be configured to include instructions for determining coefficients, performing filtering on the reference block using the first filter to which the coefficients are applied, obtaining a filtering result value, generating a prediction sample for the current block based on the filtering result value, and restoring the current block using the prediction sample.

[0043] The first filter may be characterized as being a cross filter configured to receive at least one left pixel, at least one right pixel, at least one top pixel, at least one bottom pixel, and a center pixel as inputs, and having at least one coefficient that can be applied to each pixel.

[0044] The step of determining at least one coefficient to be applied to the first filter may be characterized in that it is performed by a minimum mean squared error (MSE) minimization method based on a Wiener filter, which is calculated based on at least one of the restored pixel values ​​on the top and left of the current block.

[0045] According to the present invention, there is a beneficial effect of improving prediction performance by LIC technology and simultaneously minimizing overhead in the information capacity of a bitstream to be encoded.

[0046] According to the present invention, at least one effect of improving encoding efficiency, improving decoding efficiency, improving video quality, reducing computational load, reducing software size, reducing hardware size, and improving other performances related to encoding and decoding can be derived in video encoding and decoding.

[0047] Figure 1 is a conceptual diagram of a video communication system according to one embodiment of the present invention;

[0048] FIG. 2 is a conceptual diagram of the arrangement of an encoder and decoder in a real-time video streaming environment according to one embodiment of the present invention.

[0049] Figure 3 is a functional unit conceptual diagram of a video decoder according to one embodiment of the present invention;

[0050] Figure 5 is a conceptual diagram of a frame type according to one embodiment of the present invention;

[0051] Figure 6 is a conceptual diagram showing the structure of a video encoder according to the H.266 / VVC standard.

[0052] Figure 7 is a conceptual diagram showing the general LIC application method according to one embodiment of the present invention.

[0053] FIG. 8 is a conceptual diagram showing the shape and coefficient positions of a cross filter according to one embodiment of the present invention; and

[0054] FIG. 9 is an exemplary diagram showing an example of a template application mode according to one embodiment of the present invention.

[0055] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention.

[0056] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, the first component could be referred to as the "second component," and similarly, the second component could also be referred to as the "first component." The term "and / or" includes any combination of multiple related items described herein or any of multiple related items described herein, and is non-exclusive unless otherwise indicated. The listing of items in this application is merely an exemplary description to facilitate the spirit and possible implementation methods of the present invention, and therefore is not intended to limit the scope of embodiments of the present invention.

[0057] As used herein, "A or B" can mean "only A," "only B," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B and C."

[0058] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."

[0059] In this specification, "at least one of A and B" may mean "only A", "only B" or "both A and B". Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted identically to "at least one of A and B".

[0060] Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”

[0061] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.

[0062] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0063] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning within the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined herein.

[0064] In describing the invention in this application, embodiments may be described or illustrated in terms of unit blocks that perform the described function or functions. The blocks may be expressed in this application as one or more devices, units, modules, parts, etc. The blocks may be implemented in hardware by one or more logic gates, integrated circuits, processors, controllers, memories, electronic components, or information processing hardware implementation methods, but not limited thereto. Alternatively, the blocks may be implemented in software by application software, operating system software, firmware, or information processing software implementation methods, but not limited thereto. A single block may be implemented by being separated into multiple blocks that perform the same function, or conversely, a single block may be implemented to perform the functions of multiple blocks simultaneously. The blocks may also be implemented by being physically separated or combined according to any criteria. The blocks may also be implemented to operate in an environment where their physical locations are not specified and are separated from each other by a communication network, the Internet, a cloud service, or a communication method, but not limited thereto. All of the above implementation methods are within the scope of various embodiments that can be taken by a person skilled in the field of information and communication technology to implement the same technical idea, and therefore, any detailed implementation method should be interpreted as being included within the scope of the technical idea of ​​the invention of the present application.

[0065] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the attached drawings. To facilitate a comprehensive understanding of the present invention, the same reference numerals will be used for identical components in the drawings, and redundant descriptions of identical components will be omitted. Furthermore, the multiple embodiments are not mutually exclusive, and it is assumed that some embodiments may be combined with one or more other embodiments to form new embodiments.

[0066]

[0067] digital video codec

[0068] Figure 1 is a conceptual diagram of a video communication system according to one embodiment of the present invention. The video communication system (100) may be configured to include at least two terminals (110, 120) connected to each other via a network (105).

[0069] In one embodiment of the present invention, the above-described FIG. 1 may refer to a block diagram for configuring a unidirectional video communication network. A first terminal (110) among the terminals may encode video data in order to transmit (111) the video data through a network (105). A second terminal (120) among the terminals may be configured to receive (121) the encoded video data through a network and decode and display the same.

[0070] In another embodiment of the present invention, the above-described FIG. 1 may refer to a block diagram for configuring a two-way video communication network. For the two-way video communication, each terminal (110, 120) may be configured to encode video data acquired by itself for video transmission (112, 122) to each other terminal via the network. Each terminal may also be configured to receive (113, 123) video data transmitted by another terminal via the network, decode the same, and display the decoded video data.

[0071] The terminals (110, 120) shown in Fig. 1 may be exemplified as devices such as server computers, personal computers, portable computers, and smartphones, depending on the embodiment, but are not limited thereto, and may be any computing device generally used. For example, according to an embodiment of the present invention, each terminal (110, 120) may mean a desktop computer, a laptop computer, a tablet PC, a mobile phone, a smart phone, a personal digital assistant (PDA), a workstation, an electronic calculator, a server computer, a cloud computer, a virtual computer, a quantum computer, or any other electronic, electrical, or quantum computing device implemented in a movable or non-movable form, and in particular, it may be interpreted as any device that is designed to operate as a terminal device according to an embodiment of the present invention among such devices, is capable of operating as a terminal device according to an embodiment of the present invention, or is capable of installing and / or executing a computer program that enables the terminal device to operate as an embodiment of the present invention or perform a method corresponding to such an operation.

[0072] Each of the above terminals (110, 120) may be implemented by a plurality of functional units that are interconnected in various forms, such as a bus, a circuit, or a relationship between a routine and a subroutine, and configured to exchange information within each of the above terminals (110, 120). In addition, through the interconnection, the terminals may be configured to include a processor (130) having a calculation function and a memory (140) connected to the processor for the purpose of executing or supporting the operation of a functional unit that primarily requires calculation among the above functional units.

[0073] The processor (130) described in this specification may mean one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding.

[0074] Even if the processor (130) is expressed singularly for the sake of ease of understanding, a person of ordinary skill in the art will recognize that the processor (130) may include multiple processing elements and / or multiple types of processing elements. For example, a device according to one embodiment of the present invention may include multiple processors or one processor and one controller as the processor (130). In addition, the processor (130) may be implemented using various processing configurations, such as a parallel processor or a multi-core processor.

[0075] The processor (130) may be configured to execute an operating system (OS) and one or more software programs running on the OS. In addition, the processor may access, store, manipulate, process, and generate data in response to the execution of the software. The software program may include a computer program, code, instructions, or a combination of one or more of these, and may configure a processing device to operate as desired or independently or collectively command a processing device. The software may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave to be interpreted by the processor (130) or to provide instructions or data to the processor. The software may be distributed among multiple computer systems connected to the network (105), such as the terminals (110, 120), and stored or executed in a distributed manner.

[0076] The software may also be implemented in the form of program commands that can be executed through various computer means and recorded or stored in the memory (140). The memory (140) may be a computer-readable recording medium, and program commands, data files, data structures, etc. may be recorded singly or in combination in the computer-readable recording medium. The program commands stored in the memory (140) may be based on a command system specifically designed and configured for the embodiment of the present invention, or may follow a command system known and available to those skilled in the art of computer software, for example, a command system exemplified by the assembly language, C, C++, Java, Python, etc. It should be understood that the command system and the program commands therefrom include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by the device and / or the processor (130) according to an embodiment of the present invention using an interpreter, etc.

[0077] The computer-readable recording medium constituting the device according to one embodiment of the present invention, including the memory (140) described in this specification, may include a temporary or volatile recording medium that is maintained only while the processor is operating, such as a processor cache, a RAM, a flash memory, or a relatively non-volatile or long-term recording medium, such as a magnetic media such as a hard disk, a floppy disk, and a magnetic tape, an optical media such as a CD-ROM, a DVD, a magneto-optical media such as a floptical disk, or a solid state memory, or may include a read-only recording medium, such as a ROM arranged on hardware, and furthermore, the hardware itself configured to perform an operation equivalent to a series of program commands by a hard-wired structure by circuit wiring, and each step for performing the operation for implementing the embodiment of the present invention can be viewed as being recorded by the connection and arrangement of the hardware components, and therefore, the connection and arrangement method It is obvious to those skilled in the art that this can be seen as equivalent to the above memory (140).

[0078] The embodiments described above with respect to the processor (130) and the memory (140) are not mutually exclusive, and may be selected or combined and implemented as needed. For example, one hardware device may be configured to operate as a module composed of one or more of the software to perform the operations of an embodiment of the present invention, and vice versa. As another example, in the present specification, all or part of the operations assigned to a certain functional unit may be implemented by one or more of the software stored in a device according to an embodiment of the present invention (preferably, in any one of the recording media belonging to the category of the memory) and configured to be executed by the processor, and in this case, such a functional unit may be referred to as a functional unit "included" in the processor.

[0079] The present invention is applicable to all environments for creating a one-way or two-way video communication network, and it should be understood that the network (105) can be created by any means for transporting encoded video data between the terminals (110, 120).

[0080] In one embodiment of the present invention, the network (105) may refer to a wired or wireless communication network. In this case, depending on the embodiment, the network may be configured to communicate information using any communication standard, and the communication standard may include packet-based communication. The packet communication may be understood to mean including packets known as TCP or UDP, for example. The wired communication method of the network (105) may be a method of connecting to an external communication network by means of a telephone line of the RJ-11 standard, an Ethernet cable belonging to various categories of the RJ-45 standard, other coaxial cables, metal cables, optical cables, and various other wired media. The wireless communication method of the above network (105) may include, depending on the embodiment, a short-range wireless communication method including Bluetooth, Wi-Fi, Zigbee, and NFC (near field communication), or may include a long-range wireless communication method that may be referred to as a name of a wireless communication technology collectively called a generation name of a communication standard such as Wibro, WiMax, Global Systems for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long-Term Evolution (LTE), New Radio (NR), and other 2G, 3G, 4G, 5G, and 6G, and a name of an international standard wireless communication standard such as IMT-2000, IMT-Adcanced, IMT-2020, and IMT-2030. Of course, even if any conventional or newly developed wired or wireless communication means, method, standard, and protocol are applied to the implementation of the network (105), there is no problem in achieving the purpose of the present invention as long as it is a means configured to perform transmission and reception in an information and communication device such as the terminal (110, 120).It is also obvious that one network (105) can be configured in a mixed manner by one or more wired and / or wireless standards.

[0081] However, in another embodiment of the present invention, the network (105) may be understood to include a process of transmitting information using a computer-readable recording medium. In this case, the configuration of the network is not limited to a communication medium, and should be understood to include a process of temporarily storing and physically transporting information in a computer-readable memory and / or recording medium. The computer-readable recording medium used for transmitting information may be understood to mean a relatively non-volatile or long-term recording medium such as a magnetic media such as a hard disk, a floppy disk, and a magnetic tape, an optical media such as a CD-ROM or DVD, a magneto-optical media such as a floptical disk, or a solid state memory, which is mainly used as a means of transporting data between computing devices.

[0082] Any other means of information communication or transport, regardless of the method employed, can be considered within the scope of the present invention as long as it has a structure that supports the transmission and decoding of video data in an encoded state. Therefore, in addition to the examples listed above, any means of information communication or transport, whether known in the past or newly available, can fall within the scope of application of the present invention.

[0083] FIG. 2 is a conceptual diagram illustrating the arrangement of an encoder and decoder in a real-time video streaming environment according to one embodiment of the present invention. The streaming system (200) illustrated in FIG. 2 can be applied to video data communication networks, including, for example, digital broadcasting, video telephony, and video conferencing. However, it should be noted that technical structures identical or similar to the streaming system can be equally applied even when information is transmitted via a recording medium, as described above.

[0084] According to one embodiment of the present invention, the streaming system may include a video source (210) that generates a video stream. The video source may include a digital video acquisition means (212), which may be, for example, a digital camera or other device, for acquiring uncompressed raw video. The raw video stream (215) may have a large capacity and may therefore be compressed by a video encoder (217) coupled or connected to the video source.

[0085] The above encoder (217) may be configured as a means including hardware, software, or a combination of the two, configured to implement an image encoding method and / or an implementation method thereof according to one embodiment of the present invention.

[0086] Through the encoder (217), an encoded bitstream (219) having a reduced capacity compared to the original video stream can be output. The bitstream (219) can be provided in real time via a relay device, which may be referred to as a streaming server (220), and / or can be stored in a recording medium (225) of the streaming server (220) for subsequent use.

[0087] The streaming system (200) may include at least one streaming client (230, 240) that connects to the streaming server (220) to receive the encoded bit string (229) in real time or obtain it later. The streaming client may include a video decoder (232) that obtains the encoded bit string (229) (which may also be regarded as a copy of the bit string (219) received by the streaming server), decodes the bit string (229), and outputs the resulting video data as video data in a form that can be displayed by a display (235) or other visual, auditory, or other sensory display means.

[0088] As described above, the functions for encoding and decoding video data are collectively called a coder-and-decoder, or video codec.

[0089] FIG. 3 is a conceptual diagram of a functional unit of a video decoder according to an embodiment of the present invention. As shown in FIG. 3, a receiving unit (310) can receive at least one encoded video data to be decoded by a decoder (305). In an embodiment of the present invention, the encoded video data may be independent for each reception, and the decoding procedure of each independent video data may be independent from the decoding procedure of other video data. The encoded video data may be received by the receiving unit (310) through a hardware or software connection (315) to a device storing the same, and as described above, the storing device may be a type of streaming server located at the other end of a communication network, or may mean a physical recording medium, but is not limited thereto.

[0090] The above-described receiving unit (310) can receive the encoded video data together with other data accompanying it, such as encoded audio data or other auxiliary data, and each of the data can be separated from the video data and provided to an appropriate processing function unit (312) other than the video decoder.

[0091] When the video data is provided through a communication network, a buffer memory (320) may be coupled between the receiving unit (310) and the decoder (305) to minimize delay and disconnection according to the network environment. The buffer memory (320) may refer to a computer-readable recording medium that temporarily stores the received video data and stably supplies it to a parser (330) corresponding to the input terminal of the decoder (305). However, if the bandwidth of the communication network is sufficient, if video data is read from a recording medium in a local location that is not physically separated, or if the possibility of communication delay is not predicted in other environments, the buffer memory may be unnecessary.

[0092] The video decoder (305) may include the parser (330) as its input terminal to interpret the encoded video data. The parser may perform a function of separating (parsing) a plurality of pieces of information stored in the form of a bit string in the encoded video data according to a predetermined rule, and, if necessary, performing an entropy decoding (335) of entropy-coded video data, thereby performing a function of reconstructing symbols (338), which are paragraphs of video encoding information. The symbols (338) may include all information for controlling the operation of the decoder (305), and / or may further include information for controlling a device that is attached to and operable with the decoder (305), such as a display device. Control information for controlling the above display device may include information in a format called supplementary enhancement information (SEI) or video usability information (VUI).

[0093] As described above, the parser (330) may be configured to perform entropy decoding (335) of the encoded video data. The entropy encoding method of the encoded video data may vary depending on the encoding standard, and decoding may be performed accordingly. Representative examples of the entropy encoding standard may include variable length coding, Huffman coding, and arithmetic coding, and each of the encoding methods may be a context-adaptive or context-sensitive method depending on the standard, and may also be based on principles widely known to those skilled in the art.

[0094] The parser (330) may be configured to extract at least one partial image from the encoded video data. The definition of the partial image may vary depending on the encoding standard, and depending on the standard, one or more of the examples listed below may correspond simultaneously and overlappingly. The partial image may be defined in units such as, for example, a group of pictures (GOPs), pictures / frames, tiles, slices, macroblocks, blocks, subblocks, transform units (TUs), and prediction units (PUs).

[0095] The parser (330) may be configured to extract encoding information, such as transform coefficients, quantization parameters (QPs), and / or motion vectors, from the encoded video data. The parser (330) may be configured to perform entropy decoding (335) and parsing operations on the video data received from the buffer memory, and to selectively decode symbols (338) representing the encoding information. In addition, the parser (330) may be configured to selectively supply a specific symbol (338) to a specific decoding function unit within the decoder (305), such as an inverse quantization and inverse transform unit (340), an intra prediction unit (350), an inter prediction unit (355), or a loop filter unit (360). Control of such information supply can be determined by the information sequence contained in the encoded video, and may vary depending on the encoding standard, and is not limited within the scope of the embodiments of the present invention, and is not described in detail in this conceptual diagram.

[0096] The decoder (305) may be comprised of a number of conceptual functional units that receive and process the encoded information provided by the parser (330). It should be readily apparent that these conceptual functional units may be combined or further subdivided, depending on implementation needs. For example, they may be further separated for ease of implementation, or integrated into one for operational efficiency. In any case, each functional unit may be configured to perform close interaction with each other. However, despite the possibility of such integration or separation, the following description will be given as a combination of conceptual functional units to illustrate the video data decoding procedure applied as an embodiment of the present invention.

[0097] The decoder may include an inverse quantization and inverse transformation unit (340). The inverse quantization and inverse transformation unit (340) may be configured to receive encoding information including a method to be used for numerical transformation (transform), a block size, quantization coefficients for recovering quantized information, and distinction information of a quantization matrix that simplifies and represents the quantized coefficients from the parser (330), and may be configured to output block values ​​(341) that can be input to an aggregator (370) as a result of processing the encoding information.

[0098] In one embodiment of the present invention, the output values ​​of the inverse quantization and inverse transformation unit (340) may include intra-prediction encoded block values. The intra-predicted block values ​​may refer to values ​​that can be decoded without using prediction information from a previously decoded partial image, for example, a previous frame, but using prediction information within a partial image currently being decoded, for example, a current frame.

[0099] Prediction information within the current partial image may be provided by the intra prediction unit (350). According to an embodiment of the present invention, the intra prediction unit (350) generates a block value of the same form as the block being decoded as prediction information by using image information of a spatially adjacent area derived from a partial image currently being decoded and of which decoding has been partially completed. The partial image information may be provided (381) from a current image buffer, a so-called line buffer (380). The merging unit (370), according to an embodiment, may be configured to merge the prediction information (351) generated by the intra prediction unit (350) with the block values ​​(341) provided by the inverse quantization and inverse transformation unit (340).

[0100] In another embodiment, the output values ​​of the inverse quantization and inverse transformation unit (340) may include block values ​​subjected to motion compensation as inter-prediction encoded block values, and in some cases, block values ​​subjected to motion compensation. In this case, the inter-prediction unit (355) may extract and use sample information (386) used for motion-based prediction from a reference picture buffer (385). The information (356) derived by performing motion compensation on the sample information based on the symbols (338) included in the block values ​​as the output values ​​may be configured to be merged with the block values ​​(341) provided by the inverse quantization and inverse transformation unit (340) by the merger unit (370). In this case, the block values ​​(341) may be referred to as so-called differential or residual values.

[0101] The position information within the memory used by the inter prediction unit (355) to extract the sample information from the reference image may be determined by a motion vector provided to the inter prediction unit (355) which is composed of a combination of symbols (338) for representing, for example, X, Y, and other specific points of the reference image. The inter prediction unit (355) may also include a function for interpolating and using the sample values ​​when a so-called 'subsampling' capable motion vector is provided, and may further include a function for predicting and reinforcing the value of the motion vector.

[0102] The output values ​​(371) of the above merging unit (370) may be provided to the loop filter unit (360) and processed by various loop filtering methods. The loop filter unit (360) may be configured to receive not only the block unit output (371) of the merging unit (370) but also the symbol (338) provided from the parser (330) to control its operation. The output of the loop filter unit (360) may be output to an external display means such as the display device through an output connection (390), but may be stored (361) in a line buffer (380) for use in prediction for interpreting a subsequent intra or inter coded block value, and may also be stored in a reference image buffer (385) through this.

[0103] Certain partial images, such as frames, after their decoding is completed, can be utilized as reference images for performing predictive decoding in a subsequent decoding process. One partial image, such as a frame, can be gradually accumulated in a line buffer (380) and decoded, and when one frame is decoded, the contents of the line buffer (380) are transferred (383) to the reference image buffer (385), and a new line buffer (380) can be allocated for decoding the new frame.

[0104] The above video decoder (305) may be configured to perform a decoding operation according to a predetermined video compression technique that may be documented by various international standards or commercial standards. The standards may include, for example, international standard recommendations such as H.264, H.265, and H.266 defined by the International Telecommunication Union Standardization Sub-Division (ITU-T). Those skilled in the art will understand that each of the above recommendations is equivalent to an international standard jointly defined by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC). The encoded video data may comply with a specific bitstream syntax defined by the relevant standards, as defined and required by the video compression standard documents and standard documents, and specifically by the profiles and levels specified within such documents. In addition, the complexity of the encoded video data may be limited to a certain level to comply with the profiles and levels. For example, a profile or level may be configured to limit a maximum picture size, a maximum decoding speed, and a maximum reference picture size. These limitations may, in some embodiments, also be further constrained via metadata signals for a hypothetical reference decoder (HRD) and HRD buffer management included in the encoded video data.

[0105] According to one embodiment of the present invention, the receiver (310) may receive additional redundant data along with the encoded video. The additional data may be considered as part of the encoded video data. The additional data may include information that may be used by the decoder (305) to properly decode the data or to more accurately reconstruct an image that approximates the original image. The additional data may be provided in the form of, for example, layers for temporal, spatial, or signal-to-noise ratio (SNR) enhancement, redundant slices, redundant images, and forward error correction codes.

[0106] FIG. 4 is a functional unit conceptual diagram of a video encoder according to one embodiment of the present invention. The encoder (405) may be configured to receive original video information (402) from a video source (401) and perform encoding.

[0107] The original video information (402) may have any suitable bit depth, for example, 8 bits, 10 bits, 12 bits, etc. In addition, the original video information (402) may have any suitable color space, for example, R / G / B, Y / U / V, Y / Cb / Cr, etc. In addition, the original video source may have any suitable sampling structure corresponding to the color space, for example, Y / Cb / Cr 4:2:0, Y / Cb / Cr 4:4:4, etc. The original video source having such a predetermined format may be provided to the encoder in the form of a digital video stream.

[0108] In a one-way video communication network, the original video information (402) can be obtained from a recording medium storing a previously prepared video source. In a two-way video communication network, the original video information (402) can be obtained from an image acquisition device, such as a camera, that generates at least one video transmission stream included in the two-way video communication.

[0109] The video data including the above original video information (402) may be configured as a plurality of partial images configured to simulate motion by being played back in chronological order. The partial images may be expressed as concepts such as pictures or frames, for example. The partial images may include one or more samples depending on the type of sampling structure, color space, etc. being used. Those skilled in the art will understand that the terms "samples" and "pixels" in digital images are closely related. The operation of the encoder will be described below with reference to such samples.

[0110] According to one embodiment of the present invention, the encoder (405) may be configured to encode and compress (partial) images constituting the original video information (402) in real time (or according to other temporal requirements required according to the implementation method) into the form of encoded video information.

[0111] In the encoder (405), the control unit (450) may be a functional unit configured to control an appropriate encoding speed. The control unit (450) may be configured to control other functional units and be functionally coupled to the following functional units as described below. The parameters set by the control unit (450) may include parameters related to bitrate control, such as skip of the image, quantizer, variable values ​​for applying a picture quality optimization technique, and may also include values ​​such as the size of the image, the structure of a group of pictures (GOP), and the maximum search range of a motion vector. A person skilled in the art will be able to understand various other functions that the control unit (450) may have, and such other functions may be added or removed according to the design of a video encoder optimized for an individual system design.

[0112] According to an embodiment of the present invention, the encoder (405) may be configured to operate in a structure such as a "coding loop" well known to those skilled in the art. To simplify the description by way of example, the coding loop may be configured with an internal encoder (so-called "source coder") (410) responsible for receiving an image to be encoded and generating symbols based on at least one reference image that has been encoded in the past, and a local decoder (420) configured to be connected to the internal encoder. The local decoder (420) may be configured to perform an operation to reproduce sample data to be generated by a decoder (490) located at an actual remote location that will receive encoded video information from the encoder (405) by receiving an output of the internal encoder (410).

[0113] Video data composed of sample data reconstructed by the internal decoder (420) may be configured to be input to the reference picture buffer of the encoder (405). As described above, the internal decoder (420) is implemented to reproduce the result output by the encoder (405) and to be decoded by a remote decoder, so the video data recorded in the reference picture buffer may also be identical in bit units to the information of the reference picture buffer of the remote decoder. That is, the prediction function unit that may be included in the encoder (405) may read the same values ​​as the sample values ​​of the previous frame that the decoder will later refer to in the decoding process from the reference picture buffer of the encoder (405).

[0114] As described above, the principle of achieving matching of the reference image buffer between the encoder (405) and the decoder (490) by means of the internal decoder (420) on the encoder (405) side is well known to those skilled in the art, and a method of responding to an environment in which such an environment is not guaranteed (e.g., information loss due to communication failure, etc.) can also follow what is known to those skilled in the art.

[0115] An embodiment of the operation method of the internal decoder (420) has been described in detail above with reference to FIG. 3. The decoder of FIG. 3 may be regarded as the aforementioned "remote" decoder (490). The internal decoder (420) may be implemented excluding lossless encoding and decoding sections such as the parser (330) or entropy decoding (335). This is because the internal encoder (405) is implemented to simply reproduce the operation of a decoder located at a remote location, and thus may directly decode symbols without requiring a process of compressing and then decompressing symbols. Accordingly, the functional units preceding the parser and entropy decoder as shown in FIG. 3 may not be provided or may be implemented at least partially.

[0116] As described above, according to a preferred embodiment of the present invention, any decoder function (excluding a parser and an entropy decoder) present in the decoder can naturally exist as a substantially identical function in the corresponding encoder (405).

[0117] The operation of the encoding function unit that may be included in the above encoder (405) can be considered as the inverse of the decoder function unit. Therefore, the embodiment can be explained by performing the operation of the decoder function unit in reverse. For example, a quantization and transform function unit corresponding to the inverse quantization and inverse transform unit may be provided, and an inter prediction encoding unit corresponding to the inter prediction unit may be provided. In addition, some additional explanations will be added.

[0118] The internal encoder (410) may be configured to perform encoding on input image information, for example, an input frame, by a predictive encoding method executed by a predictive encoding unit (440) that operates by referencing at least one temporally previous encoded partial image, for example, frames, from a reference image buffer (430) from at least one reference image information, for example, video data designated as a reference frame. In this case, the encoder (405) may be configured to encode a differential between blocks of samples constituting the input image and blocks of samples constituting the reference image.

[0119] The internal decoder (420) can decode video data that can be designated as the reference picture from symbols generated by the internal encoder (410). As described above, since the video data is subject to the same decoding operation as that performed by a remote decoder, the video data used as the reference picture may be provided to the encoder (405) in a form that has undergone lossy compression and has suffered some damage, and this operation may be intended to ensure operational consistency with the decoder.

[0120] The prediction encoding unit (440) may be configured to perform a prediction search operation within the encoder (405). The prediction search operation may refer to an operation corresponding to the inter prediction or intra prediction described in the description of the decoder. For image information that is input and scheduled to be newly encoded, the prediction unit may access the reference image buffer (430) to retrieve information such as a motion vector, a block shape, and metadata that may include the same, which are information indicating points of a reference image that can function as prediction reference information suitable for the new image information, and a sample block to be actually referenced. The prediction encoding unit (440) may operate on the basis of a so-called "sample block by pixel block" to obtain appropriate prediction reference information. According to one embodiment of the present invention, at least one prediction reference information may be designated for the input image, which designates at least one reference image information stored in the reference image buffer (430), as determined based on the search results obtained by the prediction encoding unit (440).

[0121] In one embodiment of the present invention, the control unit (450) may be configured to manage the overall encoding operation of the internal encoder (410), including setting parameters used to encode video data.

[0122] All outputs of the above-described functional units may be subjected to entropy encoding (460) in order to be finally output. The entropy encoding (460) may include various entropy coding techniques, such as variable length coding, Huffman coding, and arithmetic coding, for the symbols generated by the various functional units as described above, and each encoding method may be a context-adaptive or context-sensitive method according to the standard, or may be based on principles widely known to those skilled in the art. Such entropy encoding (460) can typically achieve lossless compression, and thus can be configured to convert at least one symbol generated by the functional units into encoded video data.

[0123] The above control unit (450) may, when controlling the operation of the encoder (405), apply the type of encoding of a specific partial image during the encoding period to each partial image, such as a picture or frame. Depending on the type, the method by which the partial image is encoded may be affected. Depending on the embodiment, the type may include what is categorized as the following "frame type."

[0124] Fig. 5 is a conceptual diagram of a frame type according to one embodiment of the present invention. The following description will be made with reference to Fig. 5.

[0125] An intra (“I”) picture (510) may refer to a picture that can be encoded and decoded using only its own information without referring to other picture information in the video data through predictive encoding. The “I” picture may be designated by names such as a key frame, an independent / instantaneous decoder referh (IDR) frame, and a clean random-access (CRA) frame, depending on the video encoding standard, and the “I” pictures designated by the various names as described above may have various modifications and application methods as permitted by each standard and may be partially different from each other. In addition to those listed above, various application methods for implementing the “I” picture may be by various methods that are already known to those skilled in the art or may be newly provided.

[0126] A prediction ("P") picture (520) may refer to a picture that can be encoded and decoded through intra or inter prediction based on at least one prediction information and / or a motion vector that designates at least one reference picture to predict sample values ​​of a block constituting the picture. The "P" picture may be configured to refer to only one reference frame, or may be configured to refer to one or more reference frames, according to a video encoding standard. When referring to more than one reference frame, sample information and / or associated metadata derived from multiple reference pictures may be used to reconstruct a single block. However, in common cases, a picture designated as a "P" picture may be understood as a picture that performs reference only to a temporally preceding picture.

[0127] A bidirectional prediction ("B") picture (530) may refer to a picture that can be encoded and decoded through intra or inter prediction based on at least one prediction information and / or motion vector that designates at least two reference pictures to predict sample values ​​of blocks constituting the picture. In a common case, a picture designated as the "B" picture is distinct from a picture designated as the "P" picture, and may be understood as a picture that performs a reference without being limited to a temporally preceding picture.

[0128] Video data may be spatially divided into a plurality of sample blocks during the encoding and decoding process, and encoding may be performed in units of the blocks. The block units may include, but are not limited to, sizes such as 4x4, 8x8, 4x8, or 16x16 in units of horizontal / vertical pixels, as is widely known. The block may be encoded using a predictive encoding method with reference to any other (already encoded) blocks, as permitted and / or restricted by the type specified for each partial picture including the block. For example, the blocks of the "I" picture (510) may be encoded without using a predictive encoding method, or with reference to blocks already encoded within the same partial picture. That is, only the so-called intra prediction method may be used. In contrast, the "P" picture (520) may further reference a reference picture encoded in at least one previous time unit, and thus, inter prediction may also be used for encoding along with intra prediction. In the case of a "B" picture (530), reference can be made not only to a previously encoded picture in the encoding order but also to a later reference picture in terms of time unit. However, it is widely known that there may be blocks within a "P" picture or a "B" picture that are encoded without relying on predictive encoding.

[0129] The above video encoder (405) may be configured to perform encoding operations according to a predetermined video compression technique that may be documented by various international standards or commercial standards. Examples of the above standards may include all of those described in the above decoder.

[0130] According to one embodiment of the present invention, the transmitter (470) may buffer the encoded video data generated by the entropy encoding in order to provide / transmit the video data (ultimately to a remote decoder (490)) to a device storing the encoded video data via a hardware or software connection (495). According to an embodiment, when providing / transmitting the encoded video data from the video encoder (405), the transmitter (470) may receive and merge other data accompanying the encoded video data, for example, encoded audio data or other auxiliary data, from a separate source (480).

[0131] According to one embodiment of the present invention, the transmitter (470) may be configured to transmit additional data along with the encoded video. The additional data may be considered part of the encoded video data. The additional data may include information that can be used by a decoder to properly decode the data or to more accurately reconstruct an image that approximates the original image. Examples of the additional data may include all of the examples previously presented with respect to the receiver (310) of the decoder.

[0132] The present invention can be implemented by a digital video compression standard that is widely used and understood by those skilled in the art as described above. The digital video compression standard may include at least one of compression standards known by the standard name such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.

[0133] Fig. 6 is a conceptual diagram illustrating the structure of a video encoder according to the H.266 / VVC standard. What is depicted in Fig. 6 corresponds to the rough structure of a video encoder widely known by standard codes such as ITU-T H.266 and ISO / IEC 23090-3, and also by the name MPEG-I Part 3 or the general name versatile video coding (VVC).

[0134] According to FIG. 6, a video encoder (605) may be configured to receive uncompressed and unencoded original video data (601) as input and output an encoded bit stream (602). The video data (601) may be supplied directly to a luma mapping unit (610a) when intra-encoded, or may be supplied to a luma mapping unit (610b) via an inter-prediction unit (620) including motion vector extraction. In the intra-encoded case, the mapped luma signal may be supplied to an output merger (606) by selecting (608) at least one of an intra-prediction encoded signal via an intra-prediction unit (625) or an inter-prediction encoded signal output from the luma mapping unit (610b) via the inter-prediction unit (620). The result of the above output merger can be applied to a chroma scaling unit (615). (The operation of the luminance signal mapping unit (610) and the operation of the chroma scaling unit (615) are collectively referred to as a luma mapping / chroma scalaing (LMCS) process.) The reduced chroma signal can be provided to a transform unit (630), and the transform unit (630) can perform an adaptive color transform, particularly on the chroma signal. The coefficients derived as a result of the transform are applied to a quantization unit (640) and quantized. As a result, lossy compression is achieved, and the result of the lossy compression can be output as a bit string (602) through a multi-hypothesis CABAC (650), which is a lossless compression method.

[0135] Meanwhile, the result of the lossy compression may actually enter the decoding process by going through the processes of inverse quantization (645), inverse transform (635), and luminance signal expansion (617) to generate a coding loop. The result of the luminance signal expansion may be supplied to the internal merger (607) together with the result of selecting (608) at least one of the previously generated intra prediction encoding signal or inter prediction encoding signal. The result of the internal merger may go through inverse luma mapping (617), and then may go through processing such as a deblocking filter (660), sample adaptive offset (SAO) (670), and an adaptive loop filter (ALF) to reproduce the image quality improvement process in the decoder. The result of reproducing the operation in the decoder as described above is applied to the reference image buffer (690) and can be reused for prediction encoding by the inter prediction unit (620).

[0136] The present invention can also be used by or in combination with the enhanced compression model (ECM), which is an implementation of a next-generation video codec being developed by the joint video experts team (JVET), an international standardization expert group, to improve H.266 / VVC. According to the standardization progress document of the above JVET, document number ISO / IEC JTC 1 / SC 29 / WG 5 N 190 (also document number JVET AC2025-v1), the improved compression model may include an improved intra prediction coding method, an improved inter prediction coding method, an improved transform and transform coefficient coding method, an improved adaptive loop filtering method, a bilateral filtering method, a new sample adaptive offset (SAO) method for image quality improvement, an extended entropy coding method, and an improved gradual decoding refresh (GDR) technique, and such techniques may be included in the present invention as background techniques for implementing the present invention.

[0137]

[0138] General LIC application method

[0139] The present invention provides a method for applying an improved local illuminance compensation (LIC) technique that can be used in the video encoding and decoding field including the above-described embodiments, and a device to which such a method is applied.

[0140] Fig. 7 is a conceptual diagram illustrating the general structure of a LIC application method according to an embodiment of the present invention. To address the problem of reduced compression efficiency due to differences in illumination between frames, Weighted Prediction (WP) is applied in traditional video coding standards. The WP is applied on a frame-by-frame basis, and is a technology that compensates for brightness changes in a reference block to accurately obtain a prediction value when encoding the current frame. The WP is applied to both H.264 / AVC and H.265 / HEVC, and can be used in an explicit mode or an implicit mode depending on whether or not a brightness compensation parameter is transmitted for each slice.

[0141] While the above WP-based technology can effectively compensate for brightness variations between frames and accurately compensate for the predicted value of the reference frame for the entire video image, its technical limitations, applied frame by frame, prevent it from adequately compensating for local brightness variations. To overcome this, the LIC technology induces appropriate brightness compensation for local predicted brightness variations.

[0142] The above LIC is a technique for modeling (760) local brightness changes between a template (715) adjacent to a current block (710) and a final prediction block (770) for a current block (710) by a function (750) that approximates between a value (730) based on a current block (710) and a value (740) based on a reference sample (720) between a template (715) adjacent to a reference block (720) and a reference block (720). The parameters of the function can be represented by a scale α and an offset β that form a linear equation, which can be expressed in the form of α*p[x]+β as a linear equation for compensating for brightness changes, where p[x] can mean a reference sample pointed to by a motion vector (MV) (717) at a location x on a reference picture in an inter prediction method. In one implementation of the above LIC, when wrap around motion compensation is enabled, the MV can be clipped considering an offset that takes into account the circular pixel space structure. Since α and β can be derived based on the current block template and the reference block template, no signaling overhead is required for them except when a LIC usage indication flag is signaled within the bitstream to indicate the usage of the LIC (preferably for Advanced Motion Vector Prediction (AMVP) mode).

[0143] However, the LIC technology presented in the present invention can be equally utilized for intra prediction, and for example, when a block vector (BV) for intra prediction is used instead of the MV (717) as a prediction vector for the current block, when the MV and / or the BV is derived from a decoder, when a chained vector is used that is determined by sequentially referencing the MV and / or the BV, or when any type of prediction vector is used by any other method, the present invention does not exclude the possibility that a method based on the LIC technology according to an embodiment of the present invention may be equally applied.

[0144] Among the applied implementation methods of the present invention, the implementation method citing the method proposed in JVET Technical Contribution JVET-O0066 can be modified as follows and used for blocks that are inter-coded between screens, for example, inter-coded units (inter-coded CUs).

[0145] According to one embodiment of the present invention, neighboring samples in the intra-dimensional space can be used to derive the LIC parameters. This allows for improved LIC accuracy by utilizing information from intra-encoded neighboring blocks.

[0146] According to one embodiment of the present invention, LIC can be disabled for blocks with fewer than 32 luminance samples. This reduces unnecessary computations and improves coding efficiency for small blocks.

[0147] According to one embodiment of the present invention, in non-subblock and affine modes, LIC parameter derivation can be performed based on template block samples corresponding to the current CU, rather than partial template block samples corresponding to the first 16x16 unit from the upper left. This allows for more information to be utilized, thereby improving the accuracy of LIC.

[0148] According to one embodiment of the present invention, samples of a reference block template can be generated using motion compensation (MC) along with the MV of the corresponding block without rounding to integer pixel precision. This enables more precise prediction, thereby improving coding efficiency.

[0149] According to one embodiment of the present invention, in the case of a bi-directional inter prediction CU, two sets of LIC parameters may be derived separately for L0 and L1 prediction samples. The L0 and L1 prediction samples may refer to respective samples obtained in two prediction directions used in the bi-directional prediction. For example, according to one embodiment, the L0 prediction sample may generally refer to a prediction sample derived from a reference picture that is temporally earlier than the current picture. That is, it may represent prediction information obtained from a past frame. Furthermore, according to one embodiment, the L1 prediction sample may generally refer to a prediction sample derived from a reference picture that is temporally later than the current picture. That is, it may represent prediction information obtained from a future frame.

[0150] According to one embodiment of the present invention, the L0 and L1 prediction samples may be obtained from different reference pictures and may be located at different temporal locations. As described above, by separately deriving LIC parameters using the L0 and L1 prediction samples, more accurate brightness compensation that takes temporal directionality into account can be provided.

[0151] According to one embodiment of the present invention, the same method can be repeatedly applied to derive the L0 and L1 LIC parameters. Specifically, the L0 LIC parameter is first derived by minimizing the difference between the L0 template prediction T0 and the template T, and the samples of T can be updated by subtracting the corresponding samples of T0. Next, the L1 parameter that minimizes the difference between the L1 template prediction T1 and the updated template can be calculated. Finally, the L0 parameter can also be refined again in the same manner.

[0152] According to another embodiment of the present invention, in the process of separately deriving LIC parameters for the L0 and L1 prediction samples, weights that take temporal characteristics of each direction into account can be applied. For example, if the L0 prediction sample is temporally closer to the current picture, a higher weight can be assigned to the L0 LIC parameter. Conversely, if the L1 prediction sample is closer, a higher weight can be assigned to the L1 LIC parameter. This allows for more effective utilization of temporal correlation. The present invention proposes a method and device for improving compression performance by appropriately performing brightness compensation of LIC.

[0153]

[0154] 3x3 cross-shaped filter

[0155] According to one embodiment of the present invention, when applying the LIC method, an L-shaped template is set from restored pixel values ​​on the top and left of the current block currently being encoded / decoded, and then coefficients that can be used in a cross-shaped filter having a size of 3x3 are generated and used from the template area. Fig. 8 is a conceptual diagram showing the shape and coefficient positions of a cross filter according to one embodiment of the present invention. Referring to Fig. 8, as described above, a cross-shaped filter (900) having a size of 3x3 can be expressed by the coefficient positions of A (left), B (right), C (upper), D (lower), and E (center).

[0156] According to various implementation methods of the present invention, the value filtered through the 3x3 cross-shaped filter can be expressed by one of the following mathematical formulas.

[0157]

[0158] The above mathematical expression 1 represents basic linear filtering, which can be calculated by multiplying the pixel value at each location by a pre-specified coefficient (C1 to C5) and adding them up.

[0159]

[0160] The above mathematical expression 2 represents the addition of a nonlinear term F to the basic linear filtering, and the F may mean that the square value of the central pixel E is adjusted according to the bit depth value (bitDepth) per pixel (in this case, "bitDepth_mid" may mean the middle value among the range values ​​for expressing the pixel value according to the bit depth per pixel). The ">>" operator symbol appearing in the above mathematical expression 2 may be understood as a right binary bit shift operation. Of course, depending on the embodiment, the elements constituting the nonlinear term F may change. For example, the nonlinear term F may be based on other pixel values ​​instead of E. For another example, the nonlinear term F may be defined by an equation of the third or higher degree. For another example, constant values ​​such as "bitDepth" and "bitDepth_mid" included in the nonlinear term and calculated may be replaced by other suitable constant values.

[0161]

[0162] The above mathematical expression 3 represents the addition of a constant term G to the above mathematical expression 2, where G may be "bitDepth_mid" and may represent a median value within the expression range of the pixel value. Of course, depending on the embodiment, the elements constituting the constant term F may vary. For example, the value of the constant term F may be replaced by another suitable constant value.

[0163] According to one embodiment of the present invention, the coefficients designated for the coefficients C1 to C7 shown in the above mathematical expressions 1 to 3 can be derived from an L-shaped template that can be derived from the restored pixel values ​​on the top and / or left of the current block. Of course, the various application methods described above can be applied in common, so that when using the template, only the entire designated template area or a portion of the template area can be used for deriving the coefficients. In addition, according to one embodiment of the present invention, the coefficient derivation method of a Wiener filter that utilizes the restored pixel values ​​on the top and / or left of the current block can be utilized in the same manner or in an applied manner for deriving the coefficients. The Wiener filter can have a characteristic of deriving coefficients in a manner that minimizes the minimum mean squared error (MSE), and the coefficients can be obtained based on this characteristic.

[0164] It is obvious that the present invention is not limited to the above-described embodiments, and various modifications and applications that can be applied in relation to LIC and predictive encoding within the scope of the present invention can be equally applied. For example, the size of the filter can be adjusted, and a type of filter other than a cross filter can be used. For example, a 5x5 cross filter or a 3x3 Sobel filter may be used. In addition, it is obvious that any filter configured to receive as input a pixel group having at least one shape of a square, a diamond, and a circle, each having a width of one pixel and a height of two pixels, can be applied. In addition, extensions such as dynamically adjusting the shape or size of the template, or introducing polygonal or nonlinear functions into the linear model to model more complex illumination changes can be made. Alternatively, various methods that explicitly or implicitly determine the prediction mode can be utilized to implement the examples described above with respect to the present invention, or the present invention can be implemented to operate in conjunction with such other methods.

[0165]

[0166] How to decide on a template

[0167] The above-described LIC-based method can be configured to find a reference block using a set template and generate a linear model using an L-shaped template for the change in brightness between a current block (e.g., CU) and the reference block. The linear model can be configured in the form of a linear function of α*p[x]+β. The p[x] can be defined as a reference value as a prediction pixel, and the α and β can represent the slope and bias of the derived linear model. According to research in the standardization activities of JVET, two or more linear models can be used simultaneously. According to one embodiment of the present invention, the method for deriving the linear model parameters is not limited to a first-order linear model, but can be set as a second-order or higher-order line fitting model. This can help to model more complex brightness changes more accurately.

[0168] According to one embodiment of the present invention, the LIC template setting method can be improved to provide various modes. That is, the L-shaped template may not be used uniformly. For example, a mode that uses only the top or left template may be used, or a mode that uses the top and left templates but uses each template with the same weight or different weights may be used. Figure 9 is an exemplary diagram showing an example of a template application mode according to one embodiment of the present invention. In summary, the following modes may be included and used in the LIC template setting process, but are not limited to these examples.

[0169] (a) Mode that uses only the top pixel (815) of the current block (810) as a template

[0170] (b) Mode that uses only the left pixel (827) of the current block (820) as a template

[0171] (c) A mode that uses the top (835) and left (837) pixels of the current block (830) as templates with the same weight (1:1).

[0172] (d) A mode that uses the top (845) and left (847) pixels of the current block (840) as templates with different weights (N:M).

[0173] According to one embodiment of the present invention, the mode of using only the upper pixel of the current block as a template may use at least one pixel selected based on pixels located within a distance of at least one pixel in the upper direction of the current block as a template. In addition, according to one embodiment of the present invention, the width of the template area corresponding to the pixel located at the upper end may be extended to the left or right by the number of unused left pixels. For example, if the size of the current block is 16x16 and the left pixel is not used for the template, the width of the upper template area may be extended from 16 pixels to 32 pixels.

[0174] According to one embodiment of the present invention, the mode of using only the left pixels of the current block as templates may use pixels located within a distance of at least one pixel in the left direction of the current block as templates. Furthermore, according to one embodiment of the present invention, the height of the template region corresponding to the pixels located on the left may be extended upwards or downwards by the number of unused upper pixels. For example, if the size of the current block is 16x16 and the upper pixels are not used for the template, the height of the left template region may be extended from 16 pixels to 32 pixels.

[0175] According to one embodiment of the present invention, when setting an LIC template, at least one of one or more modes including the above-described modes may be selected and used. In this case, at least one of the methods described below may be used as a method for selecting one of the above-described one or more modes.

[0176] According to one embodiment of the present invention, as a method for selecting the mode, an encoder may select a mode that exhibits optimal bit rate-to-distortion (RD) performance by using Rate-Distortion Optimization (RDO) as an evaluation criterion, and may inform a decoder of the corresponding mode information. According to one embodiment, assuming that the four types of template modes exemplified through FIG. 9 are supported, an RDO evaluation may be performed on each of the modes (a) to (d), and the mode exhibiting the lowest encoding cost and / or good RD performance may be selected as the LIC usage mode of the current block, and the corresponding information may be included in a bitstream and transmitted to the decoder. The decoder may then parse the bitstream and configure a template according to the selected LIC usage mode. Other LIC-based prediction block-related encoding and decoding processes may be performed according to the same or similar procedures as before.

[0177] According to one embodiment of the present invention, a method of utilizing a template may be used as a method for selecting the mode. According to one embodiment, the mode may be selected by an operation based on each L-shaped template (i.e., an upper pixel and a left pixel) composed of surrounding pixels of a current block and a reference block. According to one embodiment, an optimal template mode may be selected by utilizing an upper restoration value based on the L-shaped template of the current block and an upper value of the reference block (first mode), utilizing a left restoration value based on the L-shaped template of the current block and an upper value of the reference block (second mode), utilizing the upper and left restoration values ​​based on the L-shaped template of the current block and the upper and left values ​​of the reference block with the same weight (third mode, equal weight), or utilizing the upper and left restoration values ​​based on the L-shaped template of the current block and the upper and left values ​​of the reference block with different weights (fourth mode, different weight). When selecting the above optimal template mode, the SAD (Sum of Absolute Differences), SATD (Sum of Absolute Transformed Differences), or MR-SAD (Mean-removal SAD) values ​​and / or their average values, which are calculated by utilizing values ​​based on the templates (and / or templates of the templates) of the current block and the reference block, may be utilized as evaluation criteria. According to a preferred embodiment of the present invention, the above-described evaluation method may be configured to operate identically in the encoder and decoder based on a pixel area that has been previously encoded / decoded in the encoding / decoding process. Therefore, when applying this method, the template mode may be determined implicitly, so there is no need to record separate information in the bit string regarding which specific mode was used, and the decoder may be configured to determine the mode by itself.

[0178] According to one embodiment of the present invention, as a method for selecting the mode, the LIC mode can be explicitly determined by utilizing information of surrounding blocks. According to one embodiment, if there are available blocks on the top and left, the LIC mode of the current block can be explicitly determined by utilizing prediction information of the available surrounding blocks. According to one embodiment of the present invention, the explicit determination method can include a method of forming a candidate set and selecting at least one candidate belonging to the candidate set as the LIC mode. The method for forming the candidate set can be variously provided. For example, if there is a block using LIC around the current LIC block, the corresponding LIC mode can be included in the candidate set. In addition, usage history information on the LIC mode of a block previously used as the LIC mode can be stored and included in the candidate set. In addition, even if LIC is not used in the surrounding blocks, if at least one surrounding block is encoded by an intra-angular prediction mode, an appropriate LIC mode can be included in the candidate group in consideration of the angle of the directional prediction mode. For example, if the directional prediction mode has a horizontal angle, a mode that mainly refers to pixels in the horizontal direction, as shown in (b) of FIG. 9, can be included in the candidate group as an LIC mode more suitable for the direction. In addition, the candidate group can also be configured by considering information related to an inter-screen prediction method. When configuring the candidate group, according to one embodiment of the present invention, if a specific number of candidate group values ​​cannot be selected when configuring the candidate group, the modes of the candidate group can be arbitrarily added according to a predetermined order or information.According to one embodiment of the present invention, when the candidate group is formed, the template (or template of the template) (e.g., one line from the top, one line from the left, one line from the top and one line from the left, etc.) can be used to re-order the order of the candidate group. The re-ordering can be performed based on the encoding information required cost calculated according to the template matching result, and according to one embodiment, the re-ordering can be performed in order of low cost. According to one embodiment, the information required cost can be calculated by SAD, SATD, or MR-SAD, and the re-ordering can also be performed based on this cost.

[0179] Finally, once the above candidate group is formed, the optimal LIC mode can be selected through RDO, etc., and the index thereof can be included in a bit string and transmitted to the decoder. The decoder can be configured to form the candidate group in the same manner as the encoder, and to reconstruct the LIC mode used in the encoder by utilizing the information transmitted through the bit string.

[0180] According to one embodiment of the present invention, as a method for selecting the mode, the LIC mode can be implicitly determined by utilizing information of neighboring blocks. According to one embodiment, if there are available blocks on the top and left, the LIC mode of the current block can be implicitly determined by utilizing available neighboring prediction information. For example, when the current block is encoded by LIC, if there are neighboring blocks using LIC, the LIC mode of the current block can be determined based on the frequency of the LIC mode used in the neighboring blocks. In addition, even if LIC is not used in neighboring blocks, if at least one neighboring block is encoded by an intra-angular prediction mode, an appropriate LIC mode can be determined by considering the angle of the directional prediction mode. For example, if the directional prediction mode has a horizontal angle, a mode that mainly refers to pixels in the horizontal direction, as shown in FIG. 9 (b), can be determined as the LIC mode of the current block as a more appropriate LIC mode for the direction. According to a preferred embodiment of the present invention, the above-described LIC mode selection method can be configured to operate identically in the encoder and decoder based on the pixel area previously encoded / decoded during the encoding / decoding process. Therefore, when applying this method, there is no need to record separate information in the bit string regarding which specific mode was used, and the decoder can be configured to automatically determine the LIC mode.

[0181] According to one embodiment of the present invention, in the method for deriving the mode on the decoder side using the implicit decision method in the above-described embodiment, it is obvious that various decoder-side mode derivation methods known in the art or newly provided can be used in common or applied.

[0182] According to one embodiment of the present invention, in order to set up an LIC template, a mode may be defined in which both the top and the left are used with the same weight, or the top and the left are used with different weights. Various methods may be used to define the weights. According to one embodiment, the same weights as those used in a conventional LIC may be applied to generate α and / or β of the linear model. According to another embodiment, the top and left regions may be divided, and weights may be applied separately according to the regions to generate α and / or β. According to one embodiment of the present invention, the weights for each of the divided regions may be determined in advance through experiments, etc., and the determined weight values ​​may be used in a table, and an index designating a specific value of the table may be determined explicitly or implicitly and provided to a decoder.

[0183] According to one embodiment of the present invention, the table may be preferably specified in a form in which the sum of one weight set is a power of 2. This may have the purpose of easily replacing the division operation in the weighted sum operation based on the weight with a shift operation. For example, if the weight set is {3, 5}, the sum is 8 (2 3 ), the division operation required for the above weight set can be replaced with a 3-bit right binary bit shift operation.

[0184] According to one embodiment of the present invention, the weight value can be applied in a form that is determined on-the-fly using information from surrounding blocks. It will be readily understood that this method corresponds to a type of implicit signaling method. For example, if there are available blocks on the top and left, the weight value can be determined using information including inter-prediction information, intra-prediction information, and block size of the available surrounding blocks. Even in this case, similarly to the above, the weight value can be determined as a value that is a power of 2 for the purpose of easily replacing the division operation with a shift operation.

[0185] According to one embodiment of the present invention, in order to reduce the complexity of a method of selecting one of the above-described modes when setting an LIC template, a method of reducing the number of pixels used in the template by limiting the pixels to some pixels rather than all pixels within the template area may be applied. According to one embodiment, when generating α and β after finding a reference block, only some pixels rather than all may be used. According to one embodiment, when finding a reference block, all pixels of the template may be additionally used, and when generating α and β, only some pixels rather than all may be used. According to one embodiment, when finding a reference block, pixels of a template reduced to include only some pixels rather than all may be additionally used, and when generating α and β, all pixels may be used.

[0186] As described above, for the purpose of reducing complexity, methods for reducing the pixels used in the template to use only some, rather than all, of the pixels may include, but are not limited to, the following methods.

[0187] According to one embodiment of the present invention, when using the top, left, or top and left full templates, the number of pixels is reduced by utilizing a subsampling method corresponding to a multiple of 2, such as 2:1, 4:1, etc., and the reduced pixels can be used to find a reference block as described above or to generate α and β.

[0188] According to one embodiment of the present invention, when only the upper template is used, the number of pixels is reduced by horizontally dividing the area of ​​the entire upper template into N and selecting M pixels, and the reduced pixels can be used to find a reference block as described above or to generate α and β. According to an embodiment, the N and the M can be varied depending on the block size and the level of desired complexity. According to a preferred embodiment of the present invention, at least one of the N or M can be defined based on a power of 2 for ease of calculation.

[0189] For example, the entire upper template may be horizontally divided into four sections, and two pixels corresponding to the 1 / 4 and 3 / 4 sections may be selected, and the selected pixels may be used to find a reference block or to generate α and β. For another example, the entire upper template may be horizontally divided into eight sections, and three pixels corresponding to the 2 / 8, 4 / 8, and 6 / 8 sections may be selected, and the selected pixels may be used to find a reference block or to generate α and β. For another example, the entire upper template may be horizontally divided into eight sections, and four pixels corresponding to the 1 / 8, 3 / 8, 5 / 8, and 7 / 8 sections may be selected, and the selected pixels may be used to find a reference block or to generate α and β. For another example, when using only the upper template, the entire upper template area is divided into eight horizontal sections, and five pixels corresponding to sections 0 / 8, 2 / 8, 4 / 8, 6 / 8, and 8 / 8 are selected, and the selected pixels can be used to find a reference block or to generate α and β.

[0190] According to one embodiment of the present invention, when only the left template is used, the number of pixels is reduced by vertically dividing the area of ​​the entire left template into N and selecting M pixels, and the reduced pixels can be used to find a reference block as described above or to generate α and β. According to an embodiment, the N and the M can be varied depending on the block size and the level of desired complexity. According to a preferred embodiment of the present invention, at least one of the N or M can be defined based on a power of 2 for ease of calculation.

[0191] For example, the entire left template may be vertically divided into four sections, and two pixels corresponding to the 1 / 4 and 3 / 4 sections may be selected, and the selected pixels may be used to find a reference block or to generate α and β. For another example, the entire left template may be vertically divided into eight sections, and three pixels corresponding to the 2 / 8, 4 / 8, and 6 / 8 sections may be selected, and the selected pixels may be used to find a reference block or to generate α and β. For another example, the entire left template may be vertically divided into eight sections, and four pixels corresponding to the 1 / 8, 3 / 8, 5 / 8, and 7 / 8 sections may be selected, and the selected pixels may be used to find a reference block or to generate α and β. For another example, the entire left template area can be vertically divided into 8 parts, and 5 pixels corresponding to areas 0 / 8, 2 / 8, 4 / 8, 6 / 8, and 8 / 8 can be selected, and the selected pixels can be used to find a reference block or to generate α and β. In the description of the various methods described above, the entire upper template can mean an upper template that extends to the left or right from the width of the current block by the height of the current block. In addition, the entire left template can mean a left template that extends upwards or downwards from the height of the current block by the width of the current block.

[0192]

[0193] Encoder and decoder

[0194] It is obvious that the encoding method according to the present invention can be applied equally to an encoder and a decoder. The encoding method according to the present invention can be used as one of the methods for generating sample values ​​of a prediction block for performing intra- and / or inter-picture prediction in an encoder. When residual signal information is encoded by a difference value with respect to such a prediction block, the decoder can obtain decoded samples by generating sample values ​​of the same prediction block using a corresponding and / or symmetrical method and combining the difference value with such a prediction block. Additionally, as exemplified through the internal decoder (420) and the coding loop including the same in FIG. 4, this decoding process can be implemented in the same manner within the encoder to predict the state of the decoder.

[0195] The encoding method according to the present invention described above can be implemented through an encoder as a device. The encoder as a device can be implemented in a form that maintains the conventional encoder structure exemplified through FIGS. 1 to 6 above or applies a predetermined change thereto. However, the form of implementation is not necessarily limited to that exemplified, and any form of encoder structure that can function as a video encoder should be considered an encoder established by the present invention as long as it implements the technical idea of ​​the present invention.

[0196] In addition, the method for decoding the encoding result according to the present invention described above can be implemented through a decoder as a device. The decoder as a device can be implemented in a form that maintains the conventional decoder structure exemplified through FIGS. 1 to 6 above or applies a predetermined change thereto. However, the form of implementation is not necessarily limited to that exemplified, and any form of decoder structure that can function as a video decoder should be considered a decoder established according to the present invention as long as it implements the technical idea of ​​the present invention.

[0197] A person skilled in the art will readily understand that a bit string encoded by the above-described method and device can be decoded by applying a method symmetrical and / or reverse to the encoding method. In one embodiment, when reading information for decoding from the encoded bit string, at least one variable length coded phrase included in the encoded bit string can be interpreted, and furthermore, in one embodiment, the variable length coding can be performed by an entropy coding method. The technical details and application method of implementing such a decoding procedure can be readily understood from the above-described encoding procedure.

[0198] The encoder and / or decoder described herein may correspond to a device including a processor and a memory, each of which may be implemented as a computing device. The processor that may be included in the encoder and / or decoder described herein may mean one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding.

[0199] Even if the processor is expressed singularly for ease of understanding, those skilled in the art will appreciate that the processor may include multiple processing elements and / or multiple types of processing elements. For example, a device according to one embodiment of the present invention may include multiple processors or one processor and one controller as the processor. Furthermore, the processor may be implemented using various processing configurations, such as a parallel processor or a multi-core processor.

[0200] The processor may be configured to execute an operating system (OS) and one or more software programs running on the operating system. Furthermore, the processor may access, store, manipulate, process, and generate data in response to the execution of the software.

[0201] The software may include a computer program, code, instructions, or a combination of one or more of these, and may be configured to control the processor to perform a desired operation and to issue commands to the processor, either independently or collectively. The software may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for interpretation by the processor or for providing commands or data to the processor. The software may also be distributed over networked computer systems, and stored or executed in a distributed manner.

[0202] The software may also be implemented in the form of program commands that can be executed through various computer means and recorded or stored in the memory. The memory may be a computer-readable recording medium, and program commands, data files, data structures, etc. may be recorded singly or in combination in the computer-readable recording medium. The program commands stored in the memory may be based on a command system specifically designed and configured for the embodiment of the present invention, or may follow a command system known and available to those skilled in the art of computer software, such as a command system exemplified by the assembly language, C, C++, Java, Python, etc. It should be understood that the command system and the program commands therefrom include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by the device and / or the processor according to an embodiment of the present invention using an interpreter or the like.

[0203] The computer-readable recording medium constituting the device according to one embodiment of the present invention, including the memory described in this specification, may include a temporary or volatile recording medium that is maintained only while the processor is operating, such as a processor cache, a RAM, a flash memory, or a relatively non-volatile or long-term recording medium, such as a magnetic media such as a hard disk, a floppy disk, and a magnetic tape, an optical media such as a CD-ROM, a DVD, a magneto-optical media such as a floptical disk, or a solid state memory, or may include a read-only recording medium, such as a ROM arranged on hardware, and furthermore, the hardware itself configured to perform an operation equivalent to a series of program commands by a hard-wired structure by circuit wiring, and each step for performing the operation for implementing an embodiment of the present invention can be viewed as being recorded by the connection and arrangement of the hardware components, and thus the connection and arrangement method is the memory and It is obvious to a person skilled in the art that they can be considered equivalent.

[0204] The embodiments described above with respect to the processor and the memory are not mutually exclusive and may be selected or combined and implemented as needed. For example, a single hardware device may be configured to operate as a module composed of one or more of the software to perform the operations of an embodiment of the present invention, and vice versa. As another example, in the present specification, all or part of the operations assigned to a certain functional unit may be implemented by one or more of the software stored in a device according to an embodiment of the present invention (preferably, in any one of the recording media belonging to the category of the memory) and configured to be executed by the processor. In such a case, such a functional unit may be referred to as a functional unit "included" in the processor.

[0205]

[0206] Although the present invention has been described with reference to drawings and embodiments, as already mentioned above, it does not mean that the scope of protection of the present invention is limited to the drawings or embodiments presented above, and it will be understood that a person skilled in the relevant technical field can modify and change the present invention in various ways within a scope that does not depart from the spirit and scope of the present invention described in the claims of the present invention patent.

Claims

1. A method for decoding an encoded video bit stream, A step of obtaining at least one prediction vector for the current block currently being decoded from the above bit string, A step of determining at least one reference block based on the above prediction vector; A step of defining at least one of a first template including at least one pixel adjacent to the current block as a reference block, and a second template including at least one pixel adjacent to the reference block as a reference block; A step of extracting parameter information related to local illumination compensation (LIC) by referring to at least one pixel information derived from at least one of the first template and the second template; A step of determining at least one coefficient to be applied to the first filter based on the above parameter information; A step of performing filtering on the reference block using the first filter to which the above coefficients are applied, thereby obtaining a filtering result value; A step of generating a prediction sample for the current block based on the filtering result value; and A decryption method, comprising: a step of restoring the current block using the prediction sample; 2. In paragraph 1, The above first filter is, A decoding method, characterized in that the cross filter is configured to receive at least one left pixel, at least one right pixel, at least one top pixel, at least one bottom pixel, and a center pixel as inputs, and has at least one coefficient that can be applied to each of the pixels.

3. In paragraph 2, The above first filter is, A decoding method characterized by a 3x3 sized cross filter configured to receive as input five pixels including a left pixel, a right pixel, an upper pixel, a lower pixel, and a center pixel, and having five coefficients that can be applied to each of the pixels.

4. In paragraph 3, The above first filter is, A decoding method characterized in that it operates by a mathematical formula expressed as , wherein Pred is a filtering result value of the first filter, A represents a left pixel, B represents a right pixel, C represents an upper pixel, D represents a lower pixel, E represents an input value of a center pixel, and C1, C2, C3, C4, and C5 represent coefficients applied to each of the inputs.

5. In paragraph 3, The above first filter is, It operates by a mathematical formula expressed as, Pred is a filtering result value of the first filter, A represents a left pixel, B represents a right pixel, C represents an upper pixel, D represents a lower pixel, E represents an input value of a central pixel, F represents a nonlinear term based on E, and V mid is the median constant, V above div A decryption method, characterized in that C1, C2, C3, C4, C5, and C6 represent coefficients applied to each of the inputs, and is a leveling division constant.

6. In paragraph 5, Above V mid represents the median value of the bit depth for the pixels of the current block ("bitDepth_mid"), and the V div A decryption method, characterized in that division by is replaced by a right binary bit shift operation by the bit depth value (">>bitDepth").

7. In paragraph 3, The above first filter is, It operates by a mathematical formula expressed as, Pred is a filtering result value of the first filter, A represents a left pixel, B represents a right pixel, C represents an upper pixel, D represents a lower pixel, E represents an input value of a central pixel, F represents a nonlinear term based on E, and V mid is the median constant, V above div is the equalization division constant, V above const A decryption method, characterized in that C1, C2, C3, C4, C5, C6, and C7 represent coefficients applied to each of the inputs, and is a batch addition constant.

8. In paragraph 7, Above V mid and the above V const represents the median value of the bit depth for the pixels of the current block ("bitDepth_mid"), and the V div A decryption method, characterized in that division by is replaced by a right binary bit shift operation by the bit depth value (">>bitDepth").

9. In paragraph 1, The above first filter is, A decoding method, characterized in that the image processing filter is configured to receive as input a group of pixels having at least one shape among a square, a diamond, and a circle, each of which has a width equal to a first pixel and a height equal to a second pixel, centered on a central pixel, and has at least one coefficient that can be applied to each pixel.

10. In paragraph 1, The step of determining at least one coefficient to be applied to the first filter comprises: A decoding method characterized in that it is performed by a minimum mean squared error (MSE) minimization method based on a Wiener filter, which is calculated based on at least one of the restored pixel values ​​on the top and left of the current block.

11. A method for encoding a bit stream generated from a video, A step of determining at least one prediction vector for the current block currently being encoded, A step of determining at least one reference block based on the above prediction vector; A step of defining at least one of a first template including at least one pixel adjacent to the current block as a reference block, and a second template including at least one pixel adjacent to the reference block as a reference block; A step of extracting parameter information related to local illumination compensation (LIC) by referring to at least one pixel information derived from at least one of the first template and the second template; A step of determining at least one coefficient to be applied to the first filter based on the above parameter information; A step of performing filtering on the reference block using the first filter to which the above coefficients are applied, thereby obtaining a filtering result value; A step of generating a prediction sample for the current block based on the filtering result value; a step of encoding the current block using the above prediction sample; and An encoding method, comprising: a step of recording the encoding result in the bit string; 12. In paragraph 11, The above first filter is, An encoding method, characterized in that the cross filter is configured to receive at least one left pixel, at least one right pixel, at least one top pixel, at least one bottom pixel, and a center pixel as inputs, and has at least one coefficient that can be applied to each of the pixels.

13. In paragraph 12, The above first filter is, An encoding method characterized by a 3x3 sized cross filter configured to receive as input five pixels including a left pixel, a right pixel, an upper pixel, a lower pixel, and a center pixel, and having five coefficients that can be applied to each of the pixels.

14. In paragraph 13, The above first filter is, An encoding method characterized in that it operates by a mathematical formula expressed as , wherein Pred is a filtering result value of the first filter, A represents a left pixel, B represents a right pixel, C represents an upper pixel, D represents a lower pixel, E represents an input value of a center pixel, and C1, C2, C3, C4, and C5 represent coefficients applied to each of the inputs.

15. In paragraph 13, The above first filter is, It operates by a mathematical formula expressed as, Pred is a filtering result value of the first filter, A represents a left pixel, B represents a right pixel, C represents an upper pixel, D represents a lower pixel, E represents an input value of a central pixel, F represents a nonlinear term based on E, and V mid is the median constant, V above div An encoding method, characterized in that C1, C2, C3, C4, C5, and C6 represent coefficients applied to each of the inputs, and is a leveling division constant.

16. In paragraph 15, Above V mid represents the median value of the bit depth for the pixels of the current block ("bitDepth_mid"), and the V div An encoding method characterized in that division by is replaced by a right binary bit shift operation by a bit depth value (">>bitDepth").

17. In paragraph 13, The above first filter is, It operates by a mathematical formula expressed as, Pred is a filtering result value of the first filter, A represents a left pixel, B represents a right pixel, C represents an upper pixel, D represents a lower pixel, E represents an input value of a central pixel, F represents a nonlinear term based on E, and V mid is the median constant, V above div is the equalization division constant, V above const An encoding method, characterized in that C1, C2, C3, C4, C5, and C6 represent coefficients applied to each of the inputs, and is a batch addition constant.

18. In paragraph 17, Above V mid and the above V const represents the median value of the bit depth for the pixels of the current block ("bitDepth_mid"), and the V div An encoding method characterized in that division by is replaced by a right binary bit shift operation by a bit depth value (">>bitDepth").

19. In paragraph 11, The step of determining at least one coefficient to be applied to the first filter comprises: A decoding method characterized in that it is performed by a minimum mean squared error (MSE) minimization method based on a Wiener filter, which is calculated based on at least one of the restored pixel values ​​on the top and left of the current block.

20. A decoder device configured to decode a video bit stream encoded by a computing device, A receiver that receives a bit string; An output section that outputs the decrypted video; A reference buffer that stores information about at least one decrypted picture; processor; and a memory for storing instructions executable by the processor; A decoder device, comprising: instructions for obtaining at least one prediction vector for a current block currently being decoded from the bit string; determining at least one reference block based on the prediction vector; defining at least one of a first template including at least one pixel adjacent to the current block as a reference block, and a second template including at least one pixel adjacent to the reference block as a reference block; extracting parameter information related to local illumination compensation (LIC) with reference to at least one pixel information derived from at least one of the first template and the second template; determining at least one coefficient to be applied to a first filter based on the parameter information; performing filtering on the reference block using the first filter to which the coefficients are applied, thereby obtaining a filtering result value; generating a prediction sample for the current block based on the filtering result value; and reconstructing the current block using the prediction sample.