Adaptive decoder-side intra prediction mode derivation method and device for video encoding and decoding

The adaptive decoder-side intra-prediction method optimizes video encoding and decoding by combining gradient analysis and texture complexity to select efficient intra prediction modes, enhancing compression performance and reducing computational complexity.

WO2026089472A1PCT designated stage Publication Date: 2026-04-30KAON GRP CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
KAON GRP CO LTD
Filing Date
2025-10-22
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing decoder-side intra prediction methods in video encoding and decoding are inefficient and do not fully leverage the potential of combining different intra-mode derivation techniques, leading to suboptimal compression performance and increased computational complexity.

Method used

An adaptive decoder-side intra-prediction mode estimation method that utilizes decoded pixel values from adjacent blocks, calculates evaluation values for candidate modes using gradient analysis and texture complexity, and selects the most efficient mode for intra prediction, incorporating techniques like Sobel filtering and multiple reference lines.

Benefits of technology

Improves encoding and decoding efficiency, reduces computational load, and enhances video quality by minimizing bit sequence waste and optimizing mode selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025016806_30042026_PF_FP_ABST
    Figure KR2025016806_30042026_PF_FP_ABST
Patent Text Reader

Abstract

An image decoding method according to one embodiment of the present invention comprises the steps of: referring to decoded pixel values of a region adjacent to the current block in order to perform intra prediction for the current block; calculating, on the basis of the decoded pixel values, evaluation values for a plurality of candidate intra prediction modes to be applied to intra prediction of the current block; selecting at least one intra prediction mode from among the plurality of candidate intra prediction modes on the basis of the evaluation values; and using the selected intra prediction mode so as to decode the current block, wherein the plurality of candidate intra prediction modes include at least one directional intra prediction mode, and the step of calculating the evaluation values is performed on the basis of at least one from among intra prediction information of at least one neighboring block spatially adjacent to the current block, a filter operation result for the decoded pixel values, and a cost calculation result between the decoded pixel values and prediction values generated by the candidate intra prediction mode.
Need to check novelty before this filing date? Find Prior Art

Description

Adaptive decoder-side intra-prediction mode estimation method and apparatus for video encoding and decoding

[0001] The present invention relates to video compression technology, and specifically to an adaptive and improved application method and utilization of decoder-side intra mode derivation (DIMD) technology, which is one of the intra prediction application methods that contribute to improving compression performance in video encoders and decoders.

[0002] The present invention may correspond to a technical field identical to at least one of the digital video compression technology standards known by standard names such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG, a technical field for improving the inherent efficiency of the standard, or a technical field for improving or replacing the standard.

[0003] Digital video encoding and decoding are widely utilized in various digital video applications. For example, devices such as video recording equipment and camcorders used for video recording activities—including digital television broadcasting, video transmission via communication networks, video calls, video conversations, and video chats, recording and provision of video content using optical media such as VCDs (video compact discs), DVDs (digital versatile discs), and Blu-rays, all procedures for the production, editing, collection, and distribution of video content, and video recording for various purposes including personal, commercial, industrial, and security purposes—are all dependent on video encoding and decoding technology.

[0004] Accordingly, embodiments that can be referred to as digital video encoders and decoders may constitute a part of a wide range of devices related to the creation, recording, and provision of digital video, including digital television, digital broadcasting systems, wireless broadcasting systems, computers in the form of notebooks / desktops / tablets, e-book readers, digital cameras, digital recording devices, digital multimedia playback devices, video game devices / terminals / consoles, mobile phones equipped with multimedia playback functions (including smartphones), equipment for video conferencing, and other devices.

[0005] Digital video encoders and decoders as described above can be implemented by digital video compression standards that are understood by and widely used by people skilled in the art. The digital video compression standards may include at least one of the compression standards known by standard names such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.

[0006] Video encoders and decoders can be implemented to encode or decode digital video information more efficiently while complying with the above specifications, or by improving or modifying them. Attempts to modify the above specifications may also lead to the development of new specifications. Among well-known examples is the so-called enhanced compression model (ECM), which is an attempt to improve and replace the conventional H.266 / VVC specifications, currently being developed by the Joint Video Experts Team (JVET), a joint international standardization group of ISO, IEC, and ITU-T.

[0007] Among the various detailed technologies applied to conventional video encoders and decoders, there exists a technology collectively referred to as decoder-side intra mode derivation (DIMD). In the process of increasing the compression ratio of an image through intra-frame prediction, i.e., intra prediction, DIMD proposes a method for calculating the intra prediction method of the current block and / or the value of the intra prediction block from the values ​​of a conventionally decoded image adjacent to the currently encoded / decoded block in order to reduce the amount of information transmitted from the encoder to the decoder regarding the intra prediction mode. Consequently, even without the encoder providing a separate signal, the decoder in DIMD mode can autonomously derive information related to intra prediction from the preceding decoding process, thereby resulting in an improvement in compression efficiency.

[0008] However, since the DIMD that serves as the background technology of the present invention has been implemented or applied by various prior art methods for estimating intra prediction modes on the decoder side, it may include technology that is not necessarily called by the characteristic name DIMD, and may be related to various application technology areas not limited thereto, such as methods for calculating the intra prediction of the current block and / or the value of the intra prediction block from the values ​​of a conventionally decoded image, for example, methods based on a list of most probable modes (MPM), template-based intra mode derivation (TIMD), occurrence-based intra coding (OBIC), and others.

[0009] The present invention aims to provide a more efficient intra-mode derivation method by combining and combining the advantages of various intra-mode derivation methods, including decoder-side intra-mode derivation (DIMD), occurrence-based intra-coding (OBIC), template-based intra-mode derivation (TIMD), template-based multiple reference line (TMRL), and others not limited thereto.

[0010] The present invention proposes an implementation method for an adaptive decoder-side intra-prediction mode estimation method and a device to solve the aforementioned technical problem.

[0011] A video decoding method according to an embodiment of the present invention for solving the aforementioned technical problem comprises: a step of referencing decoded pixel values ​​of an area adjacent to the current block to perform intra prediction for the current block; a step of calculating an evaluation value for a plurality of candidate intra prediction modes to be applied to the intra prediction of the current block based on the decoded pixel values; a step of selecting at least one intra prediction mode among the plurality of candidate intra prediction modes based on the evaluation value; and a step of decoding the current block using the selected intra prediction mode, wherein the plurality of candidate intra prediction modes includes at least one directional intra prediction mode, and the step of calculating the evaluation value is characterized by being performed based on at least one of intra prediction information of at least one surrounding block spatially adjacent to the current block, a filter operation result for the decoded pixel values, and a cost calculation result between the decoded pixel values ​​and the prediction value by the candidate intra prediction mode.

[0012] The step of calculating the above evaluation value may be characterized by including the step of deriving a first evaluation value based on the intra-prediction mode of the surrounding block and the size of the surrounding block, the step of deriving a second evaluation value based on the filter operation result, and the step of merging the first evaluation value and the second evaluation value to calculate a final evaluation value.

[0013] The step of deriving the first evaluation value may be characterized by adding a value proportional to the number of pixels of the surrounding block to the evaluation value corresponding to the directional intra prediction mode when the surrounding block is encoded in a directional intra prediction mode.

[0014] The above method may be characterized in that the filter operation includes a gradient analysis of the decoded pixel values, and the gradient analysis calculates the horizontal gradient and the vertical gradient of the decoded pixel values, derives a gradient angle from the horizontal gradient and the vertical gradient, and increases the evaluation value of a directional intra-prediction mode corresponding to the gradient angle.

[0015] The above gradient analysis may be characterized by being performed using a Sobel filter.

[0016] The decoded pixel values ​​to which the above gradient analysis is applied may be characterized by including pixel values ​​on a plurality of reference lines located at different distances from the current block.

[0017] The size of the filter used in the above slope analysis may be characterized by being determined based on the number of reference lines or the maximum distance between the reference lines.

[0018] The above plurality of candidate intra prediction modes may be characterized by further including at least one non-directional intra prediction mode.

[0019] The above method further includes the step of calculating texture complexity for the decoded pixel values, and the step of calculating the evaluation value may be characterized by calculating an evaluation value for the non-directional intra-prediction mode based on the texture complexity.

[0020] The step of calculating the above evaluation value may be characterized by selecting the non-directional intra-prediction mode without performing the filter operation and the cost calculation when the texture complexity is below a preset threshold.

[0021] A video encoding method according to an embodiment of the present invention for solving the aforementioned technical problem comprises: a step of referencing pixel values ​​reconstructed in a prior encoding process as an area adjacent to the current block to perform intra prediction for the current block; a step of calculating an evaluation value for a plurality of candidate intra prediction modes to be applied to the intra prediction of the current block based on the reconstructed pixel values; a step of selecting at least one intra prediction mode among the plurality of candidate intra prediction modes based on the evaluation value; and a step of encoding the current block using the selected intra prediction mode, wherein the plurality of candidate intra prediction modes includes at least one directional intra prediction mode, and the step of calculating the evaluation value is characterized by being performed based on at least one of intra prediction information of at least one surrounding block spatially adjacent to the current block, a filter operation result for the reconstructed pixel values, and a cost calculation result between the reconstructed pixel values ​​and the prediction value by the candidate intra prediction mode.

[0022] The step of calculating the above evaluation value may be characterized by including the step of deriving a first evaluation value based on the intra-prediction mode of the surrounding block and the size of the surrounding block, the step of deriving a second evaluation value based on the filter operation result, and the step of merging the first evaluation value and the second evaluation value to calculate a final evaluation value.

[0023] The above filter operation may be characterized by being performed on reconstructed pixel values ​​on a plurality of reference lines located at different distances from the current block.

[0024] The above plurality of candidate intra prediction modes may be characterized by further including at least one non-directional intra prediction mode.

[0025] The above method may further include the step of calculating texture complexity for the reconstructed pixel values, and the step of selecting the non-directional intra-prediction mode while omitting the step of calculating the evaluation value when the texture complexity is less than a first threshold value.

[0026] The above method may further include a step of determining whether the selected intra prediction mode can be automatically induced at the decoder side, and a step of determining whether to include the selected intra prediction mode information in the bitstream based on the result of the determination.

[0027] An image decoder device according to an embodiment of the present invention for solving the aforementioned technical problem comprises: a processor; a memory connected to the processor; a reference unit that references decoded pixel values ​​of an area adjacent to the current block to perform intra prediction for the current block; a calculation unit that calculates an evaluation value for a plurality of candidate intra prediction modes to be applied to the intra prediction of the current block based on the decoded pixel values; a selection unit that selects at least one intra prediction mode among the plurality of candidate intra prediction modes based on the evaluation value; and a decoding unit that decodes the current block using the selected intra prediction mode, wherein the plurality of candidate intra prediction modes includes at least one directional intra prediction mode, and the calculation unit calculates the evaluation value based on at least one of intra prediction information of at least one surrounding block spatially adjacent to the current block, a filter operation result for the decoded pixel values, and a cost calculation result between the decoded pixel values ​​and the prediction value by the candidate intra prediction mode.

[0028] The above-described output unit may be characterized by deriving a first evaluation value based on the intra-prediction mode of the surrounding block and the size of the surrounding block, deriving a second evaluation value based on the filter operation result, and merging the first evaluation value and the second evaluation value to calculate a final evaluation value.

[0029] The above memory may be characterized by storing an evaluation value calculated by the above calculation unit and intra-prediction mode information selected by the above selection unit, and the above calculation unit may refer to the information stored in the above memory when calculating the evaluation value of a subsequent block.

[0030] The filter operation performed by the above-mentioned output unit may be characterized by being performed on decoded pixel values ​​on a plurality of reference lines located at different distances from the current block.

[0031] According to the present invention, at least one of the following effects can be achieved in video encoding and decoding: improvement of encoding efficiency, improvement of decoding efficiency, improvement of video quality, reduction of computational load, reduction of software size, reduction of hardware size, and other improvements in performance related to encoding and decoding.

[0032] According to the present invention, by using different intra-mode derivation methods proposed in the past in a mutually compromised manner, effects such as minimizing the waste of bit sequence information for mode selection and reducing the complexity of encoder / decoder processing can be achieved.

[0033] FIG. 1 is a conceptual diagram of a video communication system according to an embodiment of the present invention,

[0034] FIG. 2 is a conceptual diagram of the arrangement of an encoder and a decoder in a real-time video streaming environment according to an embodiment of the present invention.

[0035] FIG. 3 is a conceptual diagram of a functional unit of a video decoder according to an embodiment of the present invention,

[0036] FIG. 4 is a conceptual diagram of a functional unit of a video encoder according to an embodiment of the present invention,

[0037] FIG. 5 is a conceptual diagram of a frame type according to an embodiment of the present invention,

[0038] Figure 6 is a conceptual diagram showing the structure of a video encoder according to the H.266 / VVC standard.

[0039] FIG. 7 is a conceptual diagram of the selection of an adjacent pixel set according to an embodiment of the present invention,

[0040] FIG. 8 is a conceptual diagram of slope calculation for each pixel in a template according to an embodiment of the present invention,

[0041] FIG. 9 is a conceptual diagram of the application of a reduced frequency Sobel filter according to one embodiment of the present invention,

[0042] FIG. 10 is a conceptual diagram showing a predictive fusion algorithm of DIMD according to an embodiment of the present invention,

[0043] FIG. 11 is a conceptual diagram showing a predictive fusion algorithm of DIMD according to another embodiment of the present invention,

[0044] FIG. 12 is a conceptual diagram of a combined DIMD according to an embodiment of the present invention,

[0045] FIG. 13 is a conceptual diagram illustrating a method for the combined generation of histograms according to an embodiment of the present invention,

[0046] FIG. 14 is a conceptual diagram illustrating a method for applying a variable filter size based on multiple reference lines according to an embodiment of the present invention.

[0047] FIG. 15 is a conceptual diagram illustrating a histogram generation method including a non-directional intra-prediction mode according to an embodiment of the present invention,

[0048] FIG. 16 is a conceptual diagram showing a memory reuse structure according to an embodiment of the present invention, and

[0049] FIG. 17 is a flowchart illustrating an IBC priority processing procedure according to one embodiment of the present invention.

[0050] The present invention is capable of various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the invention to specific embodiments, and it should be understood that the invention includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the invention.

[0051] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. These terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of multiple related described items or any one of the multiple related described items, and is non-exclusive unless otherwise indicated. When items are listed in this application, they are merely illustrative descriptions intended to facilitate the explanation of the spirit of the present invention and possible methods of implementation, and are therefore not intended to limit the scope of the embodiments of the present invention.

[0052] In this specification, "A or B" may mean "only A," "only B," or "both A and B." Alternatively, in this specification, "A or B" may be interpreted as "A and / or B." For example, in this specification, "A, B or C" may mean "only A," "only B," "only C," or "any combination of A, B and C."

[0053] A slash ( / ) or a comma used in this specification may mean "and / or." For example, "A / B" may mean "A and / or B." Accordingly, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B or C."

[0054] In this specification, "at least one of A and B" may mean "only A," "only B," or "both A and B." Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted as synonymous with "at least one of A and B."

[0055] Additionally, in this specification, "at least one of A, B and C" may mean "only A," "only B," "only C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."

[0056] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0057] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "having" are intended to specify the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0058] Unless otherwise defined, all terms used herein, including technical or scientific terms, are used with the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.

[0059] In describing the invention in this application, embodiments may be described or illustrated in terms of the described functions or unit blocks that perform the functions. The blocks may be expressed in this application as one or more devices, units, modules, parts, etc. The blocks may be implemented in hardware by a method of implementing one or more logic gates, integrated circuits, processors, controllers, memory, electronic components, or information processing hardware, which are not limited thereto. Alternatively, the blocks may be implemented in software by a method of implementing application software, operating system software, firmware, or information processing software, which are not limited thereto. A single block may be implemented by being separated into multiple blocks that perform the same function, or conversely, a single block may be implemented to perform the functions of multiple blocks simultaneously. The blocks may also be implemented by being physically separated or combined according to any criteria. The blocks may be implemented to operate in an environment where their physical locations are not specified and they are spaced apart from each other by a communication network, the Internet, a cloud service, or a communication method not limited thereto. Since all of the above-mentioned methods of implementation fall within the scope of various embodiments that a person skilled in the art familiar with the field of information and communication technology can adopt to realize the same technical concept, any detailed methods of implementation should be interpreted as being included within the scope of the technical concept of the invention of this application.

[0060] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the attached drawings. In describing the present invention, to facilitate overall understanding, the same reference numerals are used for identical components in the drawings, and redundant descriptions of identical components are omitted. Furthermore, it is assumed that multiple embodiments are not mutually exclusive and that some embodiments may be combined with one or more other embodiments to form new embodiments.

[0061]

[0062] digital video codecs

[0063] FIG. 1 is a conceptual diagram of a video communication system according to an embodiment of the present invention. The video communication system (100) may be configured to include at least two terminals (110, 120) connected to each other through a network (105).

[0064] In one embodiment of the present invention, FIG. 1 may represent a block diagram for configuring a unidirectional video communication network. A first terminal (110) among the terminals may encode the video data in order to transmit (111) the video data through a network (105). A second terminal (120) among the terminals may be configured to receive (121) the encoded video data through a network, decode it, and display it.

[0065] In another embodiment of the present invention, FIG. 1 may represent a block diagram for configuring a bidirectional video communication network. For the bidirectional video communication, each terminal (110, 120) may be configured to encode video data acquired by itself for video transmission (112, 122) to each other terminal passing through the network. Each terminal may also receive (113, 123) video data transmitted through the network by another terminal, decode it, and be configured to display the decoded video data.

[0066] Each terminal (110, 120) shown in FIG. 1 may be exemplified as a device such as a server computer, a personal computer, a portable computer, and a smartphone, depending on the embodiment, but is not limited thereto and may be any device corresponding to a commonly used computing device. For example, according to an embodiment of the present invention, each terminal (110, 120) may mean a desktop computer, a laptop computer, a tablet PC, a mobile phone, a smartphone, a PDA (personal digital assistant), a workstation, an electronic calculator, a server computer, a cloud computer, a virtualization computer, a quantum computer, or any other electronic, electrical, or quantum computing device implemented in a movable or inmovable form. In particular, among such devices, it may be interpreted as any device that is designed to operate as a terminal device according to an embodiment of the present invention, is capable of operating as a terminal device according to an embodiment of the present invention, or is capable of installing and / or executing a computer program that enables it to operate as a terminal device according to an embodiment of the present invention or perform a method corresponding to such operation.

[0067] Each of the above terminals (110, 120) may be implemented by a plurality of functional units configured to exchange information within each of the above terminals (110, 120) by being interconnected in various forms, such as a bus, a circuit, or a relationship between a routine and a subroutine. Additionally, through the interconnection, for the purpose of executing or supporting the operation of a functional unit among the above functional units that primarily requires the performance of calculations, the terminals may be configured to include a processor (130) having a computational function and a memory (140) connected to the processor.

[0068] The processor (130) described in this specification may mean one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions.

[0069] Even if the processor (130) is expressed in the singular for ease of understanding, a person of ordinary knowledge in the relevant technical field will know that the processor (130) may include a plurality of processing elements and / or a plurality of types of processing elements. For example, a device according to one embodiment of the present invention may include a plurality of processors or one processor and one controller as the processor (130). In addition, the processor (130) may be implemented by various processing configurations, such as a parallel processor or a multi-core processor.

[0070] The processor (130) may be configured to execute an operating system (OS) and one or more software executed on the operating system. Additionally, the processor may access, store, manipulate, process, and generate data in response to the execution of the software. The software may include a computer program, code, instructions, or a combination of one or more of these, and may configure the processing unit to operate as desired or command the processing unit independently or collectively. The software may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave to be interpreted by the processor (130) or to provide instructions or data to the processor. The software may be distributed among a plurality of computer systems connected to the network (105), such as the terminals (110, 120), and may be stored or executed in a distributed manner.

[0071] The software may also be implemented in the form of program instructions that can be executed through various computer means and may be recorded or stored in the memory (140). The memory (140) may be a computer-readable recording medium, and program instructions, data files, data structures, etc. may be recorded in the computer-readable recording medium alone or in combination. The program instructions stored in the memory (140) may be based on a command system specifically designed and configured for an embodiment of the present invention, or may follow a command system known and available to those skilled in the art of computer software, such as Assembly, C, C++, Java, Python, etc. It should be understood that the command system and the program instructions thereunder include not only machine code such as that produced by a compiler, but also high-level language code that can be executed by the device and / or the processor (130) according to an embodiment of the present invention using an interpreter, etc.

[0072] A computer-readable recording medium constituting an apparatus according to an embodiment of the present invention, including the memory (140) described in this specification, may include a temporary or volatile recording medium that is maintained only while the processor is operating, such as a processor cache, RAM, or flash memory; or may include a relatively non-volatile or long-term recording medium such as a magnetic media such as a hard disk, floppy disk, and magnetic tape; an optical recording medium such as a CD-ROM or DVD; a magneto-optical media such as a floptical disk; or a solid state memory; or may include a read-only recording medium such as a ROM placed on hardware; furthermore, the hardware itself configured to perform a series of program instructions and equivalent operations by means of a hard-wired structure by circuit wiring, and since each step for performing the operation implementing the embodiment of the present invention can be considered to be recorded by the connection and arrangement of the hardware components, the method of connection and arrangement is the same as the It is obvious to a person skilled in the art that it can be seen as equivalent to memory (140).

[0073] The embodiments described above with respect to the processor (130) and the memory (140) are not mutually exclusive and may be selected or combined as needed. For example, a hardware device may be configured to operate as a module composed of one or more of the software to perform the operation of the embodiment of the present invention, and vice versa. As another example, in this specification, all or part of the operation assigned to a function may be implemented by one or more of the software stored in a device according to an embodiment of the present invention (preferably in a recording medium falling within the category of the memory) and configured to be executed by the processor, in which case such a function may be referred to as a function "included" in the processor.

[0074] The present invention is applicable to any environment for establishing a unidirectional or bidirectional video communication network, and the network (105) should be understood as being able to be established by any means for carrying encoded video data between the terminals (110, 120).

[0075] In one embodiment of the present invention, the network (105) may refer to a wired or wireless communication network. In this case, depending on the embodiment, the network may be configured to communicate information using any communication standard, and the communication standard may include packet-based communication. The packet communication may be understood to mean, for example, packets known as TCP or UDP. The wired communication method of the network (105) may be a method of connecting to an external communication network via an RJ-11 standard telephone line, an Ethernet cable belonging to various categories of the RJ-45 standard, other coaxial cables, metal cables, optical cables, and other various wired media. According to the embodiment, the wireless communication method of the above network (105) may include a short-range wireless communication method including Bluetooth, Wi-Fi, Zigbee, and NFC (near field communication), or a long-range wireless communication method that can be referred to by the name of a wireless communication technology commonly called by the names of communication standard generations such as Wibro, WiMax, Global Systems for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long-Term Evolution (LTE), New Radio (NR), and other 2G, 3G, 4G, 5G, 6G, and other international standard wireless communication specifications such as IMT-2000, IMT-Advanced, IMT-2020, and IMT-2030. Of course, even if any other conventional or newly developed wired or wireless communication means, method, standard, and protocol are applied to the implementation of the network (105), as long as they are means configured to perform transmission and reception in information communication devices such as terminals (110, 120), there is no impediment to achieving the purpose of the present invention.It is also obvious that a network (105) can be configured in combination with one or more wired and / or wireless standards.

[0076] However, in another embodiment of the present invention, the network (105) may be understood to mean a process of information transmission using a computer-readable recording medium. In this case, the configuration of the network is not limited to communication media only, but should be understood to include a process of temporarily storing and physically transporting information in computer-readable memory and / or recording media. The computer-readable recording medium used for information transmission may be understood to mean a recording medium that is relatively non-volatile or capable of long-term recording, such as magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, or solid-state memory, which is primarily used as a means of transporting data between computing devices.

[0077] Any other means of information communication or transport applied may be considered to fall within the scope of embodiments of the present invention as long as it is structured to support decoding by transmitting video data in an encoded state. Accordingly, in addition to the examples listed above, all means of information communication or transport known in the prior art or newly provided may fall within the scope of application of the present invention.

[0078] FIG. 2 is a conceptual diagram of the arrangement of encoders and decoders in a real-time video streaming environment according to an embodiment of the present invention. The streaming system (200) exemplified by FIG. 2 can be seen as applicable to a video data communication network including, for example, digital broadcasting, video telephone, and video conferencing. However, it should be seen that a technical structure identical or similar to the streaming system can be equally applied even when information transport via a recording medium is involved, as described above.

[0079] According to one embodiment of the present invention, the streaming system may include a video source (210) that generates a video stream. The video source may include a digital video acquisition means (212) that acquires uncompressed raw video, which may be composed of, for example, a digital camera or other equipment. The raw video stream (215) may have a massive capacity and may therefore be compressed by a video encoder (217) coupled to or connected to the video source.

[0080] The encoder (217) may be composed of means including hardware, software, or a combination of both, configured to implement an image encoding method and / or a method of implementing the same according to an embodiment of the present invention.

[0081] By passing through the encoder (217), an encoded bitstream (219) with a reduced capacity compared to the original video stream can be output. The bitstream (219) can be provided in real-time via communication by a relay device, for example, which may be referred to as a streaming server (220), and / or can be stored in a recording medium (225) of the streaming server (220) for subsequent use.

[0082] The streaming system (200) may include at least one streaming client (230, 240) that connects to the streaming server (220) to receive the encoded bit sequence (229) in real time or to acquire it subsequently. The streaming client may include a video decoder (232) that acquires the encoded bit sequence (229) (which may also be considered as a copy of the bit sequence (219) received by the streaming server), decodes the bit sequence (229), and outputs the resulting video data as video data in a form that can be displayed on a display (235) or other visual, auditory, or other sensory display means.

[0083] As mentioned above, the functions for encoding and decoding video data are collectively referred to as the coder-and-decoder system, or video codec.

[0084] FIG. 3 is a conceptual diagram of a functional unit of a video decoder according to an embodiment of the present invention. As shown in FIG. 3, a receiver (310) can receive at least one encoded video data to be decoded by a decoder (305). In an embodiment of the present invention, the encoded video data may be independent for each reception, and the decoding procedure of each independent video data may be independent from the decoding procedure of other video data. The encoded video data may be received by the receiver (310) through a hardware or software connection (315) to a device that stores the data, and as described above, the device that stores the data may be a type of streaming server located on the opposite side of a communication network or may refer to a physical recording medium, but is not limited thereto.

[0085] The receiving unit (310) can receive the encoded video data along with other accompanying data, such as encoded audio data or other auxiliary data, and each of the data can be separated from the video data and provided to an appropriate processing unit (312) other than the video decoder.

[0086] When the video data is received through a communication network, a buffer memory (320) may be combined between the receiver (310) and the decoder (305) to minimize delays and interruptions caused by the network environment. The buffer memory (320) may refer to a computer-readable recording medium that temporarily stores the received video data and reliably supplies it to a parser (330) corresponding to the input terminal of the decoder (305). However, the buffer memory may be unnecessary in environments where the bandwidth of the communication network is sufficient, where video data is being read from a recording medium located at a local location that is not physically separated, or where the possibility of communication delay is not predicted.

[0087] The video decoder (305) may include the parser (330) as its input to interpret the encoded video data. The parser separates (parses) a number of pieces of information stored in the form of bit sequences in the encoded video data according to a predetermined rule, and, if necessary, performs the function of entropy decoding (335) the entropy-coded video data, thereby reconstructing the symbols (338) which are segments of video encoding information. The symbols (338) may include all information for controlling the operation of the decoder (305), and / or may further include information for controlling a device that can operate in conjunction with the decoder (305), such as a display device. Control information for controlling the above-mentioned display device may include information in a format referred to as supplementary enhancement information (SEI) or video usability information (VUI).

[0088] As described above, the parser (330) may be configured to perform entropy decoding (335) of the encoded video data. The method of entropy encoding of the encoded video data may vary depending on the standard of the encoding, and decoding may be performed accordingly. Representative examples of the entropy encoding standard may include variable length coding, Huffman coding, and arithmetic coding, and each of the encoding methods may be context-adaptive or context-sensitive depending on the standard, or may be based on principles widely known to a person skilled in the art.

[0089] The parser (330) may be configured to extract at least one partial image from the encoded video data. The definition of the partial image may vary depending on the encoding standard, and depending on the standard, one or more of the examples listed below may simultaneously overlap. The partial image may be defined in units such as, for example, group of pictures (GOPs), pictures / frames, tiles, slices, macroblocks, blocks, subblocks, transform units (TUs), and prediction units (PUs).

[0090] The above parser (330) may be configured to extract encoding information, such as transform coefficients, quantization parameters (QPs), and / or motion vectors, from the encoded video data. The above parser (330) may be configured to perform entropy decoding (335) and parsing operations on the video data received from the buffer memory, and to selectively decode symbols (338) representing the encoding information. Additionally, the above parser (330) may be configured to selectively supply specific symbols (338) to specific decoding function units within the decoder (305), such as inverse quantization and inverse transform units (340), intra prediction units (350), inter prediction units (355), or loop filter units (360). The control of such information supply may be determined by the information permutation contained in the encoded video and may vary depending on the encoding standard; as such, it is not limited to the scope of the embodiments of the present invention and is not described in detail in this conceptual diagram.

[0091] The above decoder (305) may be composed of a plurality of conceptual functional units that receive and process the encoding information from the parser (330). It is obvious that these conceptual functional units may be combined with one another or further subdivided according to implementation needs. For example, they may be further separated for ease of implementation, or integrated into one for operational efficiency. In any case, each functional unit may be configured to perform close interaction with one another. However, despite the possibility of such integration or separation, the decoding procedure of video data applied as an embodiment of the present invention will be described as a combination of conceptual functional units as described below.

[0092] The above decoder may include an inverse quantization and inverse transform unit (340). The inverse quantization and inverse transform unit (340) may be configured to receive encoding information from the parser (330), including a method to be used for numerical transformation, a block size, quantization coefficients for recovering quantized information, and separation information of a quantization matrix representing the quantization coefficients in a simplified manner, and may be configured to output block values ​​(341) that can be input to an aggregator (370) as a result of processing the encoding information.

[0093] In one embodiment of the present invention, the output values ​​of the inverse quantization and inverse transformation unit (340) may include intra-predicted encoded block values. The intra-predicted block value may mean a value that can be decoded using prediction information within a partial image currently being decoded, such as the current frame, without using prediction information from a previously decoded partial image, such as a previous frame.

[0094] Predicted information within the current partial image may be provided by an intra prediction unit (350). According to an embodiment of the present invention, the intra prediction unit (350) generates a block value of the same form as the block being decoded as predicted information using image information of a spatially adjacent region derived from a partial image that is currently being decoded and has been partially decoded. The partial image information may be provided (381) from a current image buffer, so-called line buffer (380). According to an embodiment, the merging unit (370) may be configured to merge the predicted information (351) generated by the intra prediction unit (350) with the block values ​​(341) provided by the inverse quantization and inverse transformation unit (340).

[0095] In another embodiment, the output values ​​of the inverse quantization and inverse transformation unit (340) may be inter-predicted encoded block values, and in some cases, may include block values ​​for which motion compensation has been performed. In this case, the inter-predicting unit (355) may extract and use sample information (386) used for motion-based prediction from a reference picture buffer (385). The information (356) derived by performing motion compensation on the sample information based on the symbols (338) included in the block value as the output value may be configured to be merged by the merging unit (370) with the block values ​​(341) provided by the inverse quantization and inverse transformation unit (340). In this case, the block values ​​(341) may be referred to as so-called different or residual values.

[0096] The position information in memory used by the inter prediction unit (355) to extract the sample information from the reference image can be determined by a motion vector provided to the inter prediction unit (355), which is composed of, for example, a combination of X, Y, and other symbols (338) for representing specific points of the reference image. The inter prediction unit (355) may also include a function to interpolate and use the sample values ​​when a so-called 'subsampling' possible motion vector is provided, and may further include a function to predict and reinforce the value of the motion vector.

[0097] The output values ​​(371) of the merging unit (370) are provided to the loop filter unit (360) and can be processed by various loop filtering methods. The loop filter unit (360) may be configured to receive not only the block unit output (371) of the merging unit (370) but also the symbol (338) provided by the parser (330) to control its operation. The output of the loop filter unit (360) may be output to an external display means, such as the display device, through an output connection (390), but may be stored (361) in a line buffer (380) for use in prediction to interpret subsequent intra or inter-encoded block values, and may also be stored in a reference image buffer (385) via this.

[0098] Specific partial images, such as frames, can be utilized as reference images for performing predictive decoding during a subsequent decoding process once their decoding is completed. A single partial image, such as a frame, can be accumulated step by step in a line buffer (380) and decoding can proceed. When a frame is decoded, the contents of the line buffer (380) are transferred (383) to the reference image buffer (385), and a new line buffer (380) can be allocated for decoding a new frame.

[0099] The above video decoder (305) may be configured to perform a decoding operation according to a predetermined video compression technology that may be documented by various international standard specifications or commercial specifications. The specifications may include, for example, H.264, H.265, H.266, etc., which are international standard recommendations defined by the International Telecommunication Union Standardization Committee (ITU-T). A person skilled in the art will understand that each of these recommendations is equivalent to an international standard jointly defined by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC). The encoded video data may comply with a specific bitstream syntax defined by the specifications, as defined by the profile and level specified in the video compression specification document and standard document, and specifically within such document, and as required. In addition, the complexity of the encoded video data may be limited to a certain level for compliance with the profile and level. For example, any profile or level may be configured to limit the maximum image size, the maximum decoding rate, and the maximum reference image size. These limitations may also, in some embodiments, be further limited through metadata signals regarding the management of the HRD buffer included in the hypothetical reference decoder (HRD) and the encoded video data.

[0100] According to one embodiment of the present invention, the receiver (310) may receive additional redundant data along with the encoded video. The additional data may be considered as part of the encoded video data. The additional data may include information that can be used by the decoder (305) to properly decode the data or to more accurately reconstruct an image that is close to the image before encoding. The additional data may be provided in the form, for example, time, space, or layers for signal-to-noise ratio (SNR) enhancement, redundant slices, redundant images, and forward error correction codes.

[0101] FIG. 4 is a conceptual diagram of a functional unit of a video encoder according to an embodiment of the present invention. The encoder (405) may be configured to receive original video information (402) from a video source (401) and perform encoding.

[0102] The original video information (402) may have any suitable bit depth, for example, 8 bits, 10 bits, 12 bits, etc. Additionally, the original video information (402) may have any suitable color space, for example, R / G / B, Y / U / V, Y / Cb / Cr, etc. Additionally, the original video source may have any suitable sampling structure corresponding to the color space, for example, in the form of Y / Cb / Cr 4:2:0, Y / Cb / Cr 4:4:4. The original video source having such a predetermined format may be provided to the encoder in the form of a digital video stream.

[0103] In a unidirectional video communication network, the original video information (402) may be obtained from a recording medium that stores a pre-prepared original video. In a bidirectional video communication network, the original video information (402) may be obtained from a video acquisition device, such as a camera, that generates at least one video transmission stream included in the bidirectional video communication.

[0104] The video data containing the original video information (402) may be composed of a plurality of partial images configured to simulate motion by playing them in chronological order. The partial images may be expressed as concepts such as, for example, a picture or a frame. The partial images may include one or more samples depending on the type of sampling structure, color space, etc. used. A person skilled in the art will understand that the term "sample" is closely related to "pixel" in digital images. The operation of the encoder will be explained below with reference to such samples.

[0105] According to one embodiment of the present invention, the encoder (405) may be configured to encode and compress (partial) images constituting the original video information (402) into the form of encoded video information in real time (or according to other temporal requirements as needed depending on the method of implementation).

[0106] In the above encoder (405), the control unit (450) may be a functional unit configured to control an appropriate encoding speed. The control unit (450) may be configured to control other functional units and to be functionally coupled to the functional units described below. The parameters set by the control unit (450) may include parameters related to bitrate control, such as image skip, quantizer, and variable values ​​for applying image quality optimization techniques, and may also include values ​​such as image size, the structure of a group of pictures (GOP), and the maximum search range of motion vectors. A person skilled in the art will be able to understand the various other functions that the control unit (450) may have, and such other functions may be added or removed according to the design of the video encoder optimized for the individual system design.

[0107] According to an embodiment of the present invention, the encoder (405) may be configured to operate in a structure such as a "coding loop" that is well known to a person skilled in the art. To explain it in a simplified manner for example, the coding loop may consist of an internal encoder (so-called "source coder") (410) responsible for receiving an image to be encoded and generating symbols based on at least one reference image that has been previously encoded, and a local decoder (420) configured to be connected to the internal encoder. The local decoder (420) may be configured to perform an operation to reproduce sample data to be generated by a decoder (490) located at an actual remote location that receives video information encoded from the encoder (405) by receiving the output of the internal encoder (410).

[0108] Video data composed of sample data reconstructed by the internal decoder (420) can be configured to be input into the reference image buffer of the encoder (405). As described above, since the internal decoder (420) is implemented to reproduce the result output by the encoder (405) to be decoded at a remote decoder, the video data stored in the reference image buffer can also be identical in bit unit to the information of the reference image buffer held by the remote decoder. That is, the prediction function unit that may be included in the encoder (405) can read values ​​identical to the sample values ​​of the previous frame that the decoder will refer to during the decoding process from the reference image buffer of the encoder (405).

[0109] As described above, the principle of achieving a match between the encoder (405) and the decoder (490) by means of an internal decoder (420) on the side of the encoder (405) is widely known to a person skilled in the art, and the method of responding to an environment where such an environment is not guaranteed (e.g., loss of information due to communication failure) may also follow what is known to a person skilled in the art.

[0110] An example of the operation method of the internal decoder (420) described above has been explained in detail with reference to FIG. 3. The decoder of FIG. 3 can be considered as the decoder (490) of the "remote location" described above. The internal decoder (420) may be implemented excluding lossless encoding and decoding sections such as the parser (330) or entropy decoding (335), because the internal encoder (405) is implemented simply to reproduce the operation of the decoder located at the remote location, so it is acceptable to decode the symbols immediately without requiring a process of compressing and restoring the symbols. Therefore, it is acceptable for the functional parts preceding the parser and entropy decoder shown in FIG. 3 not to be provided or at least to be implemented only partially.

[0111] As described above, according to a preferred embodiment of the present invention, any decoder function part present in the decoder (excluding the parser and entropy decoder) may naturally exist as substantially the same function part in the corresponding encoder (405).

[0112] The operation of the encoding function unit that may be included in the encoder (405) can be considered as the inverse of the decoder function unit. Therefore, the embodiment can generally be explained by performing the operation of the decoder function unit in reverse. For example, a quantization and transform function unit corresponding to the inverse quantization and inverse transform unit may be provided, and an inter-prediction encoding unit corresponding to the inter-prediction unit may be provided. In addition to this, some additional explanations will be added.

[0113] The internal encoder (410) may be configured to perform encoding for input image information, e.g., an input frame, by a predictive encoding method executed by a predictive encoding unit (440) that operates by referencing at least one portion of an image encoded in a temporally earlier order, e.g., frames, from a reference image buffer (430) from at least one reference image information, e.g., video data designated as a reference frame. In this case, the encoder (405) may be configured to encode a differential between blocks of samples constituting the input image and blocks of samples constituting the reference image.

[0114] The internal decoder (420) can decode video data that can be designated as the reference image from the symbols generated by the internal encoder (410). As described above, since this video data is identical to the decoding operation performed by the remote decoder, the video data used as the reference image may be provided to the encoder (405) in a form that has undergone lossy compression and has some damage, and this operation may be intended to match the operation with the decoder.

[0115] The prediction encoding unit (440) may be configured to perform a prediction search operation within the encoder (405). The prediction search operation may refer to an operation corresponding to the inter-prediction or intra-prediction described in the description of the decoder. For image information that is input and scheduled to be newly encoded, the prediction unit may access the reference image buffer (430) to retrieve information, such as motion vectors, block shapes, metadata that may include the same, and sample blocks to be actually referenced, which are information indicating points of reference images that can function as prediction reference information suitable for the new image information. The prediction encoding unit (440) may operate according to the so-called "sample block by pixel block" standard to obtain appropriate prediction reference information. According to one embodiment of the present invention, at least one prediction reference information pointing to at least one reference image information stored in the reference image buffer (430) may be designated for the input image, as determined based on the search results obtained by the prediction encoding unit (440).

[0116] In one embodiment of the present invention, the control unit (450) may be configured to manage the overall encoding operation of the internal encoder (410), including the setting of parameters used to encode video data.

[0117] The outputs of all the aforementioned functional units may be subject to entropy coding (460) for final output. The entropy coding (460) may include various entropy coding techniques such as variable length coding, Huffman coding, and arithmetic coding as described above for the symbols generated by the various functional units, and each of the coding methods may be context-adaptive or context-sensitive methods depending on the standard, or may be based on principles widely known to a person skilled in the art. Such entropy coding (460) can typically achieve lossless compression and may be configured to convert at least one symbol generated by the functional units into encoded video data.

[0118] The control unit (450) may apply a type of encoding of a specific partial image during the encoding interval to each partial image, such as a picture or a frame, in controlling the operation of the encoder (405). Depending on the type, the method of encoding the partial image may be affected. Depending on the embodiment, the type may include a classification of "frame types" as follows.

[0119] FIG. 5 is a conceptual diagram of a frame type according to an embodiment of the present invention. The following description will be explained together with reference to FIG. 5.

[0120] An intra ("I") image (510) may refer to an image that can be encoded and decoded using only its own information without referencing other image information within the video data through predictive encoding. The "I" image may be designated by names such as key frame, independent / instantaneous decoder update (IDR) frame, and clean random-access (CRA) frame according to the video encoding standard. The "I" image designated by such various names may have various modifications and application methods as permitted by each standard and may differ partially from one another. In addition to those listed above, various application methods for implementing the "I" image may be based on various methods that are already known to a person skilled in the art or may be newly provided.

[0121] The prediction ("P") image (520) may mean an image that can be encoded and decoded via intra- or inter-prediction based on at least one prediction information and / or motion vector pointing to at least one reference image to predict sample values ​​of a block constituting the image. The "P" image may be configured to reference only one reference frame according to the video encoding standard, or may be configured to reference one or more reference frames. In the case of referencing one or more reference frames, sample information and / or associated metadata derived from multiple reference images may be used for the reconstruction of a single block. However, in common cases, the image designated as the "P" image may be understood as an image that performs a reference limited to the temporally preceding image.

[0122] A bidirectional prediction ("B") image (530) may mean an image that can be encoded and decoded via intra or inter prediction based on at least one prediction information and / or motion vector pointing to at least two reference images to predict sample values ​​of blocks constituting the image. In common cases, the image designated as the "B" image is distinguished from the image designated as the "P" image and may be understood as an image that performs the reference, not limited to the image that precedes it in time.

[0123] Video data is spatially divided into multiple sample blocks during the encoding and decoding process, and encoding can proceed in block units. The block units include, for example, sizes such as 4x4, 8x8, 4x8, or 16x16 in units of horizontal / vertical pixels, as is widely known, but are not limited thereto. The blocks may be encoded by a predictive encoding method by referencing any other (already encoded) blocks, as allowed and / or restricted by the type specified for each partial image in which the blocks are included. For example, blocks of the "I" image (510) may not use a predictive encoding method, or may be encoded by referencing blocks that have already been encoded within the same partial image. That is, only the so-called intra-prediction method may be used. In contrast, for the "P" image (520), at least one reference image encoded in a previous time unit may be referenced, and thus inter-prediction may also be used in encoding along with intra-prediction. In the case of the "B" image (530), reference can be performed even among reference images that are encoded earlier in the encoding order but follow later in terms of time unit. However, it is widely known that there may be blocks encoded within the "P" image or the "B" image that do not rely on predictive encoding.

[0124] The above video encoder (405) may be configured to perform encoding operations according to a predetermined video compression technology that may be documented by various international standard specifications or commercial specifications. Examples of the specifications may include all those described in the decoder.

[0125] According to one embodiment of the present invention, a transmitting unit (470) may buffer the encoded video data generated by the entropy encoding in order to provide / transmit the video data (ultimately to a decoder (490) at a remote location) through a hardware or software connection (495) to a device storing the encoded video data. According to an embodiment, when providing / transmitting the encoded video data from the video encoder (405), the transmitting unit (470) may receive and merge other data accompanying the encoded video data, such as encoded audio data or other auxiliary data, from a separate source (480).

[0126] According to one embodiment of the present invention, the transmitting unit (470) may be configured to transmit additional data along with the encoded video. The additional data may be considered as part of the encoded video data. The additional data may include information that can be used by a decoder to properly decode the data, or to more accurately reconstruct an image that is close to the image before encoding. Examples of the additional data may include all the examples previously shown in relation to the receiving unit (310) of the decoder.

[0127] The present invention can be implemented by digital video compression standards that are understood by and widely used by people skilled in the art, as described above. The digital video compression standards may include at least one of the compression standards known by standard names such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.

[0128] Figure 6 is a conceptual diagram showing the structure of a video encoder according to the H.266 / VVC standard. What is shown in Figure 6 corresponds to the approximate structure of a video encoder widely known by standard codes such as ITU-T H.266 and ISO / IEC 23090-3, and by the designation MPEG-I Part 3 or the common name versatile video coding (VVC).

[0129] According to FIG. 6, a video encoder (605) may be configured to receive uncompressed and unencoded original video data (601) as input and output an encoded bit sequence (602). When the video data (601) is intra-encoded, it may be supplied directly to a luminance signal mapping unit (610a), or supplied to a luminance signal mapping unit (610b) via an inter-prediction unit (620) that includes motion vector extraction. When the intra-encoded, the mapped luminance signal may be supplied to an output merger (606) by selecting (608) at least one of the intra-prediction encoded signal via the intra-prediction unit (625) or the inter-prediction encoded signal output from the luminance signal mapping unit (610b) via the inter-prediction unit (620). The result of the output merger may be applied to a chroma scaling unit (615). (The operation of the above-mentioned illuminance signal mapping unit (610) and the operation of the above-mentioned color difference signal reduction (615) are collectively referred to as the luma mapping / chroma scaling (LMCS) process.) The reduced color difference signal can be provided to a transform unit (630), and the transform unit (630) can perform an adaptive color transform, particularly on the color difference signal. The coefficients derived as a result of the transformation are applied to a quantization unit (640) and quantized. This results in lossy compression, and the result of the lossy compression can be output as a bit sequence (602) through a multi-hypothesis context-adaptive arithmetic coding unit (650), which is a lossless compression method.

[0130] Meanwhile, the result of the lossy compression above can enter the decoding procedure by undergoing inverse quantization (645), inverse transform (635), and luminance signal expansion (617) processes to generate a coding loop. The result of the luminance signal expansion above can be supplied to an internal merger (607) along with the result of selecting (608) at least one of the previously generated intra-predicted coded signal or inter-predicted coded signal. The result of the internal merger above can undergo inverse luminance signal mapping (617) and then undergo processing such as a deblocking filter (660), a sample adaptive offset (SAO) (670), and an adaptive loop filter (ALF) to reproduce the image quality improvement process in the decoder. The result of reproducing the operation in the decoder as described above is applied to the reference image buffer (690) and can be recycled for prediction encoding by the inter prediction unit (620).

[0131] The present invention may also be utilized by or combined with the enhanced compression model (ECM), which is an implementation of a next-generation video codec being developed by the Joint Video Experts Team (JVET), an international standardization expert organization, to improve H.266 / VVC. According to the standardization history document of the above JVET, document number ISO / IEC JTC 1 / SC 29 / WG 5 N 190 (also document number JVET AC2025-v1), the improved compression model may include an improved intra-predictive coding method, an improved inter-predictive coding method, an improved transform and transform coefficient coding method, an improved adaptive loop filtering method, a bilateral filtering method, a new sample adaptive offset (SAO) method for image quality improvement, an extended entropy coding method, and an improved gradual decoding refresh (GDR) technique, and such techniques may be included in the present invention as background technology for implementing the present invention.

[0132]

[0133] General of the Decoder-Side Intra-Predicted Mode Estimation (DIMD) Application Method

[0134] The present invention provides a method for applying an improved decoder-side intra mode derivation (DIMD) technique that can be used in video encoding and decoding regions including the embodiments described above, and an apparatus to which such method is applied. As described above, DIMD proposes a method for calculating the intra-prediction direction of the current block from the values ​​of a conventionally decoded image adjacent to the currently encoded / decoded block in order to reduce the amount of information transmitted from the encoder to the decoder in the process of increasing the compression ratio of an image by intra-prediction, i.e., intra-prediction. Accordingly, even without the encoder providing a separate signal, the decoder in DIMD mode can autonomously derive directional information of the intra-prediction from the preceding decoding process, thereby resulting in an improvement in compression efficiency.

[0135]

[0136] histogram of gradients

[0137] To estimate the intra-prediction mode for a block, the DIMD method may be configured to use a histogram of gradients (HoG). This histogram can be understood as a graph representing the intensity for each prediction direction of the block allowed by the video encoding / decoding standard to which the DIMD method is applied. In general video encoding / decoding standards, multiple directional prediction modes are defined. For example, H.265 / HEVC defines 33 directional prediction modes, and H.266 / VVC defines 65 directional prediction modes. Since each directional mode represents a specific angle, the intensity may be defined in correspondence with each of these directional modes. A person skilled in the art will understand that the histogram may be adaptedly configured even when a different number or type of directional prediction modes are used in conventional or newly provided video encoding / decoding standards. In addition, a person skilled in the art will understand that in generating the above histogram, the specific number of directional prediction modes defined in the video encoding / decoding standard as described above may be used in an arbitrarily subdivided form or in a form where some are omitted. For example, in generating a histogram, the 65 directional prediction modes of H.266 / VVC may be further subdivided to be defined as 66 or more angles, or they may be partially omitted to be defined as 64 or fewer angles.

[0138] In order to generate the above histogram, a set of adjacent pixels to perform gradient analysis may first be selected. The set of adjacent pixels may be derived from pixels that have already been decoded and reconstructed. FIG. 7 is a conceptual diagram of the selection of a set of adjacent pixels according to an embodiment of the present invention. Referring to FIG. 7, in a frame (700) of an image being decoded, an area (750) that has already been decoded and reconstructed and an area (760) that has not yet been decoded may be distinguished based on the current block (710) currently being decoded. At this time, as exemplified in FIG. 7, a set of pixels surrounding the reconstructed area (750) by TW pixels (717) to the left and TH pixels (718) upward from the current block (710) may be selected as a template (715). According to an embodiment of the present invention, the TW and TH may be set to 3. According to another embodiment of the present invention, the TH and the TW may be set to the same value or different value. In one embodiment of the present invention, either the TW or the TH may not be defined or may be defined as 0. In other words, the template (715) may be defined only in the upward direction by specifying only 1 or more THs, or defined only in the left direction by specifying only 1 or more TWs.

[0139] Next, a slope analysis can be performed on the pixels included in the template (715). Through this, the slope direction of the image of the pixels within the template (715) can be determined, and it is highly likely to be the same as the slope direction of the image of the current block. Accordingly, according to one embodiment of the present invention, a horizontal and vertical Sobel filter defined as in Equation 1 can be used to estimate the slope direction of the image of the current block by measuring the slope of the template (715).

[0140]

[0141] In order to apply the Sobel filter matrices to each pixel of the above template (715), the filter may be applied by forming a window consisting of eight immediately adjacent pixels centered on each pixel of at least one pixel of the above template (715). Through this, the Sobel filter matrix M of Equation 1 x Horizontal slope value G by the application of x 를, M y Vertical slope value G by the application of y You can obtain.

[0142] FIG. 8 is a conceptual diagram of the slope calculation for each pixel within a template according to an embodiment of the present invention. The example illustrated in FIG. 8 is for a template (810) with a width of 3 pixels (817) for the current block (710) shown in FIG. 7. As shown in FIG. 8 (a), the pixels (820) corresponding to the centerline of the template (810) are designated as pixels to be analyzed for slope, and a window (830) of 3 pixels horizontally and 3 pixels vertically is formed for each reference pixel (835) to apply the Sobel filter. By applying the filter, the amplitude ("Ampl") and angle ("Angle") of the slope can be calculated for each pixel (820) on the centerline of the template (810) using slopes Gx and Gy as shown in Equation 2 below.

[0143]

[0144] Next, a histogram can be obtained based on the direction of the gradient. The histogram can be formed by an angle-specific index of the intra-prediction mode allowed by the video encoder and decoder to which the DIMD method is applied, as described above. Depending on each angle value, the histogram value of the corresponding inner angle mode can be configured to increase by Ampl. When all the pixels (820) on the centerline of the template (810) are processed, the histogram may contain the cumulative value of the gradient amplitude for each inner angle mode. FIG. 8(b) shows an example of a histogram calculated after applying the above operation to all pixel locations of the template. Referring to FIG. 8(b), it can be seen that the histogram can be displayed as an amplitude for each angle.

[0145] According to one embodiment of the present invention, other calculation methods other than those described above may be used to generate the histogram. For example, the angle may be calculated based on the angle value of the directional prediction mode of the surrounding block, which is already determined, instead of being calculated based on the direction of the slope. As another example, the amplitude may be determined and / or adjusted based on a value determined proportionally or exponentially based on the size and / or number of samples included in the template, or samples included in the surrounding block included in the template.

[0146] There is a concern that an intra-predictive coding method referencing pixels above the current block may cause complexity issues related to so-called line buffering in the decoder. Therefore, in one embodiment, for a block located at the upper boundary of a large coding / decoding unit (e.g., a coding tree unit (CTU) indicated in the VVC codec specification), gradient analysis may not be performed on pixels located in the upper part of the template.

[0147] According to one embodiment, the mode showing the highest amplitude, i.e., the peak, in the histogram can be configured to be selected as the intra-prediction mode for the current block. Additionally, if the maximum value of the histogram is 0 (meaning that slope analysis cannot be performed or that the image of the area constituting the template is flat without slope), the average (DC) mode can be selected as the intra-prediction mode for the current block.

[0148] According to one embodiment, intra-prediction for the current block can be performed using a mode that exhibits at least one upper peak in the histogram. To this end, a prediction fusion technique based on weighted summation or weighted average may be utilized.

[0149] According to one embodiment, intra-prediction for the current block can be performed by using the prediction results of a mode showing at least one upper peak in the histogram and a planar mode together. To this end, prediction fusion techniques based on weighted summation or weighted average can be utilized as described above.

[0150] According to one embodiment, when the size of the block is less than a certain size, for example, less than or equal to the size of 4x4 pixels, the application frequency of the Sobel filter can be made less frequent. FIG. 9 is a conceptual diagram of the reduced frequency application of the Sobel filter according to one embodiment of the present invention. Referring to FIG. 9(a), for a small block (910), windows (931, 932) can be constructed using one pixel (935) from the left and one pixel (936) from the top, respectively, to calculate the gradient. In addition to reducing the number of operations for calculating the gradient, such an embodiment can contribute to simplifying the selection of the best two modes in the histogram, as exemplified in FIG. 9(b).

[0151] According to one embodiment, the template may be extended in the upper right and lower left directions. If available pixels exist, in a block of horizontal W pixels and vertical H pixels, the template may be extended by up to W pixels in the upper right direction and up to H pixels in the lower left direction. According to another embodiment, the template may be excluded or disabled in a specific direction. For example, the template may be defined only in the upper direction or only in the left direction. The decision to disable such a template may be made when the block satisfies an edge condition, or based on other conditions / flags / signals regardless thereof.

[0152]

[0153] prediction fusion

[0154] A prediction fusion algorithm may refer to an algorithm that derives a single prediction value by fusing at least one prediction value. According to an embodiment of the present invention, at least one prediction value may be generated by a spike of the histogram, and the prediction fusion algorithm may be used to combine them.

[0155] FIG. 10 is a conceptual diagram illustrating a prediction fusion algorithm for DIMD according to an embodiment of the present invention. According to an embodiment of the present invention, three angles corresponding to the three highest spikes of the histogram can be detected as M1 (1011), M2 (1012), and M3 (1013). Next, by applying pixel prediction based on the three angles (1011, 1012, 1013) to reference pixels (1020) respectively (1021, 1022, 1023), intra-predicted pixel information Pred1 (1031), Pred2 (1032), and Pred3 (1033) are obtained, and then pixel information obtained by fusing (1030) each of the intra-predicted pixel information can be calculated as the final prediction value (1050) of the block. For the above fusion (1030), a weighted average of three predictor variables with respect to the spike amplitude of the histogram can be calculated (1041, 1042, 1043). In a similar manner, at least two or more varying numbers of spikes from the histogram can be selected and merged together.

[0156] FIG. 11 is a conceptual diagram illustrating a prediction fusion algorithm for DIMD according to another embodiment of the present invention. According to one embodiment, two angles corresponding to the two highest spikes of the histogram can be detected as M1 (1111) and M2 (1112). Next, pixel predictions (1121, 1122) based on the two angles (1111, 1112) are applied to reference pixels (1120) respectively to obtain Pred1 (1131) and Pred2 (1132), and Pred3 (1123) can also be obtained by conventional planar prediction. The pixel information obtained by fusing (1130) each of the intra-predicted pixel information can be calculated as the final prediction value of the block. To this end, weights (1141, 1142, 1143) for the fusion (1130) may be applied. The weighted prediction value (1143) for Pred3 (1133) may be fixed at 1 / 3, or it may be set to 21 / 64 for ease of bit-based calculation. The respective weighted prediction values ​​(1141, 1142) for Pred1 (1131) and Pred2 (1132) may be determined in proportion to the spike amplitude size within the histogram within a range where the sum is 2 / 3 (or 43 / 64). According to one embodiment of the present invention, three, four, five, or more spikes may be selected from the histogram by a method similar to or a combination of methods shown in FIGS. 11 and FIGS. 12, and may be merged with the prediction result (1143) of the planar prediction mode.

[0157] According to one embodiment of the present invention, in a block of horizontal W pixels and vertical H pixels, if the upper or left histogram size is twice as large as the other, the weights for the predictions derived from each spike may be modified. In this case, the weights may vary depending on the pixel location from which the predictions are derived within the block.

[0158] For example, if the upper histogram is twice the size of the left one, the weights can be modified as shown in Equation 3 below.

[0159]

[0160] As another example, if the left histogram is twice the size of the upper one, the weights can be modified as shown in Equation 4 below.

[0161]

[0162] wDimd in the above mathematical formulas 3 and 4 i can mean the unmodified weight, and Δi is a predefined arbitrary number, for example, 10.

[0163]

[0164] General of Merged DIMD Application Methods

[0165] A technique for realizing a merged DIMD (DIMD Merge) that uses DIMD to cite the result of the DIMD processing from a previously encoded / decoded block and reuses it in a subsequent block has also been proposed.

[0166] FIG. 12 is a conceptual diagram of a merged DIMD according to an embodiment of the present invention. Referring to FIG. 12, the merged DIMD may be configured to execute a DIMD process for the current block (1210) using DIMD information (1225, 1235) extracted from surrounding blocks (1220, 1230) of the current block (1210). In particular, a new merged histogram (1215) for the current block (1210) may be calculated based on the histograms (1225, 1235) of neighboring blocks (1220, 1230). For example, a new merged histogram of gradients (MHoG) for the current block may be calculated based on the HoG of neighboring blocks. Additionally, for example, a new Merged Histogram of Occurrences (MHoC) for the current block can be calculated based on the HoC of a neighboring block. According to one embodiment of the present invention, for the merged DIMD to be used, at least one block (1220, 1230) adjacent to the block (1210) currently being encoded / decoded must be encoded by DIMD or merged DIMD. If one adjacent block (1220, 1230) encoded by DIMD or merged DIMD is available, the histogram (1225, 1235) of that block can be used as is as the merged histogram (1215) for the current block. If two or more adjacent blocks (1220, 1230) are available, their respective histograms (1225, 1235) can be combined to derive the merged histogram (1215). According to one embodiment of the present invention, each histogram (1225, 1235) can be merged into the merged histogram (1215) by taking the average of its amplitude.

[0167]

[0168] Emergence-based intracoding (OBIC)

[0169] Occurrence-based intra coding (OBIC) is a technique that operates similarly to the DIMD method described above, and in this specification, it may be considered as an embodiment implementing the DIMD or an embodiment applied from the DIMD method. According to the OBIC, the Histogram of Occurrences (HoC) may be used as the histogram in conjunction with, compatible with, and / or replacing the HoG. The HoC may refer to a histogram that collects the frequency of occurrence of intra prediction modes possessed by samples of surrounding blocks in proportion to the sample size of surrounding blocks. The HoC and the HoG may be used in conjunction with, compatible with, and / or interchangeable in that all or most of the intra prediction modes collected by the HoC include the directional prediction modes collected by the HoG, and, for example, when generating the histogram, each amplitude value may be considered to have been generated by the frequency of occurrence in proportion to the sample size. When a histogram including HoC is generated as described above, the fusion of intra-prediction results, such as exemplified in FIG. 10 or FIG. 11, can be performed based on this histogram by the same or similar method.

[0170] Therefore, it is evident that all embodiments described in this invention can be applied to OBIC as well as DIMD within the scope that does not conflict with the technical concept of this invention. In particular, regarding the description of histograms in this invention, it will be understood that the disclosure of this invention concerning the operation and manipulation thereof can be applied in the same way when using gradient histograms (HoG) and occurrence histograms (HoC).

[0171]

[0172] Template-based Intramode Estimation (TIMD)

[0173] Template-based intra mode derivation (TIMD) technology is one of the methods designed to derive an intra prediction result value based on minimal signals or without a separate signal regarding the mode in relation to intra prediction on the decoder side, and can be considered as a type of encoding / decoding technology corresponding to a concept similar to the aforementioned DIMD in a broad sense. The aforementioned TIMD technology may refer to a method that enables efficient intra prediction based on a template derived from samples adjacent to the current block among conventionally decoded samples.

[0174] According to one embodiment of the present invention, the TIMD technology may be configured to extract a template region consisting of samples adjacent to the current block among conventionally decoded samples, and to estimate the most suitable intra-probable mode based on the values ​​of the template. According to an embodiment, the TIMD technology may operate based on a partial list of intra-probable modes configured by including a list of most probable modes (MPMs). In one embodiment, the MPM list may be composed of an indexed list including intra-probable modes of surrounding blocks and some basic modes (e.g., plane or DC probable modes, etc.).

[0175] According to one embodiment of the present invention, in the encoding and / or decoding process using the TIMD technology, the MPM list is derived by the same method shared by the encoder and the decoder, and for each intra prediction mode included in the MPM list, a prediction based on a potential template region is performed, and the prediction result can be compared with actual samples of the extracted template region. The above-described procedure may be understood as corresponding to a procedure commonly referred to as "template matching." Depending on the embodiment, the comparison may be performed using a cost function such as SATD (sum of absolute transformed differences). As a result of the comparison, the mode with the lowest cost may be selected as the intra prediction mode of the current block.

[0176] According to one embodiment of the present invention, the TIMD may be configured to select two or more modes and fuse prediction blocks derived based on each of the selected modes. For example, it may be configured to select multiple prediction modes in order of lowest cost through the template matching and to derive a final prediction block based on the result of weighted sum or weighted average of each prediction mode. The prediction fusion may be performed by a process identical or similar to the prediction fusion of the DIMD technology.

[0177] According to one embodiment of the present invention, in order to fuse prediction modes in the TIMD technique, for each intra prediction mode included in the MPM list (which may include a wide-angle mode if upper-right and / or lower-left reference samples are available), a cost function, e.g., SATD, between the prediction sample and the reconstruction sample of the template may be calculated. Based on the calculation, at least two directional intra prediction modes with the minimum cost and at least one non-directional intra prediction mode with the minimum cost (e.g., planar or DC mode) may be selected as TIMD modes. The selected TIMD modes may be configured to be prediction fused with weights after applying a Position Dependent Prediction Combination (PDPC) process.

[0178] At this time, whether the non-directional intra prediction mode is used in the prediction fusion can be determined according to conditions. According to one embodiment, the conditions may include cases where the non-directional intra prediction mode is different from two selected intra prediction modes, and / or cases where the lowest cost by the non-directional intra prediction mode is lower than the lowest cost by the directional intra prediction mode by a certain level (e.g., less than 1.5 times the directional lowest cost).

[0179] According to one embodiment, the non-directional intra prediction mode may be included in the prediction fusion only when both of the above two conditions are true. In this case, in the TIMD technology-based intra prediction fusion, the weight of each intra prediction mode may be calculated from the calculated cost (e.g., SATD cost) as shown in Equation 5 below.

[0180]

[0181] According to one embodiment, if either of the two conditions is not true, the non-directional intra prediction mode may not be included in the prediction fusion. In this case, the costs of the selected at least two directional intra prediction modes may be compared by a threshold. The comparison by the threshold may mean, for example, assuming that a first mode and a second mode are selected, comparing the costs of the two modes and determining that the cost of the second mode is less than twice the cost of the first mode. In this case, the first mode and the second mode may be combined by prediction fusion, or otherwise, it may be determined that only the first mode is used. In this case, in the TIMD technology-based intra prediction fusion, the weight of each intra prediction mode may be calculated from the calculated cost (e.g., SATD cost) as shown in Equation 6 below.

[0182]

[0183] According to one embodiment of the present invention, at least one division operation used in the prediction fusion may be configured to be processed with low computational complexity using an integerization method based on a lookup table, for example, a lookup table identical or similar to that used in the implementation of the CCLM (Cross-Component Linear Model) technology may be used. Additionally, in the TIMD-based prediction fusion process, depending on the embodiment, a fusion operation based on criteria and / or weights dependent on individual pixel positions of the derived prediction block may be applied, similar to that used in the DIMD-based prediction fusion process. In one embodiment, the criteria may be replaced with those based on the cost, e.g., the SATD cost. In one embodiment, the cost may be configured to be determined from the ratio of the normalized SATDs of the selected TIMD-based prediction modes calculated in the upper and left template regions.

[0184]

[0185] Template-based multiple reference lines (TMRL)

[0186] Template-based multiple reference line (TMRL) technology is a technical concept that can operate in combination with the aforementioned DIMD and / or TIMD technologies, and may refer to an intra-prediction method configured to utilize multiple reference lines by extending the scope of template references.

[0187] According to one embodiment of the present invention, when the TMRL is applied, the thickness and distance of the template extracted from samples adjacent to the current block may be extended. By applying the TMRL, multiple reference lines located at various distances from the current block may be configured to be used as candidates for the extended template. For example, not only a template based on a reference line composed of samples immediately adjacent to the current block, but also reference lines adjacent to the current block at intervals of 2 pixels, 3 pixels, or more may be configured to be utilized as extended templates in the encoding / decoding process. According to an embodiment, the method of referencing the plurality of reference lines may operate by providing an index for a list in which a predetermined number of reference lines is set. For example, the number of reference lines may operate in units of 1, 3, 5, 7, or 12 pixels.

[0188] According to one embodiment of the present invention, when the TMRL is applied, the operation for the template matching described above can be repeatedly applied to a plurality of reference lines. For example, the decision process of various intra prediction modes based on template matching can be refined based on a plurality of templates by comparing the costs of the intra prediction modes for each reference line and finally selecting an intra prediction mode based on the result of synthesizing the costs derived from a plurality of reference lines. In addition, according to an embodiment, if the template area can be extended in the direction of the upper right and / or lower left, a plurality of templates based on the plurality of reference lines referenced by the TMRL can also similarly be extended in the direction of the upper right and / or lower left.

[0189]

[0190] Composition of the present invention

[0191] The present invention can provide an adaptive method for using a decoder-side intra prediction mode estimation method that can be used in video encoding and decoding regions including the embodiments described above. More specifically, the present invention aims to provide a more efficient intra prediction method by combining and compromising the advantages of various intra prediction methods described above, such as DIMD, OBIC, TIMD, TMRL, and others not limited thereto.

[0192] According to one embodiment of the present invention, a plurality of decoder-side intra-prediction mode estimation methods, such as DIMD, OBIC, TIMD, and TMRL, may each be used independently, but the core technical concept of the present invention is to achieve an optimal balance between encoding efficiency and decoding complexity by selectively combining the advantages of these methods or switching to use them conditionally. To this end, each embodiment of the present invention may be configured to operate adaptively based on various conditions, such as the characteristics of the current block, encoding information of surrounding blocks, texture complexity of the template area, and the number of allowed reference lines.

[0193]

[0194] Merging generation of histograms

[0195] According to one embodiment of the present invention, in the process of generating a histogram for intra-prediction mode estimation on the decoder side, more accurate and reliable prediction mode estimation is possible by combining multiple histogram generation methods based on different principles. In conventional DIMD methods, the directionality of the template region was analyzed mainly using a gradient histogram (HoG), and in OBIC methods, an occurrence histogram (HoC) based on the frequency of occurrence of intra-prediction modes of surrounding blocks was used. However, when these methods are used independently, there was a problem in that the reliability of HoG may decrease when the template region is flat or noisy, and HoC may lead to inaccurate prediction when surrounding blocks have textures different from the current block.

[0196] According to one embodiment of the present invention, when using a method for estimating an intra-prediction mode based on a histogram, such as the DIMD and / or OBIC, the histogram generation method described above may be used in combination to determine the amplitude of the histogram and to derive the histogram, e.g., HoG or HoC. The histogram derived by the combination method may be treated as a single histogram and used to implement the HoG for the DIMD or the HoC for the OBIC and other histograms or priority-based intra-prediction modes.

[0197] According to one embodiment, when generating the histogram, if at least one surrounding block adjacent to the current block is encoded in a directional intra-prediction mode, at least one histogram value can be calculated using the sample size of the surrounding block and the angle of the directional intra-prediction mode of the surrounding block, for example, as in the method for deriving histogram values ​​for the HoC described above. Furthermore, a histogram can be constructed by limiting the samples adjacent to the current block that do not correspond to the block encoded in the directional intra-prediction mode, and through directional filtering of the samples, for example, as in the method for deriving histogram values ​​using a 3x3 Sobel filter for the HoG described above.

[0198] According to the above embodiment, since the intra-prediction mode information of surrounding blocks is already determined, it can be directly reflected in the histogram without additional gradient analysis, thereby reducing the amount of computation. Furthermore, by accumulating histogram values ​​in proportion to the sample size of surrounding blocks, greater weight can be assigned to directions that appear consistently over a wide area. Meanwhile, by performing gradient analysis, such as a Sobel filter, only on samples that are not encoded in the directional intra-prediction mode, unnecessary redundant computations can be prevented and processing efficiency can be improved.

[0199] According to another embodiment, when generating the histogram, if at least one surrounding block adjacent to the current block is encoded in a directional intra-prediction mode, a first histogram can be constructed using the sample size of the surrounding block and the angle of the directional intra-prediction mode of the surrounding block, for example, as in the method for deriving histogram values ​​for the HoC described above. Additionally, for the samples adjacent to the current block, a second histogram can be constructed through directional filtering of the samples, for example, as in the method for deriving histogram values ​​using a 3x3 Sobel filter for the HoG described above. Next, an integrated histogram is obtained by merging the first histogram and the second histogram, and the integrated histogram can be configured to be used for purposes such as the HoG or HoC.

[0200] According to the above embodiment, by independently generating the first histogram and the second histogram and then merging them, flexibility can be secured to individually evaluate and appropriately adjust the reliability of each histogram generation method. For example, a higher weight can be assigned to the first histogram if there are many surrounding blocks and they have a consistent directionality, and conversely, a higher weight can be assigned to the second histogram if the slope analysis of the template area indicates a clear directionality.

[0201] According to an embodiment, the method for obtaining the integrated histogram may utilize data processing and / or mathematical operation methods used to merge two histograms of the same / similar attributes, including simple addition, simple multiplication, simple average, weighted summation, weighted average, selection of a maximum value for the same angle, selection of a minimum value for the same angle, and, but not limited to, simple addition, simple multiplication, simple average, weighted summation, weighted average, selection of a maximum value for the same angle, and selection of a minimum value for the same angle between histogram values ​​representing each angle. Additionally, a method of selectively applying at least one of the above-described operations based on an arbitrary threshold value may also be applied.

[0202] FIG. 13 is a conceptual diagram illustrating a method for generating a histogram through merging according to an embodiment of the present invention. Referring to FIG. 13, for estimating an intra-prediction mode for a current block (1310), a first histogram (1330) based on intra-prediction mode information of surrounding blocks (1321, 1322, 1323) and a second histogram (1340) based on slope analysis of a template area (1315) may each be generated. In the first histogram (1330), histogram values ​​corresponding to the angle of the directional mode of the surrounding blocks may be accumulated in proportion to the sample size of each block. In the second histogram (1340), histogram values ​​may be accumulated based on the angle and amplitude of the slope calculated at each sample location of the template area. Next, the first histogram (1330) and the second histogram (1340) can be combined into an integrated histogram (1360) through a merging operation (1350), and a final intra prediction mode (1370) can be determined from the integrated histogram (1360). In the example of FIG. 13, the locations of the surrounding blocks (1321, 1322, 1323) are exemplary, and it should be understood that in the present invention, any block unit that is directly adjacent to or not adjacent to the current block (1310) may be the target.

[0203]

[0204] Filter operation based on multiple reference lines

[0205] According to one embodiment of the present invention, when performing a filtering operation for gradient analysis in decoder-side intra-prediction mode estimation, more accurate and efficient prediction mode estimation is possible by using a variable-size filter that adaptively responds to the setting of multiple reference lines (MRL or TMRL). In conventional DIMD methods, gradients in a template area were mainly analyzed using a fixed-size filter (e.g., a 3x3 Sobel filter), but there was a problem in that when this method was combined with an intra-prediction method utilizing multiple reference lines, it could lead to inaccurate directional estimation due to a mismatch between the reference range and the filter size.

[0206] According to one embodiment of the present invention, when using a method for estimating an intra-prediction mode based on a histogram such as DIMD and / or OBIC, a filtering method, for example, the same or similar method as deriving HoG, is used to determine the amplitude of the histogram, and the size of the filter applied to the filtering may be variable.

[0207] According to one embodiment of the present invention, when a multi-reference line-based intra-prediction mode such as MRL and / or TMRL is allowed, the size of the filter may be varied based on the number of reference lines allowed for the current block in the MRL and / or TMRL method. For example, if the number of reference lines is determined to be one of 1, 3, 5, 7, or 12 pixel units, the size of the filter may be varied as 1x1, 3x3, 5x5, 7x7, or 12x12. That is, a filter having a size equal to or smaller than the multiple reference line-based template used for encoding by the MRL / TMRL method may be used. It is obvious that the coefficients of the filter according to the size may be determined as any one of various filter coefficient values ​​configured to allow a person skilled in the art to determine the slope angle through sample filtering, including the form of the Sobel filter described above.

[0208] According to the above embodiment, by configuring the filter size to increase as the number of reference lines increases, it is possible to perform slope analysis that reflects a wider range of texture information. For example, when using reference lines up to a distance of 7 pixels from the current block, a 3x3 filter alone may not sufficiently reflect the directionality of the reference lines at a distance, but by using a 7x7 filter, consistent directionality across the entire reference range can be captured.

[0209] The size of the filter does not necessarily have to be square, and for example, the filter may be configured such that only its width is variable, only its height is variable, or only the direction adjacent to the current block is fixed in terms of size (for example, when using multiple reference lines with a maximum distance of 7 pixels, a filter with 3 pixels horizontally and 7 pixels vertically is used when filtering the upper sample, and a filter with 7 pixels horizontally and 3 pixels vertically is used when filtering the left sample, etc.), or only the direction adjacent to the current block is variable in terms of size.

[0210] According to one embodiment of the present invention, different reference distances can be reflected in each direction by independently controlling the width and height of the filter. For example, if a reference line of 7 pixels is utilized in the upper direction but only a reference line of 3 pixels is utilized in the left direction, a filter of 7 pixels in the vertical direction is used when filtering the upper sample and a filter of 3 pixels in the horizontal direction is used when filtering the left sample, thereby matching the range of reference information that can actually be utilized in each direction with the filter size.

[0211] According to one embodiment of the present invention, the coefficients of the variable-size filter can be adaptively adjusted according to the filter size. For example, if the coefficients of a 3x3 Sobel filter are in the form of [-1, 0, +1; -2, 0, +2; -1, 0, +1], the coefficients of a 5x5 filter can be expanded in a form where the weight decreases as it moves away from the center, and a similar principle can be applied to sizes of 7x7 or larger. In addition, through normalization of the filter coefficients, the amplitude scale of the gradient can be configured to be consistently maintained even if the filter size changes.

[0212] The filter operation based on the above multiple reference lines can also be linked with the merged histogram generation method described above. For example, even if a neighboring block adjacent to the current block is encoded in a directional intra-prediction mode, the sample size of the neighboring block and the number of pixels of the reference line referenced by the MRL / TMRL can be compared to derive a histogram value based on the sample size only when the neighboring block is larger than the pixels of the reference line, or derive a histogram value based on the inverse condition.

[0213] According to one embodiment of the present invention, the linkage method may operate as follows. First, after checking the maximum distance of multiple reference lines set for the current block (e.g., 7 pixels), the size of a block encoded in a directional intra-prediction mode among the surrounding blocks may be evaluated. If the width or height of the surrounding block is greater than the maximum distance, the intra-prediction mode information of the block may be considered to represent consistent directionality over a wide area and may be included in the generation of a first histogram (HoC-based). Conversely, if the size of the surrounding block is smaller than the maximum distance, the information of the block may be excluded or reflected with a low weight.

[0214] In cases where the above multiple reference lines include at least one reference line that is not directly adjacent to each other, the filter used in the filter operation may be composed only of the values ​​of the reference lines actually used. For example, if the above multiple reference lines consist of four reference lines including a reference line adjacent to the current block at a distance of 1 pixel, a reference line adjacent at a distance of 3 pixels, a reference line adjacent at a distance of 5 pixels, and a reference line adjacent at a distance of 7 pixels, at least one of the width or height of the filter may be adjusted to be 4 or less. Alternatively, at least one of the width or height of the filter may be adjusted to be 7 or less, and a filter operation method may be applied in which a filter that does not have coefficient values ​​to be derived from unusable reference lines is used, or the values ​​of unusable reference lines are filled in by methods such as interpolation.

[0215] According to one embodiment of the present invention, the filter configuration method for the discontinuous reference line can operate as follows. First, a sparse filter can be configured by assigning non-zero filter coefficients only to the locations of the reference line actually used and assigning zero coefficients to the locations of the reference line that are not used. Second, the sample values ​​of the unused reference line can be estimated from the sample values ​​of adjacent available reference lines through linear interpolation, cubic interpolation, or other interpolation methods, and then a general filter operation can be performed. Third, the filter size can be reduced to match the number of reference lines actually used, and the mapping can be adjusted so that each filter coefficient is applied only to the samples of the corresponding reference line.

[0216] FIG. 14 is a conceptual diagram illustrating a method for applying a variable filter size based on multiple reference lines according to an embodiment of the present invention. Referring to FIG. 14, a case is illustrated in which a 1-pixel distance reference line (1421), a 3-pixel distance reference line (1422), a 5-pixel distance reference line (1423), and a 7-pixel distance reference line (1424) are set for the current block (1410). In this case, the size of the filter for slope analysis may vary for each pixel distance. For example, when corresponding to the maximum distance of the reference line, which is 7 pixels (1424), it may be set as a 7x7 filter (1434). At this time, the 7x7 filter (1434) is configured to calculate the slope by applying weights to samples within a 7-pixel range in each direction relative to a center sample (1434c), thereby enabling the capture of directional information across the entire reference range. Similarly, a 5x5 filter (1433, 1433c) corresponding to 5 pixels (1423) and a 3x3 filter (1432, 1432c) corresponding to 3 pixels (1422) can be set, and each filter can perform slope analysis within a range corresponding to the reference line distance.

[0217] Meanwhile, if only a 1-pixel distance reference line (1421) is used, the size of the filter may be applied as a 1x3 (1431-1) or 3x1 (1431-2) filter that shares a common center point (1431c) but is biased in the vertical / horizontal direction. It is obvious that such bias in the rectangular length is not limited to the above 1-pixel distance, but can be applied in various forms, such as, for example, a 5x7 filter or a 3x5 filter.

[0218]

[0219] Allowing estimation of non-directional intra-prediction mode

[0220] According to one embodiment of the present invention, in the estimation of the decoder-side intra prediction mode, not only the directional intra prediction mode but also the non-directional intra prediction mode (e.g., planar or DC prediction mode) is included as a target for histogram-based estimation, thereby improving the prediction accuracy for flat regions or regions with unclear directionality. In conventional DIMD and OBIC methods, mainly only the angle information of the directional intra prediction mode was reflected in the histogram; however, this method posed a risk of selecting an inappropriate directional mode when the texture was flat or when multiple directions were mixed.

[0221] According to one embodiment of the present invention, when using a method for estimating an intra-prediction mode based on a histogram such as DIMD and / or OBIC, the method may be configured to generate a histogram that includes not only the angle of the directional intra-prediction mode but also other modes or information related to intra-prediction. For example, the histogram may include a value representing at least one non-directional intra-prediction mode (e.g., Planar or DC prediction mode).

[0222] In this case, operations such as selecting a non-directional intra-prediction mode by another decoder-side estimation method that includes calculations for a surrounding sample region (e.g., a template region) can be performed by omitting the cost calculation and comparison procedure based on TIMD or generating a histogram (e.g., HoG or HoC) as described above. In this case, by further omitting the intra-mode prediction process on the decoder side as described above, the effect of further reducing the amount of computation during encoding / decoding and the amount of information in the encoded bit sequence can be achieved.

[0223] According to one embodiment of the present invention, the following method may be used to assign histogram values ​​to the non-directional intra prediction mode. According to one embodiment, for a surrounding sample area of ​​the current block, e.g., a template area, the complexity of the texture of said area may be calculated, and a histogram amplitude value representing said plane or DC prediction mode may be assigned based on said complexity. The method for calculating the complexity of said texture may be applied in various ways, but may be calculated in the form of a single function applied to said samples, a matrix operation calculated for said samples, and / or a filter applied to said samples. For instance, a filter that is easy to measure the flatness of data values, such as a Laplacian filter, may be used to calculate the complexity of said texture.

[0224] According to one embodiment of the present invention, the Laplacian filter is a second-order derivative operator and can be used to determine the flatness of a region by measuring the change in the rate of change at each sample position. For example, the coefficients of a 3x3 Laplacian filter may be in the form of [0, -1, 0; -1, 4, -1; 0, -1, 0] or [-1, -1, -1; -1, 8, -1; -1, -1, -1], and the smaller the absolute value of the filter response, the flatter the region. After applying the Laplacian filter to all samples of the template region, the mean or standard deviation of the response values ​​can be calculated and used as the texture complexity. If the calculated complexity is below a certain threshold, the histogram value for the non-directional intra prediction mode can be set high, and if it is above the threshold, it can be set low.

[0225] According to one embodiment of the present invention, when the surrounding blocks of the current block are encoded in a non-directional intra-prediction mode, the histogram value for the non-directional intra-prediction mode may be configured to be specified based on the sample size of the surrounding blocks. This method may correspond to the same or similar method as the method for calculating the histogram value of the HoC described above, except that the target is a non-directional mode independent of angle.

[0226] According to one embodiment of the present invention, for example, if a neighboring block located at the top of the current block is encoded in a planar mode and the size of the block is 16x16, a weight of 256 (16x16) can be added to the histogram value for the planar mode. Similarly, if a neighboring block located to the left is encoded in a DC mode and the size of the block is 8x8, a weight of 64 (8x8) can be added to the histogram value for the DC mode. In this way, the non-directional intra-prediction mode information of the neighboring blocks is also reflected in the histogram, so that when the current block is located in a flat area, the likelihood of the non-directional mode being selected naturally increases.

[0227] According to one embodiment of the present invention, when a non-directional intra-prediction mode is incorporated into the histogram-based intra-prediction mode estimation as described above, it will be readily understood that when intra-prediction fusion is operated as exemplified by DIMD, OBIC, and / or TIMD, the procedures utilizing directional intra-prediction modes determined based on the histogram in the above-described embodiment may all be modified or combined with procedures utilizing non-directional intra-prediction modes determined based on the histogram according to the above-described improved implementation method. For example, any method of predictive merging prediction blocks derived from two directional prediction modes may be modified to include the possibility of predictive merging prediction blocks derived from one non-directional prediction mode and one directional prediction mode, or two non-directional prediction modes, based on the above-described improved histogram generation method. In such cases, the procedure that conventionally required non-directional prediction modes to be included in the calculation may be removed or retained at the discretion of the implementer, and the present invention may include both implementation methods.

[0228] According to one embodiment of the present invention, an encoder and / or decoder may be configured in a manner that directs the unconditional or preferential use of a non-directional intra prediction mode based on calculations regarding a surrounding block or a surrounding sample region of a current block, e.g., a template region. For example, regarding a surrounding sample region of a current block, e.g., a template region, the complexity of the texture of said region may be calculated, and based on said complexity, the method may be configured to determine whether at least one of at least one non-directional intra prediction mode is used unconditionally or preferentially for the current block. For instance, if said complexity is calculated to be below a certain threshold level, the surrounding block of the current block may be considered to be generally flat, and either a planar or DC prediction mode may be automatically applied to the current block, or a method of prediction fusion including at least one non-directional intra prediction mode may be automatically applied.

[0229] According to one embodiment of the present invention, the method for determining the priority of use may operate as follows. First, the texture complexity calculated for a template region may be compared with a first threshold. If the complexity is less than or equal to the first threshold, the histogram generation and gradient analysis procedures may be omitted, and a planar mode or a DC mode may be selected immediately. If the complexity exceeds the first threshold but is less than or equal to a second threshold (greater than the first threshold), a histogram may be generated, and additional weights may be applied to the non-directional mode to increase the likelihood of selection. If the complexity exceeds the second threshold, a general histogram-based directional mode estimation procedure may be performed.

[0230] According to the embodiment, the size of the surrounding area for calculating the complexity of the texture may be variable. For example, the size of the surrounding area (e.g., the pixel size of the template) may vary depending on the size of the current block. Additionally, when multiple reference lines are set by MRL / TMRL, multiple surrounding areas (i.e., multiple templates) may be set considering the multiple reference lines, and the improved method described above may be configured to operate based on the complexity of the texture derived from the multiple surrounding areas.

[0231] According to one embodiment of the present invention, for example, when the current block size is 4x4, the width and height of the template area can be set to 1 pixel each, and when it is 8x8, 2 pixels, and when it is 16x16, 3 pixels, and so on, the template size can be adjusted in proportion to the block size. In addition, when reference lines at distances of 1 pixel, 3 pixels, and 5 pixels are set by MRL / TMRL, the template area corresponding to each reference line can be individually set to calculate the texture complexity for each. In this case, the average, minimum value, or weighted average of the complexity calculated from multiple templates can be used to determine whether to select the final non-directional mode.

[0232] FIG. 15 is a conceptual diagram illustrating a histogram generation method including a non-directional intra prediction mode according to an embodiment of the present invention. Referring to FIG. 15, a texture complexity calculation (1520) for a template area (1515) may be performed to estimate the intra prediction mode for the current block (1510). In the complexity calculation (1520), a Laplacian filter or other flatness measurement method may be applied, and a complexity value may be derived as a result. When the complexity value is low, a histogram value (1541) for a non-directional mode (which may include a plane and / or DC) may be set high. At the same time, a histogram value (1542) for directional modes may also be calculated through a directional gradient analysis (1530). Finally, an integrated histogram (1550) including both the non-directional mode value (1541) and the directional mode value (1542) can be generated, and from this, the final intra-prediction mode (1560) can be determined.

[0233]

[0234] Memory reusability

[0235] According to one embodiment of the present invention, various intermediate and final results calculated during the decoder-side intra-prediction mode estimation process are stored in memory and reused for encoding / decoding of subsequent blocks, thereby reducing the overall amount of computation and improving processing efficiency. In conventional DIMD and OBIC methods, histogram generation and prediction mode estimation were performed independently for each block; however, due to high spatial and temporal correlations between adjacent blocks, redundant computations often occurred.

[0236] According to one embodiment of the present invention, operation values ​​for surrounding blocks used in the various embodiments of the present invention described above, for example, input values ​​and output results of filter operations, results of references and / or calculations for template regions (including multiple templates by multiple reference lines by MRL / TMRL), histogram values ​​for prediction modes, and at least a portion of said histogram may be stored in memory and reused for subsequent encoding. The reuse may mean reuse based on at least one of spatial overlap or temporal overlap. For example, they may be configured to be stored in memory for a certain period of time to be used for encoding spatially subsequent pictures, blocks, and / or samples, or for encoding temporally subsequent pictures, blocks, and / or samples (and / or in the encoding / decoding order in bit sequences).

[0237] According to one embodiment of the present invention, reuse based on spatial overlap can operate as follows. After storing the histogram, gradient analysis results, and sample values ​​of the template area calculated for the current block in memory, the information can be reused during the encoding / decoding of the next block (e.g., the right block or the bottom block) that is spatially adjacent to the current block. In particular, if the template area of ​​the next block partially overlaps with the template area of ​​the current block, redundant calculations can be prevented by reusing the gradient analysis results for the overlapping area. For example, the next block located to the right of the current block may include the left template area of ​​the current block as part of its own template area, and in this case, the gradient value for that area can be read from memory and used directly.

[0238] According to one embodiment of the present invention, reuse based on temporal overlap may operate as follows. Histogram and prediction mode information calculated for a specific block location in the current frame may be stored in memory, and then utilized as an initial value or reference value during the encoding / decoding of a block located at the same or similar location in a temporally subsequent frame. In particular, for static regions in a video sequence where there is little texture change between consecutive frames, the histogram calculated in the previous frame may have high validity in the current frame as well. In this case, the generation of the histogram in the current frame may be omitted or simplified, and the histogram of the previous frame may be used directly or partially updated.

[0239] According to one embodiment of the present invention, a storage structure for memory reuse may be configured as follows. First, the histogram calculated for each block may be stored in the form of a two-dimensional array, and the location of the block (e.g., relative coordinates within a coding tree unit) may be used as an index. Second, the gradient analysis results may store the gradient angle and amplitude calculated for each sample location, and may be managed as a global array in the form of frames or slices. Third, the sample values ​​in the template area may be managed with minimal memory by utilizing a line buffer structure, and this may be shared with the reference sample buffer used in conventional intra prediction.

[0240] According to one embodiment of the present invention, conditions for determining the validity of memory reuse may be set as follows. First, reuse may be allowed only if the number of blocks or time elapsed since the creation of the stored information is below a certain threshold. Second, reuse may be allowed only if the spatial distance between the current block and the block that created the stored information is within a certain range. Third, reuse may be allowed only if the maximum peak value of the stored histogram is above a certain threshold (i.e., if a clear directionality exists). Fourth, reuse may be allowed only if the size, quantization parameters, or other encoding settings of the current block are similar to those of the block that created the stored information.

[0241] FIG. 16 is a conceptual diagram illustrating a memory reuse structure according to an embodiment of the present invention. Referring to FIG. 16, a histogram (1615), a gradient analysis result (1616), and a template sample value (1617) calculated for a block A (1610) encoded / decoded in a specific frame (1691) may be stored in memory (1620) in conjunction with the encoding / decoding process. Next, if a block B (1630) currently being encoded / decoded is spatially adjacent to block A (1610) within the same frame (1691) (1630-1) or is located in the same / adjacent position in a temporally different frame (1692) (1630-2), information stored from memory (1620) may be read and utilized for estimating the intra-prediction mode of block B (1630). For example, if the template area of ​​Block B overlaps with the template area of ​​Block A, the histogram (1615) of Block A can be used as an initial value or a standard for weight adjustment for generating the histogram of Block B. Alternatively, the slope analysis for the overlapping area can be omitted, and the stored slope value (1616) can be used directly. Alternatively, the template area information (1617) of Block A can be used in parallel, combined, and / or replaced with the template area corresponding to Block B.

[0242]

[0243] IBC priority processing

[0244] According to one embodiment of the present invention, when a neighboring block is encoded in the Intra Block Copy (IBC) mode during the decoder-side intra prediction mode estimation process, the prediction accuracy can be improved and the amount of computation reduced by prioritizing IBC-based intra prediction over histogram-based estimation methods such as general DIMD, OBIC, and TIMD. The IBC mode is a method of prediction that references an already decoded region within the same frame as the current block, and can provide high compression efficiency in regions where textures are repeated or patterns are similar. In particular, since the existence of a neighboring block encoded in IBC mode suggests a high probability that a similar texture exists near the current block, it is reasonable to prioritize this.

[0245] According to one embodiment of the present invention, when applying various embodiments of the present invention described above, if an intra-block copy (IBC) mode is used in a surrounding block of the current block, the system may be configured to prioritize intra-mode prediction by referring to the block.

[0246] According to one embodiment of the present invention, when at least one block among the surrounding blocks is intra-predicted in IBC mode, a method of calculating the cost of intra-predicted based on IBC in advance may be applied prior to the decoder-side intra-predicted mode estimation methods including DIMD, OBIC, and TIMD. According to the embodiment, the method of calculating in advance may be applied equally in the encoder and the decoder, or it may be configured to reduce the capacity of the bit sequence by automatically determining whether to use the IBC-based intra-predicted mode preferentially in the decoder-side after being applied in the encoder, or it may be configured to reduce complexity in the decoder-side by reading a bit sequence flag indicating whether to use the IBC-based intra-predicted mode generally or optionally before a bit sequence flag indicating whether to use the decoder-side intra-predicted mode estimation.

[0247] According to one embodiment of the present invention, the IBC priority processing method may be composed of the following steps. First, surrounding blocks of the current block (e.g., top, left, top-left, top-right, bottom-left blocks) may be examined to determine if there is a block encoded in IBC mode. Second, if a surrounding block encoded in IBC mode is found, a Block Vector (BV) referenced by the block may be applied to the current block to generate a prediction block. Third, a prediction cost may be calculated by comparing the generated prediction block with the template area of ​​the current block. Fourth, if the calculated cost is below a certain threshold, other histogram-based intra-prediction mode estimation procedures may be omitted, and IBC-based prediction may be used immediately. Fifth, if the calculated cost exceeds the threshold, general procedures such as DIMD, OBIC, and TIMD may be performed, but the IBC-based prediction result may also be included as one of the candidate modes and considered in the final selection process.

[0248] According to one embodiment of the present invention, when a plurality of surrounding blocks are encoded in IBC mode, a plurality of IBC-based prediction blocks can be generated by considering all of their respective block vectors. In this case, the cost of each prediction block can be calculated individually, and then the block vector with the lowest cost can be selected, or a final prediction block can be generated by weighted averaging the prediction blocks corresponding to the plurality of block vectors. The weights can be set to be inversely proportional to the cost of each prediction block, for example, a higher weight can be assigned as the cost decreases.

[0249] According to one embodiment of the present invention, the calculation of the predicted cost can be performed in various ways. In one embodiment, the sum of absolute differences (SAD) or the sum of squared differences (SSD) between the sample values ​​of the region referenced by the block vector and the sample values ​​of the template region of the current block can be calculated. In another embodiment, the sum of absolute transformed differences (SATD) in the transform domain can be calculated similarly to that used in the TIMD described above. In yet another embodiment, a template matching technique can be optionally applied to perform a more sophisticated cost calculation, but this may be limited to an optional embodiment as it may lead to an increase in computational complexity.

[0250] According to one embodiment of the present invention, the encoder may determine the optimal intra-prediction mode through IBC priority processing and then include only a minimum number of signals in the bit sequence so that the decoder can make the same decision. For example, only a 1-bit flag indicating whether the current block uses IBC-based prediction may be transmitted, and the specific block vector may be configured to be automatically derived by the decoder from surrounding blocks. Alternatively, the decoder may automatically determine the prediction mode of the current block based on the IBC information of surrounding blocks, thereby omitting even the flag, in which case the capacity of the bit sequence is further reduced.

[0251] According to one embodiment of the present invention, the decoding complexity can be reduced by optimizing the order in which flags related to IBC priority processing are read from a bit sequence in a decoder. Specifically, by configuring the decoder to read IBC priority processing flags before general decoder-side intra-prediction mode estimation flags, subsequent complex histogram generation and template matching procedures can be terminated early when IBC-based prediction is used. This can provide a significant reduction in complexity, particularly in Screen Content Coding (SCC) scenarios where IBC mode is frequently used.

[0252] FIG. 17 is a flowchart illustrating an IBC priority processing procedure according to an embodiment of the present invention. Referring to FIG. 17, along with the intra-prediction mode estimation (1710) of the current block, it can be checked whether the surrounding block is encoded in IBC mode (1720). If there is no surrounding block encoded in IBC mode, a general mode prediction procedure (1770) (e.g., may include at least one process among DIMD / OBIC / TIMD) can be performed. If there is a surrounding block encoded in IBC mode, the block vector of the block can be applied to the current block to generate an IBC-based prediction block (1730). Next, the generated prediction block and the template area can be compared to calculate the prediction cost (1740). The calculated cost can be determined (1750), and if it is below a threshold, the IBC-based prediction can be used (1760) and the procedure can be terminated (1780). If the calculated cost exceeds a threshold, the general mode prediction procedure (1770) is performed by including the IBC-based prediction result as a candidate, and finally, the optimal intra prediction mode can be selected (1780).

[0253]

[0254] Encoder and decoder

[0255] It is evident that the method according to the present invention can be applied equally to an encoder and a decoder. Additionally, as illustrated in FIG. 4 through an internal decoder (420) and a coding loop including it, this decoding process can be implemented equally within the encoder to predict the state of the decoder.

[0256] The encoding method according to the present invention described above can be implemented through an encoder as a device. The encoder as a device may be implemented by maintaining the conventional encoder structure exemplified in FIGS. 1 to 6 or by applying a certain change therefrom, but the form of implementation is not necessarily limited to that exemplified, and any form of encoder structure capable of functioning as a video encoder should be considered an encoder established by the present invention as long as it embodies the technical concept of the present invention.

[0257] In addition, the decoding method of the encoded result according to the present invention described above can be implemented through a decoder as a device. The decoder as a device may be implemented in a form that maintains the conventional decoder structure exemplified in FIGS. 1 to 6 or applies a certain change therefrom, but the form of implementation is not necessarily limited to that exemplified, and any form of decoder structure capable of functioning as a video decoder should be considered a decoder established by the present invention as long as it embodies the technical concept of the present invention.

[0258] A person skilled in the art will readily understand that a bit sequence encoded by the method and apparatus described above can be decoded by applying a method symmetric and / or in reverse order to the encoding method. In one embodiment, when reading information for decoding from the encoded bit sequence, at least one variable-length coded phrase included in the encoded bit sequence may be interpreted, and in one embodiment, the variable-length coding may be performed by an entropy coding method. The technical details and application methods of implementing such a decoding procedure will be readily understood from the encoding procedure described above.

[0259] The processor that may be included in the encoder and / or decoder described herein may mean one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions.

[0260] Even if the above processor is expressed in the singular for ease of understanding, a person skilled in the art will understand that the above processor may include a plurality of processing elements and / or a plurality of types of processing elements. For example, an apparatus according to one embodiment of the present invention may include a plurality of processors or one processor and one controller as the processor. In addition, the processor may be implemented by various processing configurations, such as a parallel processor or a multi-core processor.

[0261] The processor may be configured to execute an operating system (OS) and one or more software executed on the operating system. Additionally, the processor may access, store, manipulate, process, and generate data in response to the execution of the software.

[0262] The software may include a computer program, code, instructions, or a combination of one or more of these, and may be configured to control the processor to operate as desired and to issue instructions to the processor independently or collectively. The software may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave in order to be interpreted by the processor or to provide instructions or data to the processor. The software may be distributed over networked computer systems and may be stored or executed in a distributed manner.

[0263] The software described above may also be implemented in the form of program instructions that can be executed through various computer means and may be recorded or stored in the memory. The memory may be a computer-readable recording medium, and program instructions, data files, data structures, etc., may be recorded on the computer-readable recording medium alone or in combination. The program instructions stored in the memory may be based on a command system specifically designed and configured for embodiments of the present invention, or may follow a command system known and available to those skilled in the art of computer software, such as Assembly, C, C++, Java, Python, etc. It should be understood that the command system and the program instructions derived therefrom include not only machine code such as that generated by a compiler, but also high-level language code that can be executed by a device and / or processor according to an embodiment of the present invention using an interpreter, etc.

[0264] A computer-readable recording medium constituting an device according to an embodiment of the present invention, including the memory described herein, may include a temporary or volatile recording medium that is maintained only while the processor is operating, such as a processor cache, RAM, or flash memory; or may include a relatively non-volatile or long-term recording medium such as a magnetic media such as a hard disk, floppy disk, and magnetic tape; an optical recording medium such as a CD-ROM or DVD; a magneto-optical media such as a floptical disk; or a solid-state memory; or may include a read-only recording medium such as a ROM placed on hardware; furthermore, the hardware itself, configured to perform operations equivalent to a series of program instructions by a hard-wired structure by circuit wiring, may also be considered as having each step for performing the operation implementing the embodiment of the present invention recorded by the connection and arrangement of the hardware components, so the method of connection and arrangement is equivalent to the memory. It is obvious to an ordinary technician that it can be seen.

[0265] The embodiments described above with respect to the processor and the memory are not mutually exclusive and may be selected or combined as needed. For example, a hardware device may be configured to operate as a module composed of one or more of the software to perform the operation of an embodiment of the present invention, and vice versa. As another example, in this specification, all or part of the operation assigned to a certain functional unit may be implemented by one or more of the software stored in a device according to an embodiment of the present invention (preferably in a recording medium falling within the category of the memory) and configured to be executed by the processor, in which case such a functional unit may be referred to as a functional unit "included" in the processor.

[0266]

[0267] Although the present invention has been described above with reference to the drawings and embodiments, as previously stated, the scope of protection of the present invention is not limited by the drawings or embodiments presented above, and those skilled in the art will understand that various modifications and changes can be made to the present invention without departing from the spirit and scope of the invention as described in the claims of the present invention.

Claims

1. In a video decoding method, A step of referencing decoded pixel values ​​of an area adjacent to the current block to perform intra prediction for the current block; A step of calculating evaluation values ​​for a plurality of candidate intra prediction modes to be applied to the intra prediction of the current block based on the above decoded pixel values; A step of selecting at least one intra prediction mode among the plurality of candidate intra prediction modes based on the above evaluation value; and The method includes the step of decoding the current block using the selected intra-prediction mode; The above plurality of candidate intra prediction modes include at least one directional intra prediction mode, and A video decoding method characterized in that the step of calculating the above evaluation value is performed based on at least one of intra prediction information of at least one surrounding block spatially adjacent to the current block, the result of a filter operation on the decoded pixel values, and the result of a cost calculation between the decoded pixel values ​​and the predicted value by the candidate intra prediction mode.

2. In Paragraph 1, The step of calculating the above evaluation value is, A step of deriving a first evaluation value based on the intra-prediction mode of the surrounding block and the size of the surrounding block; A step of deriving a second evaluation value based on the above filter operation result; and A video decoding method characterized by including the step of merging the first evaluation value and the second evaluation value to calculate a final evaluation value.

3. In Paragraph 2, An image decoding method characterized in that the step of deriving the first evaluation value is to add a value proportional to the number of pixels of the surrounding block to the evaluation value corresponding to the directional intra prediction mode when the surrounding block is encoded in a directional intra prediction mode.

4. In Paragraph 1, The above filter operation includes gradient analysis for the above decoded pixel values, and An image decoding method characterized by the above-described slope analysis calculating the horizontal slope and vertical slope of the decoded pixel values, deriving a slope angle from the horizontal slope and the vertical slope, and increasing the evaluation value of a directional intra-prediction mode corresponding to the slope angle.

5. In Paragraph 4, An image decoding method characterized in that the above gradient analysis is performed using a Sobel filter.

6. In Paragraph 4, An image decoding method characterized in that the decoded pixel values ​​to which the above gradient analysis is applied include pixel values ​​on a plurality of reference lines located at different distances from the current block.

7. In Paragraph 6, An image decoding method characterized in that the size of the filter used in the slope analysis is determined based on the number of reference lines or the maximum distance between the reference lines.

8. In Paragraph 1, An image decoding method characterized in that the plurality of candidate intra prediction modes further include at least one non-directional intra prediction mode.

9. In Paragraph 8, The method further includes the step of calculating texture complexity for the above-decoded pixel values; An image decoding method characterized in that the step of calculating the above evaluation value calculates an evaluation value for the above non-directional intra prediction mode based on the above texture complexity.

10. In Paragraph 9, An image decoding method characterized by the step of calculating the above evaluation value selecting the non-directional intra-prediction mode without performing the filter operation and the cost calculation when the texture complexity is below a preset threshold.

11. In a video encoding method, A step of referencing pixel values ​​reconstructed in a prior encoding process as an area adjacent to the current block to perform intra prediction for the current block; A step of calculating evaluation values ​​for a plurality of candidate intra prediction modes to be applied to the intra prediction of the current block based on the above-mentioned reconstructed pixel values; A step of selecting at least one intra prediction mode among the plurality of candidate intra prediction modes based on the above evaluation value; and The method includes the step of encoding the current block using the selected intra-prediction mode, wherein The above plurality of candidate intra prediction modes include at least one directional intra prediction mode, and A video encoding method characterized in that the step of calculating the above evaluation value is performed based on at least one of intra prediction information of at least one surrounding block spatially adjacent to the current block, the result of a filter operation on the reconstructed pixel values, and the result of a cost calculation between the reconstructed pixel values ​​and the predicted value by the candidate intra prediction mode.

12. In Paragraph 11, The step of calculating the above evaluation value is, A step of deriving a first evaluation value based on the intra-prediction mode of the surrounding block and the size of the surrounding block; A step of deriving a second evaluation value based on the above filter operation result; and A video encoding method characterized by including the step of merging the first evaluation value and the second evaluation value to calculate a final evaluation value.

13. In Paragraph 11, An image encoding method characterized in that the above filter operation is performed on reconstructed pixel values ​​on a plurality of reference lines located at different distances from the current block.

14. In Paragraph 11, An image encoding method characterized in that the plurality of candidate intra prediction modes further include at least one non-directional intra prediction mode.

15. In Paragraph 14, A step of calculating texture complexity for the above-mentioned reconstructed pixel values; and A video encoding method characterized by further including the step of omitting the step of calculating the evaluation value and selecting the non-directional intra-prediction mode when the texture complexity is less than a first threshold.

16. In Paragraph 11, A step of determining whether the selected intra-prediction mode can be automatically induced on the decoder side; and A video encoding method characterized by further including the step of determining whether to include the selected intra-prediction mode information in the bitstream based on the above-mentioned judgment result.

17. In a video decoder device, processor; Memory connected to the above processor; A reference unit that references decoded pixel values ​​of an area adjacent to the current block to perform intra prediction for the current block; A calculation unit that calculates evaluation values ​​for a plurality of candidate intra prediction modes to be applied to the intra prediction of the current block based on the above decoded pixel values; A selection unit for selecting at least one intra prediction mode among the plurality of candidate intra prediction modes based on the above evaluation value; and It includes a decoding unit that decodes the current block using the selected intra-prediction mode, The above plurality of candidate intra prediction modes include at least one directional intra prediction mode, and A decoder device characterized in that the above-described output unit calculates the evaluation value based on at least one of intra prediction information of at least one surrounding block spatially adjacent to the current block, the result of a filter operation on the decoded pixel values, and the result of a cost calculation between the decoded pixel values ​​and the predicted value by the candidate intra prediction mode.

18. In Paragraph 17, A decoder device characterized in that the above-described output unit derives a first evaluation value based on the intra-prediction mode of the surrounding block and the size of the surrounding block, derives a second evaluation value based on the filter operation result, and calculates a final evaluation value by merging the first evaluation value and the second evaluation value.

19. In Paragraph 17, The above memory stores the evaluation value calculated by the above calculation unit and the intra prediction mode information selected by the above selection unit, and A decoder device characterized in that the above-mentioned output unit refers to information stored in the memory when calculating the evaluation value of a subsequent block.

20. In Paragraph 17, A decoder device characterized in that the filter operation performed by the above-described output unit is performed on decoded pixel values ​​on a plurality of reference lines located at different distances from the current block.

Citation Information

Patent Citations

  • DMVR and BDOF-based inter-prediction method and apparatus

    JP2022092010A

  • In-line dumpling forming system

    KR1020220123920A

  • Construction Site Ground Displacement Monitoring Method and System

    KR1020250146054A

  • LED lights for underground parking lot

    KR102703915B1

  • A quadrupedal robot controllable with one hand

    KR102946544B1