Method and apparatus for encoding and decoding screen content and game content video

The adaptive filter control method addresses the degradation of screen and gaming content by selectively disabling filters during intra-prediction, enhancing encoding efficiency and visual quality.

WO2026095739A1PCT designated stage Publication Date: 2026-05-07KAON GRP CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
KAON GRP CO LTD
Filing Date
2025-11-04
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Conventional video encoding standards degrade the quality of screen content and gaming content by applying smoothing filters intended for natural images, leading to unnecessary computational complexity and inefficient filtering methods when mixing natural and graphic content.

Method used

An adaptive filter control method that selectively deactivates unnecessary filtering processes during intra-prediction based on the characteristics of screen and gaming content, using region information and hash block-based algorithms to identify and control filtering at various levels, such as sequence, picture, and CTU levels.

Benefits of technology

Reduces computational complexity while maintaining or improving compression performance, preserving sharp edges and visual quality of screen content, and enabling efficient encoding/decoding tailored to image characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025017868_07052026_PF_FP_ABST
    Figure KR2025017868_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention presents a video encoding method for identifying features of a computer graphic-processed video and using the corresponding features. The video generated by graphic processing exhibits clearly differing values between adjacent pixel values, and the corresponding features are features different from those of a video captured by a camera. The corresponding feature can be used in a video encoding method. The present invention is to improve encoding performance by using a method in which differing values between adjacent pixel values in a video generated by computer graphics during video encoding / decoding are smoothed.
Need to check novelty before this filing date? Find Prior Art

Description

Method and apparatus for encoding and decoding screen content and game content video

[0001] The present invention relates to the field of encoding and decoding of digital video, and to a method for encoding and decoding digital video, a method for recording such data, and components, devices, and systems for realizing such a method.

[0002] The present invention may correspond to a technical field identical to at least one of the digital video compression technology standards known by standard names such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG, a technical field for improving the inherent efficiency of the standard, or a technical field for improving or replacing the standard.

[0003] The present invention relates to an adaptive filter control method and apparatus capable of reducing encoding / decoding complexity while maintaining or improving compression performance by selectively deactivating unnecessary filtering operations during the intra-prediction process, taking into account the characteristics of images generated by computer graphics, such as screen content and gaming content.

[0004] Digital video encoding and decoding are widely utilized in various digital video applications. For example, devices such as video recording equipment and camcorders used for video recording activities—including digital television broadcasting, video transmission via communication networks, video calls, video conversations, and video chats, recording and provision of video content using optical media such as VCDs (video compact discs), DVDs (digital versatile discs), and Blu-rays, all procedures for the production, editing, collection, and distribution of video content, and video recording for personal, commercial, industrial, and security purposes—are all dependent on video encoding and decoding technology.

[0005] Accordingly, embodiments that can be referred to as digital video encoders and decoders may constitute a part of a wide range of devices related to the creation, recording, and provision of digital video, including digital television, digital broadcasting systems, wireless broadcasting systems, computers in the form of notebooks / desktops / tablets, e-book readers, digital cameras, digital recording devices, digital multimedia playback devices, video game devices / terminals / consoles, mobile phones equipped with multimedia playback functions (including smartphones), equipment for video conferencing, and other devices.

[0006] Digital video encoders and decoders as described above can be implemented by digital video compression standards that are understood by and widely used by people skilled in the art. The digital video compression standards may include at least one of the compression standards known by standard names such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.

[0007] Video encoders and decoders can be implemented to encode or decode digital video information more efficiently while complying with the above specifications, or by improving or modifying them. Attempts to modify the above specifications may also lead to the development of new specifications. Among well-known examples is the so-called enhanced compression model (ECM), which is an attempt to improve and replace the conventional H.266 / VVC specifications, currently being developed by the Joint Video Experts Team (JVET), a joint international standardization group of ISO, IEC, and ITU-T.

[0008] Meanwhile, as part of recent advancements in video encoding technology, attempts are being made to improve the compression efficiency of videos containing screen content, such as text or graphics, and videos containing gaming content. Unlike natural images, the aforementioned screen content images may contain edges occurring in arbitrary directions rather than horizontal or vertical directions, and are characterized by a complex mixture of linear and curved components within the image. Furthermore, the aforementioned gaming content images may exhibit characteristics of both natural images and screen content simultaneously. For example, gaming content may contain a mixture of parts with characteristics of natural images, such as in-game characters or backgrounds, and parts with characteristics of screen content, such as in-game interfaces or text. Consequently, it has been pointed out that smoothing-related processes designed for natural images, such as Reference Sample Filtering, Position Dependent Prediction Combination (PDPC), and Intra Prediction Fusion, may actually degrade the sharp edge characteristics of screen content.

[0009] Conventional video encoding standards are designed primarily for the compression of natural images captured by cameras and include various filtering techniques that consider the continuous distribution of pixel values ​​and noise characteristics. However, since screen content and gaming content are generated by computer graphics and have characteristics where pixel values ​​are concentrated at specific values ​​and edges are distinct, smoothing filters intended for natural images can actually degrade image quality and induce unnecessary computational complexity. Furthermore, until now, there has been a lack of an explicit method to distinguish whether an image is screen content or natural image, which has resulted in a problem where encoders and decoders cannot select a processing method suitable for the image characteristics.

[0010] Moreover, in cases where natural video and graphic content are mixed, such as news subtitles, home shopping information displays, and presentation materials, uniformly applying or disabling filters to the entire video is inefficient. This is because while the quality of natural video areas within the video can be improved through filtering, the quality of screen content areas can be degraded as edges become blurred due to filtering. Therefore, a method is needed to selectively control filters for specific areas within the video, but an efficient method for area selection and signaling for this purpose has not been established.

[0011] Furthermore, while hash block-based screen content identification algorithms introduced in the latest video codecs can automatically identify screen content through pixel hit rates, a method to effectively integrate this with filter control in intra-prediction has not been systematized. In particular, there is a need for an integrated framework that identifies screen content at various levels—such as the sequence, picture, and CTU levels—and adaptively controls filters accordingly.

[0012] The present invention identifies the unique characteristics of images processed by computer graphics and presents an efficient image encoding method utilizing these characteristics. Images generated by graphic processing are characterized by clearly visible differences between adjacent pixel values, a feature that distinguishes them from natural images captured by a camera. Considering these characteristics, the present invention can improve encoding performance by selectively disabling the filtering process during the image encoding / decoding process to prevent the sharp boundaries of pixel values ​​within the computer graphics-generated image from being degraded by smoothing.

[0013] A video decoding method according to an embodiment of the present invention for solving the aforementioned technical problem may include: a step of obtaining region information indicating that at least a portion of a current image is an arbitrary region generated by computer graphics from a bitstream; a step of identifying at least one filtering process applied in a prediction process of a current block corresponding to the arbitrary region based on the region information; a step of generating a prediction block for the current block by selectively deactivating the identified filtering process; and a step of restoring the current block using the prediction block.

[0014] The above production area corresponds to a part of the above current image, and may further include the step of additionally obtaining area information indicating the location and size of the said part of the area from the bitstream.

[0015] The above area information may be characterized by specifying the above partial area using a Coding Tree Unit (CTU) index.

[0016] The above area information may be characterized by designating the above partial area using a tile index.

[0017] The above method may be characterized in that the production area is multiple, and area information for each production area is individually signaled.

[0018] The filtering process described above may include at least one of reference sample filtering, location-based prediction combination (PDPC), and intra prediction fusion, and may be characterized in that the at least one filtering process is integrally controlled by a single flag.

[0019] The above method may further include the step of individually controlling at least one of the at least one filtering process based on an additional flag.

[0020] The above method may further include a step of automatically determining whether the current block belongs to the production area using a hash block-based algorithm, and may be characterized by determining it as the production area when the hit percentage of the hash block exceeds a predetermined threshold.

[0021] The filtering process may include at least one of overlapping block motion compensation (OBMC), geometric segmentation mode (GPM) blending, and bidirectional optical flow (BDOF).

[0022] The above method may be characterized in that the deactivation of the filtering process is adaptively determined at the coding unit (CU) level.

[0023] In the above-mentioned production area, at least one of the transform skip mode and the intra-block copy (IBC) mode may be applied preferentially.

[0024] The above area information may be characterized by being signaled in at least one bit sequence syntax unit among a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH).

[0025] The above method may be characterized in that, when a dual-tree structure is applied, the filtering process is controlled independently for the chroma block.

[0026] A video encoding method according to an embodiment of the present invention for solving the aforementioned technical problem may include: a step of determining whether at least a portion of a current image is an arbitrary region generated by computer graphics; a step of identifying at least one filtering process applied in a prediction process of a current block corresponding to the arbitrary region based on the determination result; a step of generating a prediction block for the current block by selectively deactivating the identified filtering process; a step of encoding the current block using the prediction block; and a step of including region information indicating that it is an arbitrary region in a bitstream.

[0027] The above production area corresponds to a part of the above current image, and may further include the step of including area information indicating the location and size of the said part of the area in the bitstream.

[0028] The above filtering process may include at least one of reference sample filtering, location-based prediction combination (PDPC), and intra-prediction fusion, and may be characterized by controlling the at least one filtering process with a single flag.

[0029] The above method may further include the step of determining the production area by calculating a hash value in units of 4x4 or 8x8 blocks.

[0030] An image decoding device according to an embodiment of the present invention for solving the aforementioned technical problem may include: a parsing unit configured to acquire region information indicating that at least a portion of a current image is an arbitrary region generated by computer graphics from a bitstream; a prediction unit configured to identify at least one filtering process applied in a prediction process of a current block corresponding to the arbitrary region based on the region information, and to selectively disable the identified filtering process to generate a prediction block for the current block; and a restoration unit configured to restore the current block using the prediction block.

[0031] The above-mentioned production area corresponds to a part of the above-mentioned current image, and the above-mentioned parsing unit may be configured to additionally acquire area information indicating the location and size of the above-mentioned part of the area.

[0032] The filtering process described above includes at least one of reference sample filtering, location-based prediction combination (PDPC), and intra-prediction fusion, and the prediction unit may be configured to integrally control the at least one filtering process based on a single flag.

[0033] According to the present invention, by selectively deactivating unnecessary filtering operations during the intra prediction process by considering the pixel distribution characteristics of screen content and gaming content, the computational complexity of encoders and decoders can be significantly reduced while maintaining or improving compression performance. In particular, by omitting smoothing processes such as reference sample filtering, PDPC, and intra prediction fusion, sharp edges and clear color boundaries of screen content can be preserved, which can directly contribute to improving the readability and visual quality of text or graphic elements.

[0034] Furthermore, the present invention provides a systematic method for efficiently signaling screen content information for all or part of an image, thereby enabling the application of optimal encoding / decoding methods tailored to the characteristics of each area even in complex images where natural images and graphic content are mixed. The area designation method using a CTU index or a tile index enables precise area control with minimal signaling overhead while remaining compatible with the structure of existing video codecs, thereby realizing efficient video compression in various application fields such as broadcasting, streaming, and video conferencing.

[0035] Furthermore, the combination of the hash block-based automatic determination mechanism and explicit signaling of the present invention can guarantee a standardized decoding process while providing implementation flexibility for the encoder. Adaptive filter control at various layers, such as sequences, pictures, and CTUs, can effectively respond to dynamic changes in content, and can simultaneously achieve consistent image quality and efficient compression, particularly in application fields where the proportion of screen content changes over time, such as real-time broadcasting or game streaming.

[0036] FIG. 1 is a conceptual diagram of a video communication system according to an embodiment of the present invention,

[0037] FIG. 2 is a conceptual diagram of the arrangement of an encoder and a decoder in a real-time video streaming environment according to an embodiment of the present invention.

[0038] FIG. 3 is a conceptual diagram of a functional unit of a video decoder according to an embodiment of the present invention,

[0039] FIG. 4 is a conceptual diagram of a functional unit of a video encoder according to an embodiment of the present invention,

[0040] FIG. 5 is a conceptual diagram of a frame type according to an embodiment of the present invention,

[0041] FIG. 6 is a conceptual diagram showing the structure of a video encoder according to another embodiment of the present invention,

[0042] FIG. 7 is an exemplary diagram of the configuration of screen content and gaming content according to an embodiment of the present invention.

[0043] FIG. 8 is a conceptual diagram of the difference in pixel distribution between a real-life image and screen content in the present invention,

[0044] FIG. 9 is a conceptual diagram illustrating the directional prediction mode of VVC,

[0045] FIG. 10 is a conceptual diagram of the application of a reference sample to an encoding target block in the present invention.

[0046] FIG. 11 is a conceptual diagram of the process of generating reference samples according to an embodiment of the present invention, and

[0047] FIG. 12 is a flowchart of a manufacturing area processing according to one embodiment of the present invention.

[0048] The present invention is capable of various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the invention to specific embodiments, and it should be understood that the invention includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the invention.

[0049] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. These terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of multiple related described items or any one of the multiple related described items, and is non-exclusive unless otherwise indicated. When items are listed in this application, they are merely illustrative descriptions intended to facilitate the explanation of the spirit of the present invention and possible methods of implementation, and are therefore not intended to limit the scope of the embodiments of the present invention.

[0050] In this specification, "A or B" may mean "only A," "only B," or "both A and B." Alternatively, in this specification, "A or B" may be interpreted as "A and / or B." For example, in this specification, "A, B or C" may mean "only A," "only B," "only C," or "any combination of A, B and C."

[0051] A slash ( / ) or a comma used in this specification may mean "and / or." For example, "A / B" may mean "A and / or B." Accordingly, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B or C."

[0052] In this specification, "at least one of A and B" may mean "only A," "only B," or "both A and B." Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted as synonymous with "at least one of A and B."

[0053] Additionally, in this specification, "at least one of A, B and C" may mean "only A," "only B," "only C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."

[0054] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0055] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "having" are intended to specify the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0056] Unless otherwise defined, all terms used herein, including technical or scientific terms, are used with the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.

[0057] In describing the invention in this application, embodiments may be described or illustrated in terms of the described functions or unit blocks that perform the functions. The blocks may be expressed in this application as one or more devices, units, modules, parts, etc. The blocks may be implemented in hardware by a method of implementing one or more logic gates, integrated circuits, processors, controllers, memory, electronic components, or information processing hardware, which are not limited thereto. Alternatively, the blocks may be implemented in software by a method of implementing application software, operating system software, firmware, or information processing software, which are not limited thereto. A single block may be implemented by being separated into multiple blocks that perform the same function, or conversely, a single block may be implemented to perform the functions of multiple blocks simultaneously. The blocks may also be implemented by being physically separated or combined according to any criteria. The blocks may be implemented to operate in an environment where their physical locations are not specified and they are spaced apart from each other by a communication network, the Internet, a cloud service, or a communication method not limited thereto. Since all of the above-mentioned methods of implementation fall within the scope of various embodiments that a person skilled in the art familiar with the field of information and communication technology can adopt to realize the same technical concept, any detailed methods of implementation should be interpreted as being included within the scope of the technical concept of the invention in this application.

[0058] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the attached drawings. In describing the present invention, to facilitate overall understanding, the same reference numerals are used for identical components in the drawings, and redundant descriptions of identical components are omitted. Furthermore, it is assumed that multiple embodiments are not mutually exclusive and that some embodiments may be combined with one or more other embodiments to form new embodiments.

[0059]

[0060] digital video codecs

[0061] FIG. 1 is a conceptual diagram of a video communication system according to an embodiment of the present invention. The video communication system (100) may be configured to include at least two terminals (110, 120) connected to each other through a network (105).

[0062] In one embodiment of the present invention, FIG. 1 may represent a block diagram for configuring a unidirectional video communication network. A first terminal (110) among the terminals may encode the video data in order to transmit (111) the video data through a network (105). A second terminal (120) among the terminals may be configured to receive (121) the encoded video data through a network, decode it, and display it.

[0063] In another embodiment of the present invention, FIG. 1 may represent a block diagram for configuring a bidirectional video communication network. For the bidirectional video communication, each terminal (110, 120) may be configured to encode video data acquired by itself for video transmission (112, 122) to each other terminal passing through the network. Each terminal may also receive (113, 123) video data transmitted through the network by another terminal, decode it, and be configured to display the decoded video data.

[0064] Each of the terminals (110, 120) shown in FIG. 1 may be exemplified as a device such as a server computer, a personal computer, a portable computer, and a smartphone, depending on the embodiment, but is not limited thereto. The present invention is applicable to any environment for establishing a unidirectional or bidirectional video communication network, and the network (105) should be understood to be formed by any means for carrying encoded video data between the terminals (110, 120).

[0065] In one embodiment of the present invention, the network (105) may refer to a wired or wireless communication network. In this case, depending on the embodiment, the network may be configured to communicate information using any communication standard, and the communication standard may include packet-based communication. The term "packet communication" may be understood to mean, for example, packets known as TCP or UDP.

[0066] However, in another embodiment of the present invention, the network (105) may be understood to include a process of information transmission using a recording medium. In this case, the configuration of the network is not limited to communication media only, but should be understood to include a process of temporarily storing and physically transporting information on a hard disk, solid-state disk (SSD), flash memory, CD (compact disc), DVD (digital versatile disc), Blu-ray disc, and other mechanical, electronic, or optical recording media.

[0067] Any other means of information communication or transport applied may be considered to fall within the scope of embodiments of the present invention as long as it is structured to support decoding by transmitting video data in an encoded state. Accordingly, in addition to the examples listed above, all means of information communication or transport known in the prior art or newly provided may fall within the scope of application of the present invention.

[0068] FIG. 2 is a conceptual diagram of the arrangement of encoders and decoders in a real-time video streaming environment according to an embodiment of the present invention. The streaming system (200) exemplified by FIG. 2 can be seen as applicable to a video data communication network including, for example, digital broadcasting, video telephone, and video conferencing. However, it should be seen that a technical structure identical or similar to the streaming system can be equally applied even when information transport via a recording medium is involved, as described above.

[0069] According to one embodiment of the present invention, the streaming system may include a video source (210) that generates a video stream. The video source may include a digital video acquisition means (212) that acquires uncompressed raw video, which may be composed of, for example, a digital camera or other equipment. The raw video stream (215) may have a massive capacity and may therefore be compressed by a video encoder (217) coupled to or connected to the video source.

[0070] The encoder (217) may be composed of means including hardware, software, or a combination of both, configured to implement an image encoding method and / or a method of implementing the same according to an embodiment of the present invention.

[0071] By passing through the encoder (217), an encoded bitstream (219) with a reduced capacity compared to the original video stream can be output. The bitstream (219) can be provided in real-time via communication by a relay device, for example, which may be referred to as a streaming server (220), and / or can be stored in a recording medium (225) of the streaming server (220) for subsequent use.

[0072] The streaming system (200) may include at least one streaming client (230, 240) that connects to the streaming server (220) to receive the encoded bit sequence (229) in real time or to acquire it subsequently. The streaming client may include a video decoder (232) that acquires the encoded bit sequence (229) (which may also be considered as a copy of the bit sequence (219) received by the streaming server), decodes the bit sequence (229), and outputs the resulting video data as video data in a form that can be displayed on a display (235) or other visual, auditory, or other sensory display means.

[0073] As mentioned above, the functions for encoding and decoding video data are collectively referred to as the coder-and-decoder system, or video codec.

[0074] FIG. 3 is a conceptual diagram of a functional unit of a video decoder according to an embodiment of the present invention. As shown in FIG. 3, a receiver (310) can receive at least one encoded video data to be decoded by a decoder (305). In an embodiment of the present invention, the encoded video data may be independent for each reception, and the decoding procedure of each independent video data may be independent from the decoding procedure of other video data. The encoded video data may be received by the receiver (310) through a hardware or software connection (315) to a device that stores the data, and as described above, the device that stores the data may be a type of streaming server located on the opposite side of a communication network or may refer to a physical recording medium, but is not limited thereto.

[0075] The receiving unit (310) can receive the encoded video data along with other accompanying data, such as encoded audio data or other auxiliary data, and each of the data can be separated from the video data and provided to an appropriate processing unit (312) other than the video decoder.

[0076] When the video data is received through a communication network, a buffer memory (320) may be combined between the receiver (310) and the decoder (305) to minimize delays and interruptions caused by the network environment. The buffer memory (320) may refer to a computer-readable recording medium that temporarily stores the received video data and reliably supplies it to a parser (330) corresponding to the input terminal of the decoder (305). However, the buffer memory may be unnecessary in environments where the bandwidth of the communication network is sufficient, where video data is being read from a recording medium located at a local location that is not physically separated, or where the possibility of communication delay is not predicted.

[0077] The video decoder (305) may include the parser (330) as its input to interpret the encoded video data. The parser separates (parses) a number of pieces of information stored in the form of bit sequences in the encoded video data according to a predetermined rule, and, if necessary, performs the function of entropy decoding (335) the entropy-coded video data, thereby reconstructing the symbols (338) which are segments of video encoding information. The symbols (338) may include all information for controlling the operation of the decoder (305), and / or may further include information for controlling a device that can operate in conjunction with the decoder (305), such as a display device. Control information for controlling the above-mentioned display device may include information in a format referred to as supplementary enhancement information (SEI) or video usability information (VUI).

[0078] As described above, the parser (330) may be configured to perform entropy decoding (335) of the encoded video data. The method of entropy encoding of the encoded video data may vary depending on the standard of the encoding, and decoding may be performed accordingly. Representative examples of the entropy encoding standard may include variable length coding, Huffman coding, and arithmetic coding, and each of the encoding methods may be context-adaptive or context-sensitive depending on the standard, or may be based on principles widely known to a person skilled in the art.

[0079] The parser (330) may be configured to extract at least one picture from the encoded video data. The definition of the picture may vary depending on the specification of the encoding, and depending on the specification, one or more of the examples listed below may simultaneously overlap. The picture may be grouped, defined, and / or divided into encoding / decoding units such as, for example, group of pictures (GOPs), pictures / frames, tiles, slices, macroblocks, blocks, subblocks, transform units (TUs), and prediction units (PUs).

[0080] The above parser (330) may be configured to extract encoding information, such as transform coefficients, quantization parameters (QPs), and / or motion vectors, from the encoded video data. The above parser (330) may be configured to perform entropy decoding (335) and parsing operations on the video data received from the buffer memory, and to selectively decode symbols (338) representing the encoding information. Additionally, the above parser (330) may be configured to selectively supply specific symbols (338) to specific decoding function units within the decoder (305), such as an inverse quantization and inverse transform unit (340), an intra prediction unit (350), an inter prediction unit (355), or a loop filter unit (360). The control of such information supply may be determined by the information permutation contained in the encoded video and may vary depending on the encoding standard; as such, it is not limited to the scope of the embodiments of the present invention and is not described in detail in this conceptual diagram.

[0081] The above decoder (305) may be composed of a plurality of conceptual functional units that receive and process the encoding information from the parser (330). It is obvious that these conceptual functional units may be combined with one another or further subdivided according to implementation needs. For example, they may be further separated for ease of implementation, or integrated into one for operational efficiency. In any case, each functional unit may be configured to perform close interaction with one another. However, despite the possibility of such integration or separation, the decoding procedure of video data applied as an embodiment of the present invention will be described as a combination of conceptual functional units as described below.

[0082] The above decoder may include an inverse quantization and inverse transform unit (340). The inverse quantization and inverse transform unit (340) may be configured to receive encoding information from the parser (330), including a method to be used for numerical transformation, a block size, quantization coefficients for recovering quantized information, and separation information of a quantization matrix representing the quantization coefficients in a simplified manner, and may be configured to output block values ​​(341) that can be input to an aggregator (370) as a result of processing the encoding information.

[0083] In one embodiment of the present invention, the output values ​​of the inverse quantization and inverse transformation unit (340) may include intra-predicted encoded block values. The intra-predicted block value may mean a value that can be decoded using prediction information within the picture currently being decoded, such as the current frame, without using prediction information from a previously decoded picture, such as a previous frame.

[0084] The prediction information within the current picture may be provided by the intra prediction unit (350). According to an embodiment of the present invention, the intra prediction unit (350) generates a block value of the same form as the block being decoded as prediction information using picture information of a spatially adjacent region derived from a picture that is currently being decoded and has been partially decoded. The picture information may be provided (381) from a buffer for the current picture, so-called line buffer (380). According to an embodiment, the merging unit (370) may be configured to merge the prediction information (351) generated by the intra prediction unit (350) with the block values ​​(341) provided by the inverse quantization and inverse transformation unit (340).

[0085] In another embodiment, the output values ​​of the inverse quantization and inverse transformation unit (340) may be inter-predicted encoded block values, and in some cases, may include block values ​​for which motion compensation has been performed. In this case, the inter-predicted unit (355) may extract and use sample information (386) used for motion-based prediction from a reference picture buffer (385). The information (356) derived by performing motion compensation on the sample information based on the symbols (338) included in the block values ​​as output values ​​may be configured to be merged by the merging unit (370) with the block values ​​(341) provided by the inverse quantization and inverse transformation unit (340). In this case, the block values ​​(341) may be referred to as so-called different or residual values.

[0086] The position information in memory used by the inter prediction unit (355) to extract the sample information from the reference picture can be determined by a motion vector provided to the inter prediction unit (355), which is composed of, for example, a combination of X, Y, and other symbols (338) for representing specific points of the reference picture. The inter prediction unit (355) may also include a function to interpolate and use the sample values ​​when a so-called 'subsampling' possible motion vector is provided, and may further include a function to predict and reinforce the value of the motion vector.

[0087] The output values ​​(371) of the above merging unit (370) are provided to the loop filter unit (360) and can be processed by various loop filtering methods. The loop filter unit (360) may be configured to receive not only the block unit output (371) of the above merging unit (370) but also the symbol (338) provided from the parser (330) to control its operation. The output of the above loop filter unit (360) may be output to an external display means, such as the display device, through the output connection (390), but may be stored (361) in the line buffer (380) for use in prediction to interpret subsequent intra or inter-encoded block values, and may also be stored in the reference picture buffer (385) via this.

[0088] Specific pictures, such as frames, can be utilized as reference pictures for performing predictive decoding during a subsequent decoding process once their decoding is complete. A picture (or frame) can be accumulated step by step in a line buffer (380) and decoding can proceed, and when a frame is decoded, the contents of the line buffer (380) are transferred (383) to the reference picture buffer (385), and a new line buffer (380) can be allocated for decoding a new frame.

[0089] The above video decoder (305) may be configured to perform a decoding operation according to a predetermined video compression technology that may be documented by various international standard specifications or commercial specifications. The specifications may include, for example, H.264, H.265, H.266, etc., which are international standard recommendations defined by the International Telecommunication Union Standardization Committee (ITU-T). A person skilled in the art will understand that each of these recommendations is equivalent to an international standard jointly defined by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC). The encoded video data may comply with a specific bitstream syntax defined by the specifications, as defined by the profile and level specified in the video compression specification document and standard document, and specifically within such document, and as required. In addition, the complexity of the encoded video data may be limited to a certain level for compliance with the profile and level. For example, any profile or level may be configured to limit the maximum picture size, the maximum decoding rate, and the maximum reference picture size. These limitations may also, in some embodiments, be further limited through metadata signals regarding the management of the HRD buffer included in the hypothetical reference decoder (HRD) and the encoded video data.

[0090] According to one embodiment of the present invention, the receiver (310) may receive additional redundant data along with the encoded video. The additional data may be considered as part of the encoded video data. The additional data may include information that can be used by the decoder (305) to properly decode the data or to more accurately reconstruct an image that is close to the image before encoding. The additional data may be provided in the form, for example, time, space, or layers for signal-to-noise ratio (SNR) enhancement, redundant slices, redundant pictures, and forward error correction codes.

[0091] FIG. 4 is a conceptual diagram of a functional unit of a video encoder according to an embodiment of the present invention. The encoder (405) may be configured to receive original video information (402) from a video source (401) and perform encoding.

[0092] The original video information (402) may have any suitable bit depth, for example, 8 bits, 10 bits, 12 bits, etc. Additionally, the original video information (402) may have any suitable color space, for example, R / G / B, Y / U / V, Y / Cb / Cr, etc. Additionally, the original video information (402) may have any suitable sampling structure corresponding to the color space, for example, in the form of Y / Cb / Cr 4:2:0, Y / Cb / Cr 4:4:4. The original video information (402) having such a predetermined format may be provided to the encoder in the form of a digital video stream.

[0093] In a unidirectional video communication network, the original video information (402) may be obtained from a recording medium that stores a pre-prepared original video. In a bidirectional video communication network, the original video information (402) may be obtained from a video acquisition device, such as a camera, that generates at least one video transmission stream included in the bidirectional video communication.

[0094] Video data containing the original video information (402) may be composed of a plurality of pictures configured to simulate motion by playing them in chronological order. The pictures may also be expressed as concepts such as frames in addition to pictures. The pictures may include one or more samples depending on the type of sampling structure, color space, etc. used. A person skilled in the art will understand that the terms "sample" and "pixel" in digital images are closely related. The operation of the encoder will be explained below with a focus on such samples.

[0095] According to one embodiment of the present invention, an encoder (405) may be configured to encode and compress pictures (and / or information in which they are grouped or divided) constituting the original video information (402) into the form of encoded video information in real time (or according to other temporal requirements as required by the method of implementation).

[0096] In the above encoder (405), the control unit (450) may be a functional unit configured to control an appropriate encoding speed. The control unit (450) may be configured to control other functional units and to be functionally coupled to the functional units described below. The parameters set by the control unit (450) may include parameters related to bitrate control, such as picture skip, quantizer, and variable values ​​for applying image quality optimization techniques, and may also include values ​​such as picture size, the structure of a group of pictures (GOP), and the maximum search range of motion vectors. A person skilled in the art will be able to understand the various other functions that the control unit (450) may have, and such other functions may be added or removed according to the design of the video encoder optimized for the individual system design.

[0097] According to an embodiment of the present invention, the encoder (405) may be configured to operate in a structure such as a "coding loop" that is well known to a person skilled in the art. To explain it in a simplified manner for example, the coding loop may consist of an internal encoder (so-called "source coder") (410) responsible for receiving a picture to be encoded and generating symbols based on at least one reference picture that has been previously encoded, and a local decoder (420) configured to be connected to the internal encoder. The local decoder (420) may be configured to perform an operation to reproduce sample data to be generated by a decoder (490) located at an actual remote location that receives video information encoded from the encoder (405) by receiving the output of the internal encoder (410).

[0098] Video data composed of sample data reconstructed by the internal decoder (420) can be configured to be input into the reference picture buffer of the encoder (405). As described above, since the internal decoder (420) is implemented to reproduce the result output by the encoder (405) to be decoded at a remote decoder, the video data recorded in the reference picture buffer can also be identical in bit unit to the information of the reference picture buffer held by the remote decoder. That is, the prediction function unit that may be included in the encoder (405) can read values ​​identical to the sample values ​​of the previous frame that the decoder will refer to during the decoding process from the reference picture buffer of the encoder (405).

[0099] As described above, the principle of achieving a match between the encoder (405) and the decoder (490) in reference picture buffers by means of an internal decoder (420) on the side of the encoder (405) is widely known to a person skilled in the art, and the method of responding to an environment where such an environment is not guaranteed (e.g., loss of information due to communication failure) may also follow what is known to a person skilled in the art.

[0100] An example of the operation method of the internal decoder (420) described above has been explained in detail with reference to FIG. 3. The decoder of FIG. 3 can be considered as the decoder (490) of the "remote location" described above. The internal decoder (420) may be implemented excluding lossless encoding and decoding sections such as the parser (330) or entropy decoding (335), because the internal encoder (405) is implemented simply to reproduce the operation of the decoder located at the remote location, so it is acceptable to decode the symbols immediately without requiring a process of compressing and restoring the symbols. Therefore, it is acceptable for the functional parts preceding the parser and entropy decoder shown in FIG. 3 not to be provided or at least to be implemented only partially.

[0101] As described above, according to a preferred embodiment of the present invention, any decoder function present in the decoder (excluding the parser and entropy decoder) can naturally exist as substantially the same function in the corresponding encoder (405).

[0102] The operation of the encoding function unit that may be included in the encoder (405) can be considered as the inverse operation of the decoder function unit. Therefore, the embodiment can generally be explained by performing the operation of the decoder function unit in reverse. For example, a quantization and transform function unit corresponding to the inverse quantization and inverse transform unit may be provided, and an inter-prediction encoding unit corresponding to the inter-prediction unit may be provided. In addition to this, some additional explanations will be added.

[0103] The internal encoder (410) may be configured to perform encoding for input picture information, e.g., an input frame, by a predictive encoding method executed by a predictive encoding unit (440) that operates by referencing at least one picture (or frame) encoded in a temporally earlier order from video data designated as at least one reference picture information, e.g., a reference frame, from a reference picture buffer (430). In this case, the encoder (405) may be configured to encode a differential between blocks of samples constituting the input picture and blocks of samples constituting the reference picture.

[0104] The internal decoder (420) can decode video data that can be designated as the reference picture from the symbols generated by the internal encoder (410). As described above, since this video data is identical to the decoding operation performed by the remote decoder, the video data used as the reference picture may be provided to the encoder (405) in a form that has undergone lossy compression and has some damage, and this operation may be intended to match the operation with the decoder.

[0105] The prediction encoding unit (440) may be configured to perform a prediction search operation within the encoder (405). The prediction search operation may refer to an operation corresponding to the inter-prediction or intra-prediction described in the description of the decoder. For picture information that is input and scheduled to be newly encoded, the prediction unit may access the reference picture buffer (430) to retrieve information in order to obtain information such as a motion vector, a block shape, metadata that may include the same, and a sample block to be actually referenced, which are information indicating a point of a reference picture that can function as prediction reference information suitable for the new picture information. The prediction encoding unit (440) may operate according to the so-called "sample block by pixel block" standard to obtain appropriate prediction reference information. According to one embodiment of the present invention, at least one prediction reference information pointing to at least one reference picture information stored in the reference picture buffer (430) may be designated for the input picture, as determined based on the search results obtained by the prediction encoding unit (440).

[0106] In one embodiment of the present invention, the control unit (450) may be configured to manage the overall encoding operation of the internal encoder (410), including the setting of parameters used to encode video data.

[0107] The outputs of all the aforementioned functional units may be subject to entropy coding (460) for final output. The entropy coding (460) may include various entropy coding techniques such as variable length coding, Huffman coding, and arithmetic coding as described above for the symbols generated by the various functional units, and each of the coding methods may be context-adaptive or context-sensitive methods depending on the standard, or may be based on principles widely known to a person skilled in the art. Such entropy coding (460) can typically achieve lossless compression and may be configured to convert at least one symbol generated by the functional units into encoded video data.

[0108] The control unit (450) may apply a type of encoding in which a specific picture is encoded during the encoding interval to each picture (or frame) in controlling the operation of the encoder (405). Depending on the type, the way the picture is encoded may be affected. Depending on the embodiment, the type may include being classified into the following "frame types."

[0109] FIG. 5 is a conceptual diagram of a frame type according to an embodiment of the present invention. The following description will be explained together with reference to FIG. 5.

[0110] An intra ("I") picture (510) may refer to a picture that can be encoded and decoded using only its own information without referencing other picture information within the video data through predictive encoding. The "I" picture may be designated by names such as key frame, independent / instantaneous decoder refresh (IDR) frame, and clean random-access (CRA) frame according to the video encoding standard. The "I" picture designated by such various names may have various modifications and application methods as permitted by each standard and may differ partially from one another. In addition to those listed above, various application methods for implementing the "I" picture may be based on various methods that are already known to a person skilled in the art or may be newly provided.

[0111] A prediction ("P") picture (520) may mean a picture that can be encoded and decoded via intra- or inter-prediction based on at least one prediction information and / or motion vector pointing to at least one reference picture to predict sample values ​​of a block constituting the picture. The "P" picture may be configured to reference only one reference frame according to the video encoding standard, or configured to reference one or more reference frames. In the case of referencing one or more reference frames, sample information and / or associated metadata derived from multiple reference pictures may be used for the reconstruction of a single block. However, in common cases, a picture designated as a "P" picture may be understood as a picture that performs a reference limited to a temporally preceding picture.

[0112] A bidirectional prediction ("B") picture (530) may mean a picture that can be encoded and decoded via intra or inter prediction based on at least one prediction information and / or motion vector pointing to at least two reference pictures to predict sample values ​​of blocks constituting the picture. In common cases, the picture designated as the "B" picture is distinguished from the picture designated as the "P" picture and may be understood as a picture performing the reference, not limited to a picture that precedes in time.

[0113] Video data is spatially divided into multiple sample blocks during the encoding and decoding process, and encoding can proceed in block units. The block units include, for example, sizes such as 4×4, 8×8, 4×8, or 16×16 in units of horizontal / vertical pixels, as is widely known, but are not limited thereto. The blocks may be encoded by a predictive encoding method by referencing any other (already encoded) blocks, as allowed and / or restricted by the type specified for each picture in which the blocks are included. For example, blocks of the "I" picture (510) may not use a predictive encoding method, or may be encoded by referencing blocks that have already been encoded within the same part of the picture. That is, only the so-called intra-prediction method may be used. In contrast, for the "P" picture (520), at least one reference picture encoded in a previous time unit may be referenced, and thus inter-prediction may also be used in encoding along with intra-prediction. In the case of the "B" picture (530), reference can be performed even among reference pictures that are encoded earlier in the encoding order but follow later in terms of time unit. However, it is widely known that there may be blocks encoded without relying on predictive encoding within the "P" picture or the "B" picture.

[0114] The above video encoder (405) may be configured to perform encoding operations according to a predetermined video compression technology that may be documented by various international standard specifications or commercial specifications. Examples of the specifications may include all those described in the decoder.

[0115] According to one embodiment of the present invention, a transmitting unit (470) may buffer the encoded video data generated by the entropy encoding in order to provide / transmit the video data (ultimately to a decoder (490) at a remote location) through a hardware or software connection (495) to a device storing the encoded video data. According to an embodiment, when providing / transmitting the encoded video data from the video encoder (405), the transmitting unit (470) may receive and merge other data accompanying the encoded video data, such as encoded audio data or other auxiliary data, from a separate source (480).

[0116] According to one embodiment of the present invention, the transmitting unit (470) may be configured to transmit additional data along with the encoded video. The additional data may be considered as part of the encoded video data. The additional data may include information that can be used by a decoder to properly decode the data, or to more accurately reconstruct an image that is close to the image before encoding. Examples of the additional data may include all the examples previously shown in relation to the receiving unit (310) of the decoder.

[0117] The present invention can be implemented by digital video compression standards that are understood by and widely used by people skilled in the art, as described above. The digital video compression standards may include at least one of the compression standards known by standard names such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.

[0118] FIG. 6 is a conceptual diagram showing the structure of a video encoder according to another embodiment of the present invention. What is shown in FIG. 6 may be the approximate structure of a video encoder widely known by standard codes such as ITU-T H.266 and ISO / IEC 23090-3, and by the designation MPEG-I Part 3 or the common name versatile video coding (VVC).

[0119] According to FIG. 6, a video encoder (605) may be configured to receive uncompressed and unencoded original video data (601) as input and output an encoded bit sequence (602). When the video data (601) is intra-encoded, it may be supplied directly to a luminance signal mapping unit (610a), or supplied to a luminance signal mapping unit (610b) via an inter-prediction unit (620) that includes motion vector extraction. When the intra-encoded, the mapped luminance signal may be supplied to an output merger (606) by selecting (608) at least one of the intra-prediction encoded signal via the intra-prediction unit (625) or the inter-prediction encoded signal output from the luminance signal mapping unit (610b) via the inter-prediction unit (620). The result of the output merger may be applied to a chroma scaling unit (615). (The operation of the above-mentioned illuminance signal mapping unit (610) and the operation of the above-mentioned color difference signal reduction (615) are collectively referred to as the luma mapping / chroma scaling (LMCS) process.) The reduced color difference signal can be provided to a transform unit (630), and the transform unit (630) can perform an adaptive color transform, particularly on the color difference signal. The coefficients derived as a result of the transformation are applied to a quantization unit (640) and quantized. This results in lossy compression, and the result of the lossy compression can be output as a bit sequence (602) through a multi-hypothesis context-adaptive arithmetic coding unit (650), which is a lossless compression method.

[0120] Meanwhile, the result of the lossy compression above can enter the decoding procedure by undergoing inverse quantization (645), inverse transform (635), and luminance signal expansion (617) processes to generate a coding loop. The result of the luminance signal expansion above can be supplied to an internal merger (607) along with the result of selecting (608) at least one of the previously generated intra-predicted coded signal or inter-predicted coded signal. The result of the internal merger above can undergo inverse luminance signal mapping (617) and then undergo processing such as a deblocking filter (660), a sample adaptive offset (SAO) (670), and an adaptive loop filter (ALF) to reproduce the image quality improvement process in the decoder. The result of reproducing the operation in the decoder as described above is applied to the reference picture buffer (690) and can be recycled for prediction encoding by the inter prediction unit (620).

[0121] The present invention may also be utilized by or combined with an enhanced compression model (ECM), which is an implementation of a next-generation video codec being developed by the Joint Video Experts Team (JVET), an international standardization expert organization. The enhanced compression model may include an enhanced intra-predictive coding method, an enhanced inter-predictive coding method, an enhanced transform and transform coefficient coding method, an enhanced adaptive loop filtering method, a bilateral filtering method, a new sample adaptive offset (SAO) method for image quality improvement, an extended entropy coding method, and an enhanced gradual decoding refresh (GDR) technique.

[0122]

[0123] Compression technology for screen content and game content videos

[0124] The image encoding / decoding technology to which the present invention can be applied may be configured to process various types of images as input. For example, the encoder and / or decoder according to the present invention may be configured to incorporate an algorithm or method capable of efficiently processing computer graphics-processed images.

[0125] The above computer graphics-processed image may typically include screen contents and gaming contents. The above screen contents refer to an image in which graphically processed content is included in a part or the entire area of ​​the image, and the above gaming contents may refer to an image in which the entire screen is generated using computer graphics, such as a game screen.

[0126] FIG. 7 is an exemplary diagram illustrating the configuration of screen content and gaming content according to an embodiment of the present invention. Generally, screen content is a form in which specific information is displayed in a part or the entire area within a video. For example, as shown in FIG. 7 (a), it may be a video displaying presentation materials across the entire screen, or a video displaying subtitles or product information across a part of the screen. In particular, representative examples of screen content include subtitles at the bottom of a news broadcast, a product information display area in a home shopping broadcast, presentation materials, and web browser screen captures. Additionally, gaming content, as shown in FIG. 7 (b), is a video in which the entire screen is generated using computer graphics. Gaming content is composed of 3D or 2D graphics rendered by a game engine, and features characters, backgrounds, interface elements, etc., all generated digitally.

[0127] However, since the field where full-screen computer graphics rendering is frequently performed is the gaming field, the category of video having such characteristics is collectively referred to as "game content." It should be understood that computer-generated video that does not necessarily have computer games as its main content or purpose is also included in the aforementioned game content in a comprehensive sense. Therefore, the game content encoding / decoding technology in the present invention is defined as a function that provides an efficient encoding / decoding process for video having certain characteristics, and it can be made clear that it is not limited to a relationship with "games."

[0128] Real-world footage, captured by a camera, contains a significant amount of noise. This noise can occur during the process of converting light brightness signals (electrical charge) into (analog) electrical signals via CCD or CMOS image sensors, and then converting them back into digital signals. This is often caused by factors such as charge loss or increased dark current resulting from sensor overheating. Furthermore, even subtle changes in light intensity or slight variations in airflow can lead to significant differences in pixel values ​​for the same color within an image. In contrast, images generated through graphic processing have colors artificially created, making it highly likely that pixel values ​​for the same color within the image will be identical.

[0129] FIG. 8 is a conceptual diagram illustrating the difference in pixel distribution between real-world images and screen content in the present invention. As shown in FIG. 8(a), the pixel value distribution of a movie image, as an example of a real-world image, takes the form of a continuous graph similar to an analog signal; whereas, as shown in FIG. 8(b), the pixel value distribution of a graphically processed game image is concentrated on a few sample pixel values ​​similar to a digital signal. This indicates that while the pixel values ​​of real-world images, which naturally have added noise, possess diverse values, arbitrarily produced images show a simple shape in their distribution as they are biased toward specific pixel values.

[0130] International standard video compression technologies such as H.265 / HEVC and H.266 / VVC can be used for video compression methods for screen content and game content. The aforementioned video compression standards were primarily developed to be suitable for real-world video captured by cameras. However, recently, the development of compression technologies that consider the characteristics of screen content and game content is actively underway, and the present invention can be considered to be part of this effort.

[0131] The conventional international standard for video compression, VVC (versatile video coding), provides a total of 67 prediction modes as intra-frame prediction methods, including planar mode, DC mode, and 65 directional intra prediction modes. Fig. 9 is a conceptual diagram illustrating the directional prediction modes of VVC. Referring to Fig. 9, modes 0 and 1 are non-angular modes without directionality, such as Planar mode and DC mode, while modes 2 through 66 represent 65 directional prediction modes based on angular angles.

[0132] In various embodiments of the present invention, the directional intra-prediction mode may have various types of angle prediction modes. For example, the directional intra-prediction mode may include 33 directional modes, 65 directional modes, 129 directional modes, and / or 257 directional modes. The number of directional modes is not limited to the above examples and may be specified as at least one arbitrary number according to the embodiment. According to the embodiment, the number of directional modes may be configured by combining two or more numbers, for example, 65 basic modes and 129 extended modes may be configured to be used complementarily.

[0133] In various embodiments of the present invention, the directional intra-prediction mode may be configured to be applied selectively according to the block size. For example, the same number of directional modes may be applied for all block sizes. As another example, a different number of directional modes may be applied depending on the block size, such as applying a larger number of directional modes to large blocks and a smaller number of directional modes to small blocks. As yet another example, a specific directional mode may be applied only to a specific block size, such as applying only 33 directional modes to a 4×4 block and applying 65 or more directional modes to a block of 8×8 or larger. As yet another example, the application of directional modes may be dynamically determined according to the size and shape of the block, such as applying horizontal or vertical modes more finely depending on the width-to-height ratio of the block.

[0134] In various embodiments of the present invention, the directional intra-prediction mode may be signaled by a direct index value for each directional mode, or by a difference value from a reference mode value, such as the Most Probable Mode (MPM), derived during the encoding / decoding process. The reference mode may mean that it is determined, derived, calculated, or predicted by a preceding encoding / decoding process.

[0135] To generate a prediction block for a prediction mode within the screen, surrounding already restored samples are used, and these samples can be referred to as reference samples. FIG. 10 is a conceptual diagram of the application of reference samples to a block to be encoded in the present invention. Referring to FIG. 10, reference samples for a block to be encoded of size N×N are shown, and (2N×2) + 1 reference samples may be required on the left (L, left), top (T, top), and top-left (TL, top-left).

[0136] In one embodiment of the present invention, the process of generating reference samples can be broadly divided into three steps. First, a step of filling possible reference samples from surrounding information; second, a step of generating unfilled reference samples; and third, a step of filtering and smoothing the reference samples may be included.

[0137] FIG. 11 is a conceptual diagram of the process of generating reference samples according to an embodiment of the present invention. In FIG. 11 (a), the first step described above, which is filling possible reference samples from surrounding information, can be described as a process of filling a reference area (1120) when a spatially adjacent block around the current encoding target block (1110) has already been encoded / decoded and can be used as a reference sample. In FIG. 11 (b), the second step described above, which is generating unfilled reference samples, can be described as a step of generating a remaining reference sample value (1130) using the previously filled available reference sample value to fill the reference samples that were not filled in the preceding step.

[0138] The third step described above, which is to filter and smooth reference samples, can be configured to selectively determine whether to filter based on the intra mode of the target block, the size of the current block, and the partition structure; if filtering is applied, an interpolation filter can be applied to filter the reference samples. That is, smoothing operations can be performed on the reference samples filled or / or generated in the preceding steps.

[0139] Based on the completed reference samples as described above, one prediction mode among a plurality of prediction modes as shown in FIG. 9 can be selected. Additionally, a prediction / reference block for the selected prediction mode can be generated using the values ​​of the reference samples. The generated prediction block value can redefine or update the reference pixel value according to the selected prediction mode and the position of the reference pixel value within the prediction block. This can be referred to as a position-dependent prediction combination (PDPC). A final reference block can be generated through the aforementioned multiple processes.

[0140] In one embodiment of the present invention, considering the characteristics of screen content, the filter may be selectively disabled for screen content based on the fact that the aforementioned filtering process may actually degrade image quality. For example, a method may be applied to automatically identify screen content using a hash block-based screen content identification algorithm and to disable filtering for the corresponding area. The hash block may refer to a structural means that enables efficient comparison at the hash level rather than the pixel level by deriving a single hash value from a block unit of size 4×4 or 8×8. Furthermore, the structure of the hash block described above may be understood as sharing, reusing, or improving upon the same or similar technology applied to other encoding / decoding technologies utilizing similar structures, such as Intra Block Copy (IBC) and Intra Template Matching Prediction (IntraTMP). In this case, since screen content has uniform pixel values ​​and a high hash hit percentage, it can be identified as screen content if it exceeds a threshold.

[0141]

[0142] Composition of the present invention

[0143] The present invention presents an image encoding / decoding method utilizing features of screen content and game content images. In particular, it presents an in-frame prediction method utilizing features of the pixel value distribution of screen content and game content images. Furthermore, it presents a signaling method for a partial or entire area within an arbitrarily created image among screen content and game content. For example, it proposes a method for indicating whether the entire area within the image is the created part, or whether only a partial area (e.g., the bottom of the image) is the created part.

[0144] By clearly signaling the production area within the video, encoding and decoding methods suitable for screen content and game content can be used in the said area (hereinafter, production area). As a result, encoding / decoding complexity can also be reduced by improving compression performance and eliminating unnecessary coding tools.

[0145] Existing international standard video compression technologies were primarily developed to be suitable for live-action footage captured by cameras. Consequently, some coding tools may not be suitable for screen or game content. However, until now, there has been no explicit method to indicate whether the video to be encoded is screen or game content; consequently, encoding and decoding processes were performed using the same coding tools applied to live-action footage, even when the video was screen or game content.

[0146] The present invention provides a new method for clearly indicating a production area (arbitrary content or arbitrary region) in screen content and / or gaming content through various embodiments.

[0147]

[0148] First embodiment

[0149] According to one embodiment of the present invention, an arbitrary_content_flag indicating that the target image is screen / game content can be used to indicate that the entire area of ​​the image is a production area. If the value of the flag is "on", it can be configured to mean that the entire area of ​​the target image is a production area. Here, the arbitrary_content_flag is a high-level syntax (HLS) and can be placed in one or more locations among SPS (sequence parameter set), PPS (picture parameter set), PH (picture header), and SH (slice header). For example, if the production area information is the entire area within every frame in the entire sequence, the information can be placed in SPS, but if the production area information is the entire area within a specific frame in the entire sequence (where the specific frame has no production area), it can be placed in one or more locations among PPS, PH, and SH.

[0150] According to one embodiment of the present invention, a hash block-based automatic determination mechanism may be further utilized. An encoder can automatically determine whether content is screen content by calculating the hit rate of hash blocks at the sequence level. Specifically, a block unit of size 4×4 or 8×8 can be converted into a single hash value, and the frequency of occurrence of the same hash value within the entire frame can be calculated. If this hit rate exceeds a predefined threshold, the corresponding sequence or picture can be classified as screen content and the arbitrary_content_flag can be automatically set. This automatic determination may be applied optionally depending on the implementation of the encoder and may be used in parallel with manual setting.

[0151] Table 1 below is an example of an SPS syntax element configured according to the present invention.

[0152] seq_parameter_set_rbsp() {arbitrary_content_flag}

[0153] According to one embodiment of the present invention, a method for disabling filters applied to screen content can be signaled more efficiently in conjunction with the arbitrary_content_flag. Referring to some of the latest video codec implementations, each is controlled independently using three separate flags (sps_ref_sample_filtering_disabled_flag, sps_pdpc_disabled_flag, sps_intra_fusion_disabled_flag) for reference sample filtering, PDPC, and intra-predictive fusion. However, since these filters tend to be disabled together in actual screen content processing, the present invention can reduce signaling overhead by integrating them. For example, when the arbitrary_content_flag is enabled, the three filters can be controlled collectively through a single integrated flag, or additional signaling can be performed only when individual control is optionally required. Various embodiments of the present invention may include additional flags for controlling the disabling of individual filters. For example, at least one individual flag among reference_sample_filter_disabled_flag, pdpc_disabled_flag, and intra_fusion_disabled_flag can be signaled. As another example, at least one of the aforementioned individual flags may be configured to be additionally signaled when arbitrary_content_flag is enabled. This can provide the flexibility to selectively disable only specific filters depending on the characteristics of the screen content. Additionally, signaling overhead may be reduced by managing all or some of the aforementioned flags as flags with integrated semantics.

[0154] In more diverse embodiments of the present invention, additional flags may be included for controlling the activation and / or deactivation of at least one filter for intra-frame prediction and / or for controlling the activation and / or deactivation of at least one filter for inter-frame prediction. For example, individual flags such as intra_filter_disabled_flag and inter_filter_disabled_flag may be configured to signal. For example, if the intra_filter_disabled_flag value is "on", one or more filters for intra-frame prediction, such as reference sample filtering, PDPC, and intra-prediction fusion, may be disabled, and for another example, if the inter_filter_disabled_flag value is "on", one or more filters for inter-frame prediction, such as overlapped block motion compensation (OBMC), geometric partitioning mode (GPM) blending, and bi-directional optical flow (BDOF), may be configured to be disabled.

[0155] FIG. 12 is a flowchart of a production area processing according to an embodiment of the present invention. According to an embodiment of the present invention, information of a reference sample is first filled from surrounding information (S1210), and then an unfilled reference sample is generated (S1220). At this time, the arbitrary_content_flag value is read (S1230), and if it is "on," the filtering / smoothing process of reference samples for all encoding target blocks (S1235) can be omitted. Depending on the embodiment, not only reference sample filtering but also prediction post-processing processes such as PDPC (Position Dependent Prediction Combination) and intra prediction fusion can be disabled.

[0156]

[0157] Second embodiment

[0158] According to one embodiment of the present invention, a flag (e.g., "arbitrary_content_flag") indicating that the target image is screen / game content can be used to indicate that a portion of the image is a production area. If the flag value is "on", a portion of the target image is a production area, and the system may be configured to additionally signal area information regarding the production area.

[0159] In the description of the present embodiment, information regarding a rectangular production area was displayed based on 2D CTU (coding tree unit) index information on the horizontal and vertical axes, but this is a representative example for explaining the present invention and is not limited thereto. The present invention may be configured to include any method capable of displaying a screen-processed area (subtitles, product information, etc.) within an image, i.e., a production area.

[0160] For example, in addition to the CTU index, a method using a Tile index may be additionally included. The Tile may refer to a rectangular area designed for parallel processing as a unit larger than the CTU. The Tile may be arranged to divide the picture into multiple rectangular areas, and each Tile may be encoded / decoded independently. The Tile-based signaling has less signaling overhead than the CTU-based method and is particularly efficient in applications such as cube maps of 360-degree video or multi-view video. For example, when using a Tile index, a tile_idx based on the raster scan order may be used, and the production area may be specified using boundary area specification parameters such as tile_idx_top_left and tile_idx_bottom_right.

[0161] As another example, the raster scan order of the top-left and bottom-right positions of a rectangular production area may be displayed as CTU index information based on the index. Additionally, since there may be one or more production areas within the image, an embodiment may be included in which the number of production areas is signaled and the corresponding number of production areas are displayed by any method as described above. For example, in an actual broadcasting environment, multiple production areas may exist simultaneously. For example, in a news broadcast, the headline at the top, subtitles at the bottom, and weather information on the side may each constitute separate screen content production areas. In such cases, a region-specific control flag, such as region_filter_control_flag[i], may be additionally signaled to control filter disabling independently for each area.

[0162] One or more of the above arbitrary_content_flag, area information for the production area, and number information for the production area may be placed as HLS at one or more of SPS, PPS, PH, and SH. For example, if the production area information is the same for all frames in the entire sequence, the information may be placed at SPS, but if the production area information is different for every frame in the entire sequence, it may be placed at one or more of PPS, PH, and SH.

[0163] Table 2 below is an example of an SPS syntax element configured according to the present invention.

[0164] seq_parameter_set_rbsp() {arbitrary_content_flagif(arbitrary_content_flag) {arbitReg_ctu_top_left_xarbitReg_ctu_top_left_yarbitReg_width_minus1arbitReg_height_minus1}}

[0165] Table 3 below is another example of an SPS syntax element configured according to the present invention.

[0166] seq_parameter_set_rbsp( ) {arbitrary_content_flagif( arbitrary_content_flag ) {num_ arbitReg_minus1for( i = 0; num_arbitReg_minus1 > 0 && i <= num_arbitReg_minus1; i++ ) {arbitReg_top_left_x[ i]arbitReg_top_left_y[ i]arbitReg_width_minus1[ i]arbitReg_height_minus1[ i]}}}

[0167] In a syntax structure such as Table 2 and / or 3, num_arbitReg_minus1 may mean displaying a value of the number of crafting areas minus 1. arbitReg_top_left_x may mean displaying the top-left position of the crafting area as a CTU index based on the horizontal axis. arbitReg_top_left_y may mean displaying the top-left position of the crafting area as a CTU index based on the vertical axis. arbitReg_width_minus1 may mean displaying the width of the crafting area as a value of CTU index minus 1. arbitReg_height_minus1 may mean displaying the height of the crafting area as a value of CTU index minus 1.

[0168] According to the embodiment, a flag indicating the region designation method (e.g., "arbitrary_region_type") may be added to the syntax of Table 2 and / or Table 3. Additionally, if the index is implemented based on tiles as described above, the syntax including arbitReg_top_left_x and / or arbitReg_top_left_y may be designated by the tile index rather than the CTU index. Also, in the case of tile-based, syntax elements of size designation such as arbitReg_width_minus1 and / or arbitReg_height_minus1 may be ignored as they correspond to information that can already be derived from the specified size of the tile. How the syntax is interpreted may be determined in advance by the specification or by a separate signal flag such as arbitrary_region_type.

[0169] According to one embodiment of the present invention, it is possible to check whether the location of a block to be encoded belongs to a production area; if it belongs to a production area, the filtering / smoothing process of reference samples is omitted, and if it does not belong to a production area, it is treated as a real-world image and the filtering / smoothing process of reference samples is performed. Additionally, according to one embodiment of the present invention, if the encoding block spans the boundary between a production area and a real-world area, the filter strength can be adjusted adaptively based on the proportion of the production area within the block and / or based on a predetermined threshold value. For example, a method can be used in which the filter is completely disabled if 50% or more of the block belongs to a production area, and a weak filter is applied if it is less than that.

[0170]

[0171] Third embodiment

[0172] According to one embodiment of the present invention, without explicitly signaling information about the production area based on HLS, the encoder determines during the encoding process whether the coding target block (coding unit, CU) belongs to screen / game content, and if it is determined to be a production area, it may signal a bit sequence by setting a flag (e.g., filtering_interpolation_disabled_flag) that disables filtering / smoothing of reference samples to "on". Here, the flag may be signaled restrictively when the HLS flag value indicating that it is screen / game content is "on", but it is not necessary to do so. According to an embodiment, if the value of the "filtering_interpolation_disabled_flag" flag is "on", the PDPC process may also be omitted.

[0173] The encoder according to the present embodiment may be configured to automatically determine screen content at the CU level using a hash block-based algorithm. Specifically, hash values ​​for 4×4 or 8×8 sub-blocks within the current CU may be calculated, and the frequency of occurrence of the same hash value may be measured. At this time, if this frequency exceeds a threshold (e.g., 70%), the corresponding CU may be classified as belonging to the production area of ​​screen content, and the value of filtering_interpolation_disabled_flag may be set. Such adaptive control at the CU level is particularly effective in the case of screen content where real-world areas and production areas are intermingled within the image.

[0174] According to one embodiment of the present invention, when a flag (e.g., filtering_interpolation_disabled_flag) for disabling filtering / smoothing of reference samples is "on" without explicitly signaling information about the production area based on HLS, the encoder may be configured to independently analyze and / or determine whether the block to be encoded (CU) belongs to screen / game content during the encoding process, and if it is determined to be the production area, to disable at least one filter among filters for intra-screen prediction (e.g., reference sample filtering, PDPC, intra-predictive fusion).

[0175] According to one embodiment of the present invention, additionally, spatial correlation can be utilized by referring to the screen content determination results of surrounding CUs. For example, if both the left and top CUs are determined to be screen content, the current CU is also likely to be screen content, so the threshold for classifying the production area can be lowered to improve determination accuracy.

[0176] Table 4 below is an example of a syntax element of CU according to an embodiment of the present invention.

[0177] coding_unit(x0, y0, cbWidth, cbHeight, cqtDepth, treeType, modeType) {if(arbitrary_content_flag)filtering_interpolation_disabled_flag}

[0178] According to one embodiment of the present invention, for granular control of individual filters, signals related to the deactivation of individual filters, such as reference_sample_filter_disabled_flag, pdpc_disabled_flag, and intra_fusion_disabled_flag, may appear sequentially or selectively depending on the case. In one embodiment, the above phrases may be configured so that at least one may appear only when filtering_interpolation_disabled_flag is signaled as true.

[0179] According to one embodiment of the present invention, the filtering_interpolation_disabled_flag may have an integrated semantic value instead of being considered as an individual flag. That is, when the flag is set to true, at least one smoothing-related processing unsuitable for screen content, such as reference sample filtering, PDPC, and intra-fusion, may be configured to be selectively or collectively disabled. This enables efficient processing even when real-world areas and production areas are mixed within the image through adaptive control at the CU level.

[0180] According to another embodiment of the present invention, flags may be included to control the deactivation of at least one filter for intra-frame prediction and the deactivation of at least one filter for inter-frame prediction, respectively. For example, individual flags such as intra_filter_disabled_flag and inter_filter_disabled_flag may be signaled. In this case, if the value of intra_filter_disabled_flag is "on", at least one filter among reference sample filtering, PDPC, and intra-prediction fusion, which are filters for intra-frame prediction, may be deactivated, and if the value of inter_filter_disabled_flag is "on", at least one filter among overlapped block motion compensation (OBMC), geometric partitioning mode (GPM) blending, and bi-directional optical flow (BDOF), which are filters for inter-frame prediction, may be configured to be deactivated.

[0181] According to the embodiment, additional flags for granular control of individual filters may be conditionally signaled along with the integration flag. For example, if you want to selectively utilize only some filters in specific screen content, you can switch to an individual control mode via selective_filter_control_flag.

[0182]

[0183] Extension of implementation method

[0184] Although the present invention exemplarily describes a method for omitting the filtering / smoothing process of reference samples among in-screen prediction methods for screen / game content through multiple embodiments, the configuration method of the present invention is not limited to the embodiments described above, and the present invention can be equally applied to various processes such as filtering, interpolation, and smoothing applied to pixel values, predicted values, restored values, etc.

[0185] For example, regarding intra-frame prediction methods, specific processes such as PDPC, gradient PDPC, intra-frame prediction fusion, extrapolation filter-based intra-frame prediction (EIP) mode, and subpixel unit interpolation for block vectors (BV) for intra-frame block copy (IBC) may be included in the scope of application of the present invention.

[0186] For example, in the inter-frame prediction method, specific processes such as overlapped block motion compensation (OBMC), geometric partitioning mode (GPM) blending, and bi-directional optical flow (BDOF) may be included in the scope of application of the present invention.

[0187] In addition, in the inter-frame prediction method of one embodiment, for granular control of individual filters, signals related to the disablement of individual filters, such as obmc_disabled_flag, gpm_blending_disabled_flag, and bdof_disabled_flag, may appear consecutively and optionally in bit sequences depending on the case.

[0188] The aforementioned in-frame or inter-frame prediction encoding methods may operate to include interpolation or smoothing during the processing. However, if one examines the pixel value distribution of screen / game content images (e.g., Fig. 8(b)), it can be understood that the filtering / smoothing process may be unsuitable for screen / game content images and may be omitted, as the pixel values ​​are sampled rather than linearly distributed.

[0189] In addition to disabling various filtering-related functions in predictive coding, optimization in the transform and quantization processes can also be applied based on the present invention. For example, since screen content mainly contains horizontal / vertical edges, Discrete Sine Transform (DST) may be more efficient than Discrete Cosine Transform (DCT). Furthermore, since the transform skip mode is frequently used due to the characteristics of screen content, a method to increase the priority of the transform skip mode can be applied in the production area.

[0190] Furthermore, while the present invention only mentions examples of omitting encoding / decoding processes suitable for existing live-action footage but unsuitable for game footage in the case of game content, the invention can also be implemented in a direction that additionally includes encoding / decoding processes suitable for game content through a flag value indicating game content. For example, if IBC is considered one of the tools suitable for game content, applying it additionally even in cases where the method was not originally applied can contribute to improving encoding efficiency and / or final image quality.

[0191] The present invention can be implemented by various embodiments and may include all processes of adding coding tools suitable for the production area or removing unsuitable tools by placing a flag indicating a production area different from existing live-action images.

[0192] This invention identifies the characteristics of a computer graphics-processed image and presents an image encoding method utilizing said characteristics. Images generated by graphics processing clearly exhibit differences between surrounding pixel values, and these characteristics differ from images captured by a camera. By clearly designating and signaling the graphics-processed areas within the image, encoding and decoding methods suitable for screen content and game content can be applied in those areas. Consequently, encoding and decoding efficiency can be improved by enhancing compression performance and selectively excluding unnecessary processing steps. In particular, in situations where complexity reduction is the primary goal, the burden of encoding and decoding processing can be significantly reduced while maintaining comparable compression performance.

[0193]

[0194] Encoder and decoder

[0195] It is evident that the encoding method according to the present invention can be applied identically in an encoder and a decoder. First, the encoding method according to the present invention can be used as one of the methods for generating sample values ​​of a prediction block to perform in-frame and / or inter-frame prediction in an encoder. When residual signal information is encoded by the difference value with respect to such a prediction block, the decoder can generate sample values ​​of the same prediction block using a corresponding and / or symmetric method, and obtain decoded samples by combining the difference value with such a prediction block. Additionally, as illustrated in FIG. 4 through an internal decoder (420) and a coding loop including it, this decoding process may correspond to being implemented identically within the encoder to predict the state of the decoder.

[0196] The encoding method according to the present invention described above can be implemented through an encoder as a device. The encoder as a device may be implemented by maintaining the conventional encoder structure exemplified in FIGS. 1 to 6 or by applying a certain change therefrom, but the form of implementation is not necessarily limited to that exemplified, and any form of encoder structure capable of functioning as a video encoder should be considered an encoder established by the present invention as long as it embodies the technical concept of the present invention.

[0197] In addition, the decoding method of the encoded result according to the present invention described above can be implemented through a decoder as a device. The decoder as a device may be implemented in a form that maintains the conventional decoder structure exemplified in FIGS. 1 to 6 or applies a certain change therefrom, but the form of implementation is not necessarily limited to that exemplified, and any form of decoder structure capable of functioning as a video decoder should be considered a decoder established by the present invention as long as it embodies the technical concept of the present invention.

[0198] A person skilled in the art will readily understand that a bit sequence encoded by the method and apparatus described above can be decoded by applying a method symmetric and / or in reverse order to the encoding method. In one embodiment, when reading information for decoding from the encoded bit sequence, at least one variable-length coded phrase included in the encoded bit sequence may be interpreted, and in one embodiment, the variable-length coding may be performed by an entropy coding method. The technical details and application methods of implementing such a decoding procedure will be readily understood from the encoding procedure described above.

[0199] The encoder and / or decoder described herein may correspond to a device comprising a processor and memory, preferably implemented as a computing device. The processor that may be included in the encoder and / or decoder described herein may mean one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions.

[0200] Even if the above processor is expressed in the singular for ease of understanding, a person skilled in the art will understand that the above processor may include a plurality of processing elements and / or a plurality of types of processing elements. For example, an apparatus according to one embodiment of the present invention may include a plurality of processors or one processor and one controller as the processor. In addition, the processor may be implemented by various processing configurations, such as a parallel processor or a multi-core processor.

[0201] The processor may be configured to execute an operating system (OS) and one or more software executed on the operating system. Additionally, the processor may access, store, manipulate, process, and generate data in response to the execution of the software.

[0202] The software may include a computer program, code, instructions, or a combination of one or more of these, and may be configured to control the processor to operate as desired and to issue instructions to the processor independently or collectively. The software may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave in order to be interpreted by the processor or to provide instructions or data to the processor. The software may be distributed over networked computer systems and may be stored or executed in a distributed manner.

[0203] The software described above may also be implemented in the form of program instructions that can be executed through various computer means and may be recorded or stored in the memory. The memory may be a computer-readable recording medium, and program instructions, data files, data structures, etc., may be recorded on the computer-readable recording medium alone or in combination. The program instructions stored in the memory may be based on a command system specifically designed and configured for embodiments of the present invention, or may follow a command system known and available to those skilled in the art of computer software, such as Assembly language, C, C++, Java, Python, etc. It should be understood that the command system and the program instructions derived therefrom include not only machine code such as that generated by a compiler, but also high-level language code that can be executed by a device and / or processor according to an embodiment of the present invention using an interpreter, etc.

[0204] A computer-readable recording medium constituting an device according to an embodiment of the present invention, including the memory described herein, may include a temporary or volatile recording medium that is maintained only while the processor is operating, such as a processor cache, RAM, or flash memory; or may include a relatively non-volatile or long-term recording medium such as a magnetic media such as a hard disk, floppy disk, and magnetic tape; an optical recording medium such as a CD-ROM or DVD; a magneto-optical media such as a floptical disk; or a solid-state memory; or may include a read-only recording medium such as a ROM placed on hardware; furthermore, the hardware itself, configured to perform operations equivalent to a series of program instructions by a hard-wired structure by circuit wiring, may also be considered as having each step for performing the operation implementing the embodiment of the present invention recorded by the connection and arrangement of the hardware components, so the method of connection and arrangement is equivalent to the memory. It is obvious to an ordinary technician that it can be seen.

[0205] The embodiments described above with respect to the processor and the memory are not mutually exclusive and may be selected or combined as needed. For example, a hardware device may be configured to operate as a module composed of one or more of the software to perform the operation of an embodiment of the present invention, and vice versa. As another example, in this specification, all or part of the operation assigned to a certain functional unit may be implemented by one or more of the software stored in a device according to an embodiment of the present invention (preferably in a recording medium falling within the category of the memory) and configured to be executed by the processor, in which case such a functional unit may be referred to as a functional unit "included" in the processor.

[0206]

[0207] Although the present invention has been described above with reference to the drawings and embodiments, as previously stated, the scope of protection of the present invention is not limited by the drawings or embodiments presented above, and those skilled in the art will understand that various modifications and changes can be made to the present invention without departing from the spirit and scope of the invention as described in the claims of the present invention.

Claims

1. In a video decoding method, A step of obtaining region information from a bitstream indicating that at least a portion of the current image is an arbitrary region generated by computer graphics; A step of identifying at least one filtering process applied in the prediction process of the current block corresponding to the production area based on the above area information; A step of generating a prediction block for the current block by selectively disabling the identified filtering process; and An image decoding method comprising the step of restoring the current block using the prediction block.

2. In Paragraph 1, The above production area corresponds to a part of the above current image, and A video decoding method further comprising the step of additionally obtaining region information indicating the location and size of the portion of the region from the bitstream.

3. In Paragraph 2, An image decoding method characterized by specifying the above-mentioned area information using a coding tree unit (CTU) index.

4. In Paragraph 2, An image decoding method characterized by specifying the above-mentioned area information using a tile index.

5. In Paragraph 2, A video decoding method characterized in that the above-mentioned production areas are multiple, and area information for each production area is individually signaled.

6. In Paragraph 1, The above filtering process is, It includes at least one of reference sample filtering, location-based prediction combination (PDPC), and intra prediction fusion, An image decoding method characterized in that at least one filtering process is integrally controlled by a single flag.

7. In Paragraph 6, A video decoding method further comprising the step of individually controlling at least one of the at least one filtering process based on an additional flag.

8. In Paragraph 1, It further includes a step of automatically determining whether the current block belongs to a production area using a hash block-based algorithm; An image decoding method characterized by determining the production area when the hit percentage of the above hash block exceeds a predetermined threshold.

9. In Paragraph 1, The above filtering process is, An image decoding method comprising at least one of overlapping block motion compensation (OBMC), geometric segmentation mode (GPM) blending, and bidirectional optical flow (BDOF).

10. In Paragraph 1, An image decoding method characterized by adaptively determining whether to disable the filtering process at the coding unit (CU) level.

11. In Paragraph 1, A video decoding method characterized by applying at least one of a transform skip mode and an intra-block copy (IBC) mode preferentially in the above-mentioned production area.

12. In Paragraph 1, A video decoding method characterized in that the above-mentioned region information is signaled in at least one bit sequence syntax unit among a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH).

13. In Paragraph 1, A video decoding method characterized by the fact that, when a dual-tree structure is applied, the filtering process is controlled independently for the chroma block.

14. In a video encoding method, A step of determining whether at least a portion of the current image is an arbitrary region generated by computer graphics; A step of identifying at least one filtering process applied in the prediction process of the current block corresponding to the production area based on the above judgment result; A step of generating a prediction block for the current block by selectively disabling the identified filtering process; A step of encoding the current block using the prediction block; and A video encoding method comprising the step of including area information indicating the above-mentioned production area in a bitstream.

15. In Paragraph 14, The above production area corresponds to a part of the above current image, and A video encoding method further comprising the step of including region information indicating the location and size of the aforementioned partial region in the bitstream.

16. In Paragraph 14, The above filtering process includes at least one of reference sample filtering, location-based prediction combination (PDPC), and intra-prediction fusion, and A video encoding method characterized by integrating and controlling at least one filtering process with a single flag.

17. In Paragraph 14, A video encoding method further comprising the step of determining the production area by calculating a hash value in units of 4×4 or 8×8 blocks.

18. In a video decoding device, A parsing unit configured to obtain region information from a bitstream indicating that at least a portion of the current image is an arbitrary region generated by computer graphics; A prediction unit configured to identify at least one filtering process applied in the prediction process of a current block corresponding to the production area based on the above area information, and to selectively disable the identified filtering process to generate a prediction block for the current block; and An image decoding device comprising a restoration unit configured to restore the current block using the prediction block.

19. In Paragraph 18, The above production area corresponds to a part of the above current image, and An image decoding device configured such that the above parsing unit further acquires area information indicating the location and size of the above partial area.

20. In Paragraph 18, The above filtering process includes at least one of reference sample filtering, location-based prediction combination (PDPC), and intra-prediction fusion, and An image decoding device configured such that the prediction unit is configured to integrally control the at least one filtering process based on a single flag.

Citation Information

Patent Citations

  • Video coding method, apparatus and computer program

    JP2022540536A

  • Intra block copy mode for screen content coding

    KR1020180026396A

  • Soil nailing structure for reinforcing a slope and construction method using the same

    KR1020260054974A

  • Multi-leg single-stage converters based on dual transformer

    KR102905581B1

  • KR20220163448A