Method and apparatus for video encoding and decoding based on improved geometric partitioning prediction
The method addresses the challenge of achieving high compression ratios and maintaining image quality by employing a geometric partitioning prediction technique that splits blocks into regions for flexible prediction schemes, thereby improving compression performance and image quality.
Patent Information
- Application Number
- PCT/KR2024/020654
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-18
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-26
AI Technical Summary
Existing video encoding and decoding technologies face challenges in achieving high compression ratios while maintaining image quality, particularly in efficiently processing and predicting various block shapes in digital video content.
The proposed method introduces a syntax structure for geometric partitioning prediction, allowing for the efficient encoding and decoding of blocks of various shapes by using a geometric partitioning mode (GPM) that splits blocks into regions for intra-picture and inter-picture prediction, enabling flexible prediction schemes within a single coding unit.
This approach enhances prediction precision and improves compression performance by allowing different prediction modes to be applied to different regions within a coding unit, resulting in better image quality and reduced data size.
Smart Images

Figure KR2024020654_26062025_PF_FP_ABST
Abstract
Description
Method and device for video encoding and decoding based on improved geometric segmentation prediction
[0001] The present invention relates to the field of encoding and decoding of digital video, and relates to a method for encoding and decoding digital video, a method for recording such data, and components, devices, and systems for realizing such a method.
[0002] The present invention may be in the same technical field as at least one of the digital video compression technology standards known by the names of standards such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG, or may be in the technical field for improving the inherent efficiency of the standards, or may be in the technical field for improving or replacing the standards.
[0003] Digital video encoding and decoding are widely used in various digital video applications. For example, digital television broadcasting, video transmission over communication networks, video calls / video conversations / video chats, recording and providing video content using optical media including video compact discs (VCDs), digital versatile discs (DVDs), and Blu-Rays, all processes for producing, editing, collecting, and distributing video content, and devices such as video recording devices and camcorders for various purposes, including personal, commercial, industrial, and security purposes, all depend on video encoding and decoding technologies.
[0004] Accordingly, implementations that may be referred to as digital video encoders and decoders may form part of a wide range of devices related to the generation, recording, and provision of digital video, including digital televisions, digital broadcasting systems, wireless broadcasting systems, computers in the form of notebooks / desktops / tablets, e-book readers, digital cameras, digital recording devices, digital multimedia playback devices, video game devices / terminals / consoles, mobile phones (including smartphones) with multimedia playback capabilities, equipment for video conferencing, and other devices.
[0005] The above digital video encoders and decoders can be implemented by a digital video compression standard that is widely used and understood by those skilled in the art. The digital video compression standard may include at least one of compression standards known by a standard name such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.
[0006] Video encoders and decoders can be implemented to more efficiently encode or decode digital video information while complying with the above standards, or by improving or modifying the above standards. Attempts to modify the above standards can also lead to the development of new standards. A well-known example is the so-called enhanced compression model (ECM), an attempt to improve and replace the existing H.266 / VVC standard, currently being developed by the Joint Video Experts Team (JVET), a joint international standardization group of ISO, IEC, and ITU-T.
[0007] Digital video is known to require a significant amount of information to describe its content in its uncompressed state. Therefore, recording or transmitting this information in its original form can be inefficient. Therefore, digital video is compressed using various methods prior to recording or transmission. These compression methods include lossy and lossless encoding. Lossy encoding sacrifices some image quality to achieve high compression performance, while lossless encoding sacrifices some compression performance to prevent image quality degradation. Regardless of the encoding method, there is a need to implement a technology that achieves a high compression ratio while minimizing image quality degradation to meet the growing demand for high-quality digital video within the constraints of limited memory storage capacity and communication transmission bandwidth.
[0008] As described above, the encoding process for compression requires various operations, such as spatial segmentation of digital video, segmentation and / or processing in color channels, removal of spatial redundancy, removal of temporal redundancy, tracking of motion vectors within the video, encoding of differential images, quantization, coefficient scan, run-length coding, entropy coding, and loop filtering. These encoding operations generally consume computing resources and take a certain amount of time to complete. Similarly, the decoding operation for the encoding operation also requires certain computing resources and a certain amount of time. The main goal of video encoding and decoding technology is to ensure that the above resource consumption and time consumption do not interfere with the production, recording, distribution, and viewing of digital videos.
[0009] Accordingly, the present invention provides a new technology that can contribute to at least one of the following: improvement of encoding efficiency, improvement of decoding efficiency, improvement of video quality, reduction of computational amount, reduction of software size, reduction of hardware size, and improvement of other performances related to encoding and decoding in the technical tasks in the field of video encoding and decoding.
[0010] The present invention proposes a syntax structure for dividing blocks of various shapes to be encoded / decoded and expressing the division shapes of the blocks.
[0011] According to an embodiment of the present invention for solving the above-described technical problem, a video decoding method based on geometric partitioning prediction may include a step of reading at least one piece of syntax information from an encoded bit stream, a step of obtaining information on use of a GPM (geometric partitioning mode) of a current block currently being decoded, a step of obtaining GPM index information for the current block, a step of decoding the GPM index information to obtain GPM partitioning shape information, a step of dividing the current block into a first region and a second region based on the GPM partitioning shape information, a step of obtaining prediction mode information for each of the first region and the second region, and a step of executing prediction decoding for each of the first region and the second region based on the prediction mode.
[0012] The step of decoding the GPM index information to obtain GPM segmentation shape information may include a step of executing a variable length decoding function, and the variable length decoding function may be characterized by including at least one decoding method among Golomb-Rice encoding, truncated-binary encoding, truncated-rice encoding, and entropy encoding.
[0013] The above GPM index information may be characterized in that it is derived from a list reordered according to the cost of template matching.
[0014] The above GPM index information may be characterized in that it is derived from a candidate list generated based on the GPM partitioning form of blocks spatially adjacent to the current block.
[0015] The method may further include a step of obtaining size information of a most-probable GPM list, and a step of obtaining flag information indicating that the GPM index information is for one of the most-probable GPM list and a remaining GPM list.
[0016] The step of obtaining the GPM partition shape information may be characterized in that, when the size of the optimal GPM list is 1 and flag information indicating that the GPM index information is for the optimal GPM list is obtained, the GPM partition shape information is obtained based on the optimal GPM list without decrypting the GPM index information.
[0017] The above optimal GPM list may be characterized by including at least one of an index of a GPM partition shape of surrounding blocks spatially adjacent to the current block, and another GPM partition index having a partition shape similar to the GPM partition shape.
[0018] The above prediction mode may include at least one of intra-screen prediction and inter-screen prediction, and may be characterized in that use of the intra-screen prediction is restricted in at least one of the first region and the second region.
[0019] The above prediction mode may be characterized by including at least one of IBC (intra block copy), merge prediction mode, and AMVP (adaptive motion vector prediction).
[0020] According to an embodiment of the present invention for solving the above-described technical problem, a video encoding method based on geometric partitioning prediction may include the steps of: determining whether to use GPM (geometric partitioning mode) of a current block currently being encoded; determining GPM partitioning shape information for the current block; encoding the GPM partitioning shape information to determine GPM index information; dividing the current block into a first region and a second region based on the GPM partitioning shape information; determining prediction mode information for each of the first region and the second region; performing prediction encoding by encoding the current block based on the GPM partitioning shape and the prediction mode information, and generating a residual signal between an original signal of the current block and a prediction signal; and encoding at least one of information on whether to use GPM, the GPM index information, prediction mode information for each of the first region and the second region, and information on the residual signal as syntax information in a bit stream.
[0021] The step of obtaining GPM index information by encoding the above GPM division type information may include a step of executing a variable length encoding function, and the variable length encoding function may be characterized by including at least one encoding method among Golomb-Rice encoding, truncated-binary encoding, truncated-rice encoding, and entropy encoding.
[0022] The above GPM division form may be determined based on the cost of template matching, and the template matching may be performed using a template area including a restoration area on the top and left of the current block.
[0023] The above template area may be divided by the dividing line, and the template matching may be performed using different reference template areas in the first area and the second area.
[0024] The above GPM index information may be characterized in that it is derived from a candidate list generated based on the GPM partitioning form of blocks spatially adjacent to the current block.
[0025] The method may further include a step of obtaining size information of a most-probable GPM list, and the step of determining GPM index information by encoding GPM partitioning type information may be characterized in that the GPM index information is determined based on index information when the GPM partitioning type belongs to either the most-probable GPM list or a remaining GPM list, and the step of encoding as syntax information in the bit string may be characterized in that signal information indicating whether the GPM index information belongs to either the most-probable GPM list or the remaining GPM list is further encoded in the bit string.
[0026] The above optimal GPM list may be characterized by including at least one of an index of a GPM partition shape of surrounding blocks spatially adjacent to the current block, and another GPM partition index having a partition shape similar to the GPM partition shape.
[0027] The above prediction mode may include at least one of intra-screen prediction and inter-screen prediction, and may be characterized in that use of the intra-screen prediction is restricted in at least one of the first region and the second region.
[0028] The above prediction mode may be characterized by including at least one of IBC (intra block copy), merge prediction mode, and AMVP (adaptive motion vector prediction).
[0029] It may be characterized in that a combination of the prediction mode information for each of the first region and the second region is encoded as fixed-length information in the bit string.
[0030] According to an embodiment of the present invention for solving the above-described technical problem, a video decoder device based on geometric partitioning prediction may include a bitstream parser configured to read at least one piece of syntax information from an encoded bitstream, a GPM information processing unit configured to obtain information on use of a GPM (geometric partitioning mode) of a current block currently being decoded from the syntax information, obtain GPM index information, and decode the GPM index information to generate GPM partitioning shape information, a block partitioning unit configured to divide the current block into a first region and a second region based on the GPM partitioning shape information, a prediction mode information obtaining unit configured to obtain prediction mode information for each of the first region and the second region, and a prediction decoding unit configured to perform prediction decoding for each of the first region and the second region based on the prediction mode.
[0031] The present invention proposes various block partition shapes and a method for efficiently signaling them. The present invention proposes a geometric partition mode (GPM) for both intra-picture and inter-picture prediction, and also proposes a method for signaling two PU regions within a single CU, one for intra-picture prediction and the other for inter-picture prediction. Through the prediction process of various block shapes, prediction precision is increased, which is expected to enhance compression performance.
[0032] The present invention seeks to improve video compression technology by proposing various block partition shapes and a method for efficiently signaling them. In particular, the GPM proposed in the present invention is applicable to both intra-frame and inter-frame prediction, significantly increasing prediction flexibility.
[0033] The present invention also presents an innovative signaling method that enables the application of different prediction methods to two prediction unit (PU) regions within a single coding unit (CU). This allows one region to perform intra-picture prediction, while the other region performs inter-picture prediction, allowing the selection of an optimal prediction method tailored to the characteristics of each region.
[0034] Through these diverse block-shaped prediction processes, prediction precision increases, which is expected to improve overall compression performance. Consequently, the present invention is expected to contribute to improving the quality of various multimedia services by enhancing the efficiency of image compression technology and enabling the provision of better-quality images at smaller data sizes.
[0035] Figure 1 is a conceptual diagram of a video communication system according to one embodiment of the present invention;
[0036] FIG. 2 is a conceptual diagram of the arrangement of an encoder and decoder in a real-time video streaming environment according to one embodiment of the present invention.
[0037] Figure 3 is a functional unit conceptual diagram of a video decoder according to one embodiment of the present invention;
[0038] Figure 4 is a functional unit conceptual diagram of a video encoder according to one embodiment of the present invention;
[0039] Figure 5 is a conceptual diagram of a frame type according to one embodiment of the present invention;
[0040] FIG. 6 is a conceptual diagram showing the structure of a video encoder according to another embodiment of the present invention;
[0041] Figure 7 is a conceptual diagram showing an example of a fast encoding unit for splitting an encoding tree unit in a video compression codec.
[0042] Figure 8 is an example diagram showing the aspect of block division in a video compression codec.
[0043] Figure 9 is an example diagram showing a tree structure for block division in a video compression codec.
[0044] Figure 10 is an example diagram showing various block division forms according to the application of geometric division mode in a video compression codec.
[0045] FIG. 11 is a conceptual diagram showing an image encoding area to which GPM is applied according to one embodiment of the present invention;
[0046] FIG. 12 is a flowchart showing the process of bit string syntax processing to derive a GPM index according to one embodiment of the present invention.
[0047] FIG. 13 is a flowchart showing a process of bit string syntax processing to derive a GPM index according to one embodiment of the present invention.
[0048] FIG. 14 is a conceptual diagram for the application of inter-screen and intra-screen prediction modes based on GPM according to one embodiment of the present invention; and
[0049] FIG. 15 is a conceptual diagram for the application of inter-screen and intra-screen prediction modes based on GPM according to one embodiment of the present invention.
[0050] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention.
[0051] Although terms such as “first,” “second,” etc. may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, the first component could be referred to as the second component, and similarly, the second component could also be referred to as the first component. The term “and / or” includes any combination of multiple related listed items or any of multiple related listed items, and is non-exclusive unless otherwise indicated. The listing of items in this specification is merely an exemplary description to easily explain the spirit and possible implementation methods of the invention herein, and therefore is not intended to limit the scope of embodiments of the present invention.
[0052] As used herein, "A or B" can mean "only A," "only B," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B and C."
[0053] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0054] In this specification, "at least one of A and B" may mean "only A", "only B" or "both A and B". Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted identically to "at least one of A and B".
[0055] Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”
[0056] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0057] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this specification, it should be understood that the terms "comprises" or "has" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0058] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0059] In describing the invention herein, embodiments may be described or illustrated in terms of unit blocks that perform the described function or functions. The blocks may be expressed herein as one or more devices, units, modules, parts, etc. The blocks may be implemented in hardware by one or more logic gates, integrated circuits, processors, controllers, memories, electronic components, or information processing hardware implementation methods, but not limited thereto. Alternatively, the blocks may be implemented in software by application software, operating system software, firmware, or information processing software implementation methods, but not limited thereto. A single block may be implemented by being separated into multiple blocks that perform the same function, or conversely, a single block may be implemented to perform the functions of multiple blocks simultaneously. The blocks may also be implemented by being physically separated or combined according to any criteria. The blocks may be implemented to operate in an environment where their physical locations are not specified and are separated from each other by a communication network, the Internet, a cloud service, or a communication method, but not limited thereto. All of the above implementation methods are within the scope of various embodiments that can be taken by a person skilled in the field of information and communication technology to implement the same technical idea, and therefore, any detailed implementation method should be interpreted as being included within the scope of the technical idea of the invention in this specification.
[0060] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the attached drawings. To facilitate a comprehensive understanding of the present invention, the same reference numerals will be used for identical components in the drawings, and redundant descriptions of identical components will be omitted. Furthermore, the multiple embodiments are not mutually exclusive, and it is assumed that some embodiments may be combined with one or more other embodiments to form new embodiments.
[0061]
[0062] digital video codec
[0063] Figure 1 is a conceptual diagram of a video communication system according to one embodiment of the present invention. The video communication system (100) may be configured to include at least two terminals (110, 120) connected to each other via a network (105).
[0064] In one embodiment of the present invention, the above-described FIG. 1 may refer to a block diagram for configuring a unidirectional video communication network. A first terminal (110) among the terminals may encode video data in order to transmit (111) the video data via a network (105). A second terminal (120) among the terminals may be configured to receive (121) the encoded video data via a network and decode and display the same.
[0065] In another embodiment of the present invention, the above-described FIG. 1 may refer to a block diagram for configuring a two-way video communication network. For the two-way video communication, each terminal (110, 120) may be configured to encode video data acquired by itself for video transmission (112, 122) to each other terminal via the network. Each terminal may also be configured to receive (113, 123) video data transmitted by another terminal via the network, decode the same, and display the decoded video data.
[0066] The terminals (110, 120) shown in Fig. 1 may be exemplified as devices such as server computers, personal computers, portable computers, and smartphones, depending on the embodiment, but are not limited thereto. The present invention is applicable to all environments for establishing a one-way or two-way video communication network, and it should be understood that the network (105) may be established by any means for transporting encoded video data between the terminals (110, 120).
[0067] In one embodiment of the present invention, the network (105) may refer to a wired or wireless communication network. Depending on the embodiment, the network may be configured to communicate information using any communication standard, which may include packet-based communication. The packet communication may be understood to include packets, for example, known as TCP or UDP.
[0068] However, in another embodiment of the present invention, the network (105) may be understood to include a process of information transmission using a recording medium. In this case, the configuration of the network is not limited to a communication medium, and should be understood to include a process of temporarily storing and physically transporting information on a hard disk, a solid state disk (SSD), a flash memory, a compact disc (CD), a digital versatile disc (DVD), a Blu-ray disc, and other mechanical, electronic, or optical recording media.
[0069] Any other means of information communication or transport, regardless of the method employed, can be considered within the scope of the present invention as long as it has a structure that supports the transmission and decoding of video data in an encoded state. Therefore, in addition to the examples listed above, any means of information communication or transport, whether known in the past or newly available, can fall within the scope of application of the present invention.
[0070] FIG. 2 is a conceptual diagram illustrating the arrangement of an encoder and decoder in a real-time video streaming environment according to one embodiment of the present invention. The streaming system (200) illustrated in FIG. 2 can be applied to video data communication networks, including, for example, digital broadcasting, video telephony, and video conferencing. However, it should be noted that technical structures identical or similar to the streaming system can be equally applied even when information is transmitted via a recording medium, as described above.
[0071] According to one embodiment of the present invention, the streaming system may include a video source (210) that generates a video stream. The video source may include a digital video acquisition means (212), which may be, for example, a digital camera or other device, for acquiring uncompressed raw video. The raw video stream (215) may have a large capacity and may therefore be compressed by a video encoder (217) coupled or connected to the video source.
[0072] The above encoder (217) may be configured as a means including hardware, software, or a combination of the two, configured to implement an image encoding method and / or an implementation method thereof according to one embodiment of the present invention.
[0073] Through the encoder (217), an encoded bitstream (219) having a reduced capacity compared to the original video stream can be output. The bitstream (219) can be provided in real time via a relay device, which may be referred to as a streaming server (220), and / or can be stored in a recording medium (225) of the streaming server (220) for subsequent use.
[0074] The streaming system (200) may include at least one streaming client (230, 240) that connects to the streaming server (220) to receive the encoded bit string (229) in real time or obtain it later. The streaming client may include a video decoder (232) that obtains the encoded bit string (229) (which may also be considered as a copy of the bit string (219) received by the streaming server), decodes the bit string (229), and outputs the resulting video data as video data in a form that can be displayed by a display (235) or other visual, auditory, or other sensory display means.
[0075] As described above, the functions for encoding and decoding video data are collectively called a coder-and-decoder, or video codec.
[0076] FIG. 3 is a conceptual diagram of a functional unit of a video decoder according to an embodiment of the present invention. As shown in FIG. 3, a receiving unit (310) can receive at least one encoded video data to be decoded by a decoder (305). In an embodiment of the present invention, the encoded video data may be independent for each reception, and the decoding procedure of each independent video data may be independent from the decoding procedure of other video data. The encoded video data may be received by the receiving unit (310) through a hardware or software connection (315) to a device storing the same, and as described above, the storing device may be a type of streaming server located at the other end of a communication network, or may mean a physical recording medium, but is not limited thereto.
[0077] The above-described receiving unit (310) can receive the encoded video data together with other data accompanying it, such as encoded audio data or other auxiliary data, and each of the data can be separated from the video data and provided to an appropriate processing function unit (312) other than the video decoder.
[0078] When the video data is provided through a communication network, a buffer memory (320) may be coupled between the receiving unit (310) and the decoder (305) to minimize delay and disconnection according to the network environment. The buffer memory (320) may refer to a computer-readable recording medium that temporarily stores the received video data and stably supplies it to a parser (330) corresponding to the input terminal of the decoder (305). However, if the bandwidth of the communication network is sufficient, if video data is read from a recording medium in a local location that is not physically separated, or if the possibility of communication delay is not predicted in other environments, the buffer memory may be unnecessary.
[0079] The video decoder (305) may include the parser (330) as its input terminal to interpret the encoded video data. The parser may perform a function of separating (parsing) a plurality of pieces of information stored in the form of a bit string in the encoded video data according to a predetermined rule, and, if necessary, performing an entropy decoding (335) of entropy-coded video data, thereby performing a function of reconstructing symbols (338), which are paragraphs of video encoding information. The symbols (338) may include all information for controlling the operation of the decoder (305), and / or may further include information for controlling a device that is attached to and operable with the decoder (305), such as a display device. Control information for controlling the above display device may include information in a format called supplementary enhancement information (SEI) or video usability information (VUI).
[0080] As described above, the parser (330) may be configured to perform entropy decoding (335) of the encoded video data. The entropy encoding method of the encoded video data may vary depending on the encoding standard, and decoding may be performed accordingly. Representative examples of the entropy encoding standard may include variable length coding, Huffman coding, and arithmetic coding, and each of the encoding methods may be a context-adaptive or context-sensitive method depending on the standard, and may also be based on principles widely known to those skilled in the art.
[0081] The parser (330) may be configured to extract at least one picture from the encoded video data. The definition of the picture may vary depending on the encoding standard, and depending on the standard, one or more of the examples listed below may correspond simultaneously and overlappingly. The picture may be grouped, defined, and / or divided into encoding / decoding units such as, for example, groups of pictures (GOPs), pictures / frames, tiles, slices, macroblocks, blocks, subblocks, transform units (TUs), and prediction units (PUs).
[0082] The parser (330) may be configured to extract encoding information, such as transform coefficients, quantization parameters (QPs), and / or motion vectors, from the encoded video data. The parser (330) may be configured to perform entropy decoding (335) and parsing operations on the video data received from the buffer memory, and to selectively decode symbols (338) representing the encoding information. In addition, the parser (330) may be configured to selectively supply a specific symbol (338) to a specific decoding function unit within the decoder (305), such as an inverse quantization and inverse transform unit (340), an intra prediction unit (350), an inter prediction unit (355), or a loop filter unit (360). Control of such information supply can be determined by the information sequence contained in the encoded video, and may vary depending on the encoding standard, and is not limited within the scope of the embodiments of the present invention, and is not described in detail in this conceptual diagram.
[0083] The decoder (305) may be comprised of a number of conceptual functional units that receive and process the encoded information provided by the parser (330). It should be readily apparent that these conceptual functional units may be combined or further subdivided, depending on implementation needs. For example, they may be further separated for ease of implementation, or integrated into one for operational efficiency. In any case, each functional unit may be configured to perform close interaction with each other. However, despite the possibility of such integration or separation, the following description will be given as a combination of conceptual functional units to illustrate the video data decoding procedure applied as an embodiment of the present invention.
[0084] The decoder may include an inverse quantization and inverse transformation unit (340). The inverse quantization and inverse transformation unit (340) may be configured to receive encoding information including a method to be used for numerical transformation (transform), a block size, quantization coefficients for recovering quantized information, and distinction information of a quantization matrix that simplifies and represents the quantized coefficients from the parser (330), and may be configured to output block values (341) that can be input to an aggregator (370) as a result of processing the encoding information.
[0085] In one embodiment of the present invention, the output values of the inverse quantization and inverse transformation unit (340) may include intra-prediction encoded block values. The intra-predicted block values may refer to values that can be decoded using prediction information within a picture currently being decoded, for example, the current frame, but not using prediction information from a previously decoded picture, for example, the previous frame.
[0086] The prediction information within the current picture may be provided by the intra prediction unit (350). According to an embodiment of the present invention, the intra prediction unit (350) generates a block value of the same form as the block being decoded as the prediction information by using picture information of a spatially adjacent area derived from a picture currently being decoded and of which decoding has been partially completed. The picture information may be provided (381) from a buffer for the current picture, a so-called line buffer (380). The merging unit (370), according to an embodiment, may be configured to merge the prediction information (351) generated by the intra prediction unit (350) with the block values (341) provided by the inverse quantization and inverse transformation unit (340).
[0087] In another embodiment, the output values of the inverse quantization and inverse transformation unit (340) may include block values subjected to motion compensation as inter-prediction encoded block values, and in some cases, block values subjected to motion compensation. In this case, the inter-prediction unit (355) may extract and use sample information (386) used for motion-based prediction from a reference picture buffer (385). The information (356) derived by performing motion compensation on the sample information based on the symbols (338) included in the block values as the output values may be configured to be merged with the block values (341) provided by the inverse quantization and inverse transformation unit (340) by the merger unit (370). In this case, the block values (341) may be referred to as so-called differential or residual values.
[0088] The position information within the memory used by the inter prediction unit (355) to extract the sample information from the reference picture may be determined by a motion vector provided to the inter prediction unit (355) and composed of a combination of symbols (338) for indicating, for example, X, Y, and other specific points of the reference picture. The inter prediction unit (355) may also include a function for interpolating and using the sample values when a so-called 'subsampling'-capable motion vector is provided, and may further include a function for predicting and reinforcing the value of the motion vector.
[0089] The output values (371) of the above merging unit (370) may be provided to the loop filter unit (360) and processed by various loop filtering methods. The loop filter unit (360) may be configured to receive not only the block unit output (371) of the merging unit (370) but also the symbol (338) provided from the parser (330) to control its operation. The output of the loop filter unit (360) may be output to an external display means such as the display device through an output connection (390), but may be stored (361) in a line buffer (380) for use in prediction for interpreting a subsequent intra or inter coded block value, and may also be stored in a reference picture buffer (385) through this.
[0090] Certain pictures, such as frames, after their decoding is completed, can be utilized as reference pictures for performing predictive decoding in a subsequent decoding process. A picture (or frame) can be gradually accumulated in a line buffer (380) and decoded. When the decoding of a frame is completed, the contents of the line buffer (380) are transferred (383) to the reference picture buffer (385), and a new line buffer (380) can be allocated for decoding the new frame.
[0091] The above video decoder (305) may be configured to perform a decoding operation according to a predetermined video compression technique that may be documented by various international standards or commercial standards. The standards may include, for example, international standard recommendations such as H.264, H.265, and H.266 defined by the International Telecommunication Union Standardization Sub-Division (ITU-T). Those skilled in the art will understand that each of the above recommendations is equivalent to an international standard jointly defined by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC). The encoded video data may comply with a specific bitstream syntax defined by the relevant standards, as defined and required by the video compression standard documents and standard documents, and specifically by the profiles and levels specified within such documents. In addition, the complexity of the encoded video data may be limited to a certain level to comply with the profiles and levels. For example, a profile or level may be configured to limit a maximum picture size, a maximum decoding speed, and a maximum reference picture size. These limitations may, in some embodiments, also be further limited by metadata signals for a hypothetical reference decoder (HRD) and HRD buffer management included in the encoded video data.
[0092] According to one embodiment of the present invention, the receiver (310) may receive additional redundant data along with the encoded video. The additional data may be considered as part of the encoded video data. The additional data may include information that may be used by the decoder (305) to properly decode the data or to more accurately reconstruct an image that approximates the original image. The additional data may be provided in the form of, for example, layers for temporal, spatial, or signal-to-noise ratio (SNR) enhancement, redundant slices, redundant pictures, and forward error correction codes.
[0093] Figure 4 is a functional unit conceptual diagram of a video encoder according to one embodiment of the present invention. The encoder (405) may be configured to receive original video information (402) from a video source (401) and perform encoding.
[0094] The original video information (402) may have any suitable bit depth, for example, 8 bits, 10 bits, 12 bits, etc. In addition, the original video information (402) may have any suitable color space, for example, R / G / B, Y / U / V, Y / Cb / Cr, etc. In addition, the original video information (402) may have any suitable sampling structure corresponding to the color space, for example, Y / Cb / Cr 4:2:0, Y / Cb / Cr 4:4:4, etc. The original video information (402) having such a predetermined format may be provided to the encoder in the form of a digital video stream.
[0095] In a one-way video communication network, the original video information (402) can be obtained from a recording medium storing a previously prepared video source. In a two-way video communication network, the original video information (402) can be obtained from an image acquisition device, such as a camera, that generates at least one video transmission stream included in the two-way video communication.
[0096] The video data including the above original video information (402) may be configured as a plurality of pictures configured to simulate motion by playing them in chronological order. In addition to pictures, the pictures may also be expressed in concepts such as frames. The pictures may include one or more samples depending on the type of sampling structure, color space, etc. being used. Those skilled in the art will understand that the samples are closely related terms to pixels in digital images. The operation of the encoder will be described below with reference to such samples.
[0097] According to one embodiment of the present invention, the encoder (405) may be configured to encode and compress pictures (and / or grouped or segmented information) constituting the original video information (402) in real time (or according to other temporal requirements required according to the implementation method) into the form of encoded video information.
[0098] In the encoder (405), the control unit (450) may be a functional unit configured to control an appropriate encoding speed. The control unit (450) may be configured to control other functional units and be functionally coupled to the following functional units as described below. The parameters set by the control unit (450) may include parameters related to bitrate control, such as picture skip, quantizer, and variable values for applying image quality optimization techniques, and may also include values such as picture size, group of pictures (GOP) structure, and maximum search range of motion vectors. A person skilled in the art will be able to understand various other functions that the control unit (450) may have, and such other functions may be added or removed according to the design of a video encoder optimized for individual system design.
[0099] According to an embodiment of the present invention, the encoder (405) may be configured to operate in a structure such as a "coding loop" well known to those skilled in the art. To simplify the description by way of example, the coding loop may be configured with an internal encoder (so-called "source coder") (410) which is responsible for receiving a picture to be encoded and generating symbols based on at least one reference picture that has been encoded in the past, and a local decoder (420) configured to be connected to the internal encoder. The local decoder (420) may be configured to perform an operation to reproduce sample data to be generated by a decoder (490) located at an actual remote location that will receive encoded video information from the encoder (405) by receiving an output of the internal encoder (410).
[0100] Video data composed of sample data reconstructed by the internal decoder (420) may be configured to be input to the reference picture buffer of the encoder (405). As described above, the internal decoder (420) is implemented to reproduce the result output by the encoder (405) and to be decoded by a remote decoder, so the video data recorded in the reference picture buffer may also be identical in bit unit to the information of the reference picture buffer of the remote decoder. That is, the prediction function unit that may be included in the encoder (405) may read the same values as the sample values of the previous frame that the decoder will later refer to in the decoding process from the reference picture buffer of the encoder (405).
[0101] As described above, the principle of achieving matching of the reference picture buffers between the encoder (405) and the decoder (490) by means of the internal decoder (420) on the encoder (405) side is well known to those skilled in the art, and a method of responding to an environment in which such an environment is not guaranteed (e.g., information loss due to communication failure, etc.) can also follow what is known to those skilled in the art.
[0102] An embodiment of the operation method of the internal decoder (420) has been described in detail above with reference to FIG. 3. The decoder of FIG. 3 may be regarded as the aforementioned "remote" decoder (490). The internal decoder (420) may be implemented excluding lossless encoding and decoding sections such as the parser (330) or entropy decoding (335). This is because the internal encoder (405) is implemented to simply reproduce the operation of a decoder located at a remote location, and thus may directly decode symbols without requiring a process of compressing and then decompressing symbols. Accordingly, the functional units preceding the parser and entropy decoder as shown in FIG. 3 may not be provided or may be implemented at least partially.
[0103] As described above, according to a preferred embodiment of the present invention, any decoder function (excluding a parser and an entropy decoder) present in the decoder can naturally exist as a substantially identical function in the corresponding encoder (405).
[0104] The operation of the encoding function unit that may be included in the above encoder (405) can be considered as the reverse operation of the decoder function unit. Therefore, the embodiment can be explained by performing the operation of the decoder function unit in reverse. For example, a quantization and transform function unit corresponding to the inverse quantization and inverse transform unit may be provided, and an inter prediction encoding unit corresponding to the inter prediction unit may be provided. In addition, some additional explanations will be added.
[0105] The internal encoder (410) may be configured to perform encoding on input picture information, for example, an input frame, by a prediction encoding method executed by a prediction encoding unit (440) that operates by referencing at least one reference picture information, for example, at least one temporally previous encoded picture (or frame) from a reference picture buffer (430) from video data designated as a reference frame. In this case, the encoder (405) may be configured to encode a differential between blocks of samples constituting the input picture and blocks of samples constituting the reference picture.
[0106] The internal decoder (420) can decode video data that can be designated as the reference picture from symbols generated by the internal encoder (410). As described above, since the video data is identical to the decoding operation performed by the remote decoder, the video data used as the reference picture may be provided to the encoder (405) in a form that has undergone lossy compression and has been partially damaged, and this operation may be intended to ensure operational consistency with the decoder.
[0107] The prediction encoding unit (440) may be configured to perform a prediction search operation within the encoder (405). The prediction search operation may refer to an operation corresponding to the inter prediction or intra prediction described in the description of the decoder. For picture information that is input and scheduled to be newly encoded, the prediction unit may access the reference picture buffer (430) to retrieve information such as a motion vector, a block shape, and metadata that may include the same, which are information indicating a point of a reference picture that can function as prediction reference information suitable for the new picture information, and a sample block to be actually referenced. The above prediction encoding unit (440) may operate on the basis of the so-called "sample block by pixel block" criteria in order to obtain appropriate prediction reference information. According to one embodiment of the present invention, at least one prediction reference information designating at least one reference picture information stored in the reference picture buffer (430) may be designated for the input picture, as determined based on the search results obtained by the prediction encoding unit (440).
[0108] In one embodiment of the present invention, the control unit (450) may be configured to manage the overall encoding operation of the internal encoder (410), including setting parameters used to encode video data.
[0109] All outputs of the above-described functional units may be subjected to entropy encoding (460) in order to be finally output. The entropy encoding (460) may include various entropy coding techniques, such as variable length coding, Huffman coding, and arithmetic coding, for the symbols generated by the various functional units as described above, and each encoding method may be a context-adaptive or context-sensitive method according to the standard, or may be based on principles widely known to those skilled in the art. Such entropy encoding (460) can typically achieve lossless compression, and thus can be configured to convert at least one symbol generated by the functional units into encoded video data.
[0110] The above control unit (450) may, when controlling the operation of the encoder (405), apply to each picture (or frame) the type of encoding in which a specific picture is encoded during the encoding period. Depending on the type, the method by which the picture is encoded may be affected. Depending on the embodiment, the type may include what is categorized as the following "frame type."
[0111] Fig. 5 is a conceptual diagram of a frame type according to one embodiment of the present invention. The following description will be made with reference to Fig. 5.
[0112] An intra (“I”) picture (510) may refer to a picture that can be encoded and decoded using only its own information without referring to other picture information in video data through predictive encoding. The “I” picture may be designated by names such as a key frame, an independent / instantaneous decoder referh (IDR) frame, and a clean random-access (CRA) frame, depending on the video encoding standard, and the “I” pictures designated by the various names as described above may have various modifications and application methods as permitted by each standard and may be partially different from each other. In addition to those listed above, various application methods for implementing the “I” picture may be by various methods that are already known to those skilled in the art or may be newly provided.
[0113] A prediction ("P") picture (520) may refer to a picture that can be encoded and decoded through intra or inter prediction based on at least one prediction information and / or a motion vector that designates at least one reference picture to predict sample values of a block constituting the picture. The "P" picture may be configured to refer to only one reference frame, or may be configured to refer to one or more reference frames, according to a video encoding standard. When referring to more than one reference frame, sample information and / or associated metadata derived from multiple reference pictures may be used to reconstruct a single block. However, in common cases, a picture designated as a "P" picture may be understood as a picture that performs reference only to a temporally preceding picture.
[0114] A bidirectional prediction ("B") picture (530) may refer to a picture that can be encoded and decoded through intra or inter prediction based on at least one piece of prediction information and / or a motion vector that designates at least two reference pictures in order to predict sample values of blocks constituting the picture. In a common case, a picture designated as the "B" picture is distinguished from a picture designated as the "P" picture, and may be understood as a picture that performs a reference without being limited to a temporally preceding picture.
[0115] Video data may be spatially divided into a plurality of sample blocks during the encoding and decoding process, and encoding may be performed in units of the blocks. The block units may include, but are not limited to, sizes such as 4x4, 8x8, 4x8, or 16x16 in units of horizontal / vertical pixels, as is widely known. The block may be encoded using a predictive encoding method with reference to any other (already encoded) blocks, as permitted and / or restricted by the type specified for each picture in which the block is included. For example, the blocks of the "I" picture (510) may be encoded without using a predictive encoding method, or with reference to blocks that have already been encoded within the same partial picture. That is, only the so-called intra prediction method may be used. In contrast, the "P" picture (520) may further reference a reference picture encoded in at least one previous time unit, and thus, inter prediction may also be used for encoding along with intra prediction. In the case of a "B" picture (530), reference can be made not only to a previously encoded picture in the encoding order but also to a later reference picture in terms of time unit. However, it is widely known that there may be blocks within a "P" picture or a "B" picture that are encoded without relying on predictive encoding.
[0116] The above video encoder (405) may be configured to perform encoding operations according to a predetermined video compression technique that may be documented by various international standards or commercial standards. Examples of the above standards may include all of those described in the above decoder.
[0117] According to one embodiment of the present invention, the transmitter (470) may buffer the encoded video data generated by the entropy encoding in order to provide / transmit the video data (ultimately to a remote decoder (490)) to a device storing the encoded video data via a hardware or software connection (495). According to an embodiment, when providing / transmitting the encoded video data from the video encoder (405), the transmitter (470) may receive and merge other data accompanying the encoded video data, for example, encoded audio data or other auxiliary data, from a separate source (480).
[0118] According to one embodiment of the present invention, the transmitter (470) may be configured to transmit additional data along with the encoded video. The additional data may be considered part of the encoded video data. The additional data may include information that can be used by a decoder to properly decode the data or to more accurately reconstruct an image that approximates the original image. Examples of the additional data may include all of the examples previously presented with respect to the receiver (310) of the decoder.
[0119] The present invention can be implemented by a digital video compression standard that is widely used and understood by those skilled in the art as described above. The digital video compression standard may include at least one of compression standards known by the standard name such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.
[0120] Fig. 6 is a conceptual diagram illustrating the structure of a video encoder according to another embodiment of the present invention. What is depicted in Fig. 6 may be a rough structure of a video encoder widely known as a standard code such as ITU-T H.266 and ISO / IEC 23090-3, and also known as MPEG-I Part 3 or versatile video coding (VVC).
[0121] According to FIG. 6, a video encoder (605) may be configured to receive uncompressed and unencoded original video data (601) as input and output an encoded bit stream (602). The video data (601) may be supplied directly to a luma mapping unit (610a) when intra-encoded, or may be supplied to a luma mapping unit (610b) via an inter-prediction unit (620) including motion vector extraction. In the intra-encoded case, the mapped luma signal may be supplied to an output merger (606) by selecting (608) at least one of an intra-prediction encoded signal via an intra-prediction unit (625) or an inter-prediction encoded signal output from the luma mapping unit (610b) via the inter-prediction unit (620). The result of the above output merger can be applied to a chroma scaling unit (615). (The operation of the luminance signal mapping unit (610) and the operation of the chroma scaling unit (615) are collectively referred to as a luma mapping / chroma scalaing (LMCS) process.) The reduced chroma signal can be provided to a transform unit (630), and the transform unit (630) can perform an adaptive color transform, particularly on the chroma signal. The coefficients derived as a result of the transform are applied to a quantization unit (640) and quantized. As a result, lossy compression is achieved, and the result of the lossy compression can be output as a bit string (602) through a multi-hypothesis CABAC (650), which is a lossless compression method.
[0122] Meanwhile, the result of the lossy compression may actually enter the decoding process by going through the processes of inverse quantization (645), inverse transform (635), and luminance signal expansion (617) to generate a coding loop. The result of the luminance signal expansion may be supplied to the internal merger (607) together with the result of selecting (608) at least one of the previously generated intra prediction encoding signal or inter prediction encoding signal. The result of the internal merger may go through inverse luma mapping (618), and then may go through processing such as a deblocking filter (660), sample adaptive offset (SAO) (670), and an adaptive loop filter (ALF) to reproduce the image quality improvement process in the decoder. The result of reproducing the operation in the decoder as described above is applied to the reference picture buffer (690) and can be reused for prediction encoding by the inter prediction unit (620).
[0123] The present invention can also be utilized by or incorporated into an enhanced compression model (ECM), which is an implementation of a next-generation video codec currently being developed by the Joint Video Experts Team (JVET), an international standardization expert organization. The enhanced compression model can include an enhanced intra prediction coding method, an enhanced inter prediction coding method, an enhanced transform and transform coefficient coding method, an enhanced adaptive loop filtering method, a bilateral filtering method, a new sample adaptive offset (SAO) method for improving image quality, an extended entropy coding method, and an improved gradual decoding refresh (GDR) technique.
[0124]
[0125] CTU and CU split configuration
[0126] FIG. 7 is a conceptual diagram showing an example of a fast coding unit for splitting a coding tree unit in a video compression codec. Referring to FIG. 7, in a digital video codec including the VVC, a coding / decoding unit of a video, for example, a coding tree unit (CTU) (700) that may have a predetermined height (710) and width (720), may be configured to be split into various square and / or rectangular shapes to generate a coding unit (CU) (e.g., 750, etc.) that becomes an individual coding / decoding unit block. At this time, the predetermined height (710) and width (720) may be determined to be the same or different, and depending on the embodiment, the height (710) and width (720) may be determined to be the same or different, with a length of any one of 32, 64, 128, and 256 pixels. According to an embodiment, the size of the CTU (700) may be encoded as part of a high-level syntax (HLS), such as information transmitted to a decoder, and included in a packet including a syntax such as a sequence parameter set (SPS).
[0127] Fig. 8 is an exemplary diagram showing an aspect of block segmentation in a video compression codec. Referring to Fig. 8, in a digital video codec including the VVC, at least three types of methods for segmenting the CTU block can be provided. Fig. 8 exemplarily shows a quad tree (QT) (810), a horizontal or vertical binary tree (BT) (821, 822), and a horizontal or vertical ternary tree (TT) (831, 832). The segmentation method can be applied recursively and repeatedly, and thus can be used to derive a multi-layer segmentation structure similar to that exemplified in Fig. 7. Consequently, at least one CU obtained by segmenting the CTU by recursively applying the method exemplified in Fig. 8 at least once can be understood to be at least square or rectangular.
[0128] FIG. 9 is an exemplary diagram showing a tree structure for block division in a video compression codec. Referring to FIG. 9, in a digital video codec including the VVC, an example is illustrated in which the first single CTU (900) is recursively divided by various block division methods including the one illustrated in FIG. 8. In the example of FIG. 9, the CTU (900) is quarter-divided (QT) in the first division (951). Subsequently, through the second division (952), the third division (953), and the fourth division (954), the CTU (900) is recursively divided by various methods including quarter-divided (QT), bi-divided (BT), and tri-divided (TT), thereby forming a multi-type tree (MTT), through which a total of 16 individual CUs (901 to 916) illustrated in FIG. 9 can be ultimately formed. That is, the process of dividing the CTU (900) into the CUs (901 to 916) can be depicted in the form of a divided area as in (a) of FIG. 9, or in the form of a tree according to the MTT as in (b) of FIG. 9.
[0129] Table 1 shows an example of the coding unit bitstream syntax (CU bitstream syntax) used to express the QT and MTT in the VVC, which is one of the video compression codecs.
[0130] coding_tree(x0, y0, cbWidth, cbHeight, qgOnY, qgOnC, cbSubdiv, cqtDepth, mttDepth, depthOffset, partIdx, treeTypeCurr, modeTypeCurr){if((allowSplitBtVer || allowSplitBtHor || allowSplitTtVer || allowSplitTtHor || allowSplitQT) && (x0+cbWidth<=pps_pic_width_in_luma_samples) && (y0+cbHeight<=pps_pic_height_in_luma_samples))split_cu_flag…if(split_cu_flag){if((allowSplitBtVer || allowSplitBtHor || allowSplitTtVer || allowSplitTtHor) && allowSplitQT)split_qt_flagif(!split_qt_flag){if((allowSplitBtHor || allowSplitTtHor) && (allowSplitBtVer || allowSplitTtVer))mtt_split_cu_vertical_flagif((allowSplitBtVer && allowSplitTtVer && mtt_split_cu_vertical_flag) || (allowSplitBtHor && allowSplitTtHor && !mtt_split_cu_vertical_flag))mtt_split_cu_binary_flag}}}
[0131] Referring to Table 1, the "split_cu_flag" may be a flag signal value indicating whether the block is split or has been completed, the "spl_qt_flag" may be a flag signal value indicating whether the block is split in a quadruple (QT) form, the "mtt_split_cu_vertical_flag" may be a flag signal value indicating whether the form of the MTT partition for splitting the block is vertical or horizontal, and the "mtt_split_cu_binary_flag" may be a flag signal value indicating whether the form of the MTT partition for splitting the block is a binary (BT) form or a tri-partition (TT) form.
[0132] FIG. 10 is an exemplary diagram illustrating various block partitioning forms according to application of a geometric partitioning mode in a video compression codec. Referring to FIG. 10, in a digital video codec including the VVC, when at least one prediction method among inter-screen prediction and / or intra-screen prediction is used, a block may be partitioned using a geometric partition mode (GPM). In this case, the GPM may refer to any mode configured to re-partition a CU having a square and / or rectangular shape into a partition shape other than a rectangular shape (for example, a triangle or a trapezoid may be included).
[0133] Table 2 shows an example of the value composition of the bit string syntax used to express the GPM in the VVC, which is one of the video compression codecs.
[0134] merge_gpm_partition_idx01234567angleIdx00222233distanceIdx13012301merge_gpm_partition_idx89101112131415angleIdx33444455distanceIdx23012301merge_gpm_partition_idx1617181920212223angleIdx558811111111distanceIdx23130123merge_gpm_partition_idx2425262728293031angleIdx1212121213131313distanceIdx01230123merge_gpm_partition_idx3233343536373839angleIdx1414141416161818distanceIdx01231312merge_gpm_partition_idx4041424344454647angleIdx1819191920202021distanceIdx31231231merge_gpm_partition_idx4849505152535455angleIdx2121242427272728distanceIdx23131231merge_gpm_partition_idx5657585960616263angleIdx2828292929303030distanceIdx23123123
[0135] As an example, in the above VVC, a total of 64 types of GPMs may be allowed. At this time, the GPM may be defined by a method of specifying the angle and position of the dividing line by a bit string syntax structure as shown in Table 2. Specifically, in Table 2, a flag signal for selecting a specific GPM may be "merge_gpm_partition_idx", and the shape of the GPM may be determined according to an index value of the signal. According to one embodiment, the "merge_gpm_partition_idx" bit string syntax may be recorded in a bit string as fixed-length binary number information of 6 bits in length with a minimum value of 0 and a maximum value of 63. In correspondence with each index, an angle ("angleIdx") and a position ("distanceIDdx") of the dividing line may be determined, and it can be understood that this means each of the various types of GPMs shown in FIG. 10.
[0136] According to one embodiment, when one GPM mode is selected by signaling an index as described above, at least two prediction units (PUs) are generated within the CU, and a merge index value for each of the PUs is additionally transmitted, thereby enabling intra-screen and / or inter-screen prediction to be performed. For example, a merge candidate list may be constructed, and merge index information (e.g., motion vectors (MVs)) used in the GPM mode may be derived by transmitting the candidate list. In this case, according to an embodiment, the merge candidate list may be referred to as a GPM candidate list.
[0137]
[0138] Composition of the present invention
[0139] The present invention proposes a method for efficiently signaling and recording various block division forms as bit-string syntax. In particular, according to one embodiment of the present invention, the division form of the GPM, i.e., the angle and division line position information, can be expressed by a bit-string syntax expressing the corresponding division form information, and further, a bit-string syntax expressing information for intra-screen prediction and / or inter-screen prediction for each divided region can be provided by various embodiments.
[0140] More specifically, the present invention proposes a GPM method for both intra-frame prediction and / or inter-frame prediction. According to one embodiment, a method for configuring a bitstream syntax such that, for multiple PU regions within a single CU, at least one region can perform intra-frame prediction and at least one other region can perform inter-frame prediction can be implemented.
[0141] The various embodiments of the present invention described below are intended to explain the technical idea of the present invention and may correspond to a preferred embodiment of the present invention. However, the present invention is not limited to the embodiments described below, and it will be apparent to those skilled in the art that the present invention can be implemented by various other implementation methods within the scope of the technical idea of the present invention. In the embodiments described below, the first and second embodiments may represent an implementation method for efficiently signaling information about the block division shape of the GPM itself, and the third embodiment may represent an implementation method for efficiently signaling which prediction mode was used for each division area. However, the present invention does not exclude implementation methods that are identical or similar in structure or purpose.
[0142]
[0143] Example 1 (GPM candidate list)
[0144] According to one embodiment of the present invention, a method for efficiently transmitting one GPM selected from among a plurality of GPMs may be provided. Fig. 11 is a conceptual diagram illustrating an image encoding area to which a GPM is applied according to one embodiment of the present invention. Referring to Fig. 11, when a diagonal boundary line (1120) occurs within a specific CU (1110) in an image to be encoded, it can be seen that sufficiently effective predictive encoding cannot be achieved even if the diagonal boundary line (1120) is re-divided using a conventional rectangular MTT-based CTU division method. Therefore, in the exemplary CU (1110), a triangular GPM division method (1130) configured to closely follow the shape of the diagonal boundary line (1120) may be applied. In this case, an embodiment of signaling information that an encoding method such as the GPM (1130) has been selected for the CU (1110) according to one embodiment of the present invention through a bit string syntax will be described below.
[0145] Referring to the example of FIG. 11, the diagonal boundary line (1120) does not exist only in a specific CU, for example, the current encoding / decoding target CU (1110), but may also appear similarly in other surrounding CUs depending on the characteristics of the encoding target image. In this case, the diagonal boundary line (1120) may have a similar direction across a plurality of other CUs. By utilizing this point, according to an embodiment of the present invention, as a method of efficiently designating the shape of the GPM (1130) of the current CU (1110), whether or not the surrounding CUs use GPMs and the shape of the used GPMs can be utilized. That is, while the conventional GPM candidate list configuration means a merge candidate list available in the GPM, according to an embodiment of the present invention, the GPM candidate list may mean a candidate list for the split shape of the GPM itself.
[0146] According to one embodiment of the present invention, the GPM candidate list can be constructed by citing GPM partition index information of other CUs that are spatially adjacent to the current CU (1110). According to one embodiment of the present invention, the GPM candidate list can be constructed in the following manner. First, it is assumed that the maximum number of the list can be signaled and determined by the high-level bitstream syntax (HLS), or can be a fixed number that is pre-specified and shared between the encoder and decoder, and this can be selected according to the implementation method.
[0147] The above GPM candidate list can be constructed by at least one or more of the various methods described below.
[0148] 1) GPM partition index used in a spatially adjacent block (e.g., CU). In this case, the spatially adjacent block may include at least one of the blocks adjacent to the left, top, top left, top right, and bottom left of the current block.
[0149] 2) Another GPM partition index that is adjacent to the GPM partition index of the spatially adjacent block. For example, at least one other index that is adjacent by ±a from the index extracted in 1) above may be included.
[0150] 3) GPM partition index based on past history.
[0151] 4) Basic GPM index. For example, it may include at least one GPM partition index that is commonly used by general probability.
[0152] Depending on the implementation method, at least one of the above methods may be used, or a plurality of methods may be used, and when there are multiple GPM indices derived by at least one of the above methods, an arbitrary priority may be determined between the above methods, and an index up to the maximum number of the above list may be configured to be reflected in the GPM candidate list based on the priority.
[0153] In the case of the GPM indices corresponding to the above 2), it can be understood that the characteristics in which each GPM indices are adjacent to each other based on angle and distance are utilized. This will be explained with reference to Table 2 above. The GPM segmentation angles increase little by little as the index increases, and for example, the angle at index 0 and the angle at index 2 are similar to each other. In addition, it can be seen that index values having different positions of the segmentation line at the same angle of GPM segmentation are adjacent, and for example, index 2 and index 5 correspond to GPM forms having the same diagonal segmentation angle but different segmentation positions. Therefore, an additional segmentation index can be added to the GPM candidate list using index values that are adjacent ±a to the GPM segmentation index values of the surrounding CUs obtained in the above 1), for example, +1 or -1. For example, if the GPM segmentation index of the surrounding CU obtained in 1) above is “10”, in 2) above, the indices “9” and / or “11” may be additionally added to the GPM candidate list. In the present invention, the difference value (a) is not limited, and may be determined as any numerical value that can express an angle similar to the reference GPM segmentation shape and / or another shape (angle and / or segmentation line position) close to the same angle, depending on the implementation method, and if the difference value has a predetermined range, at least one or all of a plurality of similar GPM shapes belonging to the range may be included in the GPM candidate list. Depending on the embodiment, the ranges of incremental (+) numbers and subtractive (-) numbers may be determined differently.
[0154] In the case of the GPM index corresponding to the above 3), according to one embodiment of the present invention, a data space of a first in first out (FIFO) structure may be allocated to an encoder and / or a decoder, and at least one GPM split index information generated during an encoding / decoding process may be stored / input into the data space as a past history, and then at least one GPM index information stored / input into the data space in a CU to be encoded / decoded thereafter may be configured to be added to the GPM candidate list. That is, by the FIFO method, the most recently used GPM index information may be added to the GPM candidate list.
[0155] For example, referring to the example of FIG. 11, it can be seen that the angle of the diagonal boundary line (1120) in the current CU (1110) is similar to the diagonal angle (1140) previously shown in the upper left adjacent CU (1130). Therefore, when constructing the GPM candidate list of the current CU (1110), it can be configured to add the GPM partition index of the upper left adjacent CU (1130). After constructing the GPM candidate list, the candidate index actually selected based on the list can be recorded as a bit string syntax by the encoder and transmitted to the decoder. According to a preferred embodiment of the present invention, it can be understood that the GPM partition index that appears earlier in the list (i.e., has a higher priority) among the GPM candidate lists is more likely to be selected and used. Therefore, when encoding the selected candidate index in the bit string, binary encoding methods that can relatively more efficiently express small numbers can be used as a binarization method. For example, depending on the embodiment, methods including Golomb-Rice encoding, Truncated-Rice encoding, other entropy encoding methods, and other variable length encoding methods may be used.
[0156] Fig. 12 is a flowchart illustrating a process of bit string syntax processing for deriving a GPM index according to an embodiment of the present invention. Referring to Fig. 12, "gpm_enabled_flag" syntax information, which is HLS information, may be first processed (parsed) (S1210). According to an embodiment, the "gpm_enabled_flag" may be included in at least one of various elements including SPS, PPS, PH, and SH among the HLS syntax elements in the bit string. The value of the "gpm_enabled_flag" may be decoded (S1220), and the "gpm_flag" syntax information may be processed (S1230) only when the value is "on". According to an embodiment, the "gpm_flag" may be recorded in a CU unit in the bit string. By decoding the value of the above "gpm_flag" (S1240), only when the value is "on", the "gpm_partition_idx" syntax information indicating the index of the GPM form selected corresponding to the GPM candidate list may be processed (S1250). According to an embodiment, the "gpm_partition_idx" syntax information may be binarized to be recorded in a bit string using a truncation-Rice encoding method. In addition, according to an embodiment, the maximum value in the truncation-Rice encoding may be signaled and determined by a high-level bit string syntax (HLS), or may be a fixed constant that is pre-specified and shared between the encoder and the decoder. According to an embodiment of the present invention, information about a GPM partition form may be processed in a decoder through a series of processes shown in FIG. 12, and information about a prediction method and / or a prediction mode for each of a plurality of PUs may be additionally processed in the decoder thereafter.
[0157]
[0158] Example 2 (List of Optimal GPMs)
[0159] According to one embodiment of the present invention, another method for efficiently transmitting one GPM selected from among a plurality of GPMs may be provided. Similar to the first embodiment described above, the second embodiment may also be configured to use GPM partition index information of neighboring blocks (CUs). According to one embodiment of the present invention, partition indices that are highly likely to be used as GPM partition indices for the current CU may be configured as a most probable GPM (MP-GPM) list, and the remaining partition indices that do not belong to the MP-GPM list may be configured as a remaining GPM list. That is, it may be understood that the most probable GPM list is configured as a partition GPM list of a first group that is highly likely to be used by the current CU, and the remaining GPM list is configured as a partition GPM list of a second group that is less likely to be used by the current CU.
[0160] According to one embodiment of the present invention, the GPM candidate list can be constructed by citing GPM partition index information of other CUs that are spatially adjacent to the current CU. According to one embodiment of the present invention, the optimal GPM candidate list can be constructed in the following manner. First, it is assumed that the maximum number of the list can be signaled and determined by the high-level bitstream syntax (HLS), or can be a fixed number that is pre-specified and shared between the encoder and decoder, and this can be selected according to the implementation method.
[0161] The above GPM candidate list can be constructed by at least one or more of the various methods described below.
[0162] 1) GPM partition index used in a spatially adjacent block (e.g., CU). In this case, the spatially adjacent block may include at least one of the blocks adjacent to the left, top, top left, top right, and bottom left of the current block.
[0163] 2) Another GPM partition index that is adjacent to the GPM partition index of the spatially adjacent block. For example, at least one other index that is adjacent by ±a from the index extracted in 1) above may be included.
[0164] 3) GPM partition index based on past history.
[0165] 4) Basic GPM index. For example, it may include at least one GPM partition index that is commonly used by general probability.
[0166] Depending on the implementation method, at least one of the above methods may be used, or a plurality of methods may be used, and if there are multiple GPM indices derived by at least one of the above methods, an arbitrary priority may be determined among the above methods, and the indices up to the maximum number in the list may be configured to be reflected in the optimal PM candidate list based on the priority. In addition, each of the above methods may be implemented by equally applying the detailed implementation method and / or applied implementation method shown in the above-described Embodiment 1.
[0167] According to one embodiment of the present invention, a first number of optimal GPM lists may be constructed using at least one of the above methods, and a remaining GPM list may be constructed only with the remaining indexes excluding the first number of optimal GPMs from the entire GPM partition list. According to an embodiment, the first number may mean the maximum number of the optimal GPM list predetermined by various methods as described above, and may be a natural number greater than or equal to 1. According to an embodiment, all remaining indexes not included in the optimal GPM list may be sequentially shifted to construct the remaining index list. For example, if the number of indexes of the entire GPM segmentation method is 64 in total (0 to 63), the number of the optimal GPM list is set to 4, and the optimal GPM list is composed of indices 0, 4, 5, 10, the number of indexes of the remaining GPM list can be composed of 60 in total (1, 2, 3, 6, 7, 8, 9, 11, …, 63), and at this time, a new index value (0 to 59) can be assigned to the remaining GPM list. However, in the present invention, the size of the optimal GPM list and the size of the remaining GPM list are not limited, and a person skilled in the art can arbitrarily designate the sizes of the lists within the spirit and scope of the present invention and practice the present invention in the same spirit.
[0168] According to one embodiment of the present invention, after the optimal GPM list is constructed, information regarding whether the actually selected GPM index belongs to the optimal GPM list or the residual GPM list can be recorded as a bit string syntax by an encoder and transmitted to a decoder.
[0169] Fig. 13 is a flowchart illustrating a process of bit string syntax processing for deriving a GPM index according to an embodiment of the present invention. Referring to Fig. 13, "gpm_enabled_flag" syntax information, which is HLS information, may be first processed (parsed) (S1310). According to an embodiment, the "gpm_enabled_flag" may be included in at least one of various elements including SPS, PPS, PH, and SH among HLS syntax elements in the bit string. The value of the "gpm_enabled_flag" may be decoded (S1320), and only when the value is "on", the "gpm_flag" syntax information may be processed (S1330). According to an embodiment, the "gpm_flag" may be recorded in a CU unit in the bit string. By decoding the value of the above "gpm_flag" (S1340), if the value is "on", the "gpm_mpm_flag" syntax information indicating whether the GPM split index is provided by the optimal GPM list method may be processed (S1350). Next, by decoding the value of the above "gpm_mpm_flag" (S1360), if the value is "on", the "gpm_mpm_idx" syntax information indicating an index in the optimal GPM list may be processed (S1370), and if the value is "off", the "gpm_remainder" syntax information indicating an index in the remaining GPM list may be processed (S1375). According to an embodiment, the method by which the "gpm_mpm_idx" syntax information is binarized to be recorded in a bit string may use a truncation-Rice encoding method, a Golomb-Rice encoding method, another entropy encoding method, another variable length encoding method, or a fixed length encoding method, and the maximum value in any of the encoding methods may be signaled and determined by a high level bit string syntax (HLS), or may be a fixed constant that is specified in advance and shared between an encoder and a decoder.Also, according to an embodiment, the method by which the "gpm_remainder" syntax information is binarized to be recorded in a bit string may use a truncated binary encoding method, a Golomb-Rice encoding method, another entropy encoding method, another variable length encoding method, or a fixed length encoding method, and the maximum value in any one of the encoding methods may be signaled and determined by a high level bit string syntax (HLS), or may be a fixed constant that is specified in advance and shared between an encoder and a decoder.
[0170] According to one embodiment of the present invention, when the number of the optimal GPM list is defined as 1, the "gpm_mpm_flag" value indicating whether the optimal GPM list is used may be processed (S1350) and decoded (S1360), but the separate "gpm_mpm_idx" syntax information value may not be processed (S1370). That is, it may be configured so that the GPM partition information of the current CU can be derived using only the value of "gpm_mpm_flag" to which 1 bit is allocated.
[0171] According to one embodiment of the present invention, when the number of the optimal GPM list is defined as 1 and a specific condition is satisfied, the "gpm_flag" value indicating whether the current CU is GPM-partitioned may be processed (S1330) and decoded (S1340), but a separate "gpm_mpm_flag" value may not be processed (S1350) and decoded (S1360), and the "gpm_mpm_idx" value may not be processed (S1370). Here, the specific condition may include a case where a plurality of surrounding blocks (CUs) of the current CU are GPM-partitioned. In other words, it may be configured so that whether the current CU is GPM-partitioned and the partition information can be derived only from the value of "gpm_flag" to which 1 bit is allocated.
[0172] According to one embodiment of the present invention, information about a GPM division form can be processed in a decoder through a series of processes shown in FIG. 13, and information about a plurality of PU-specific prediction methods and / or prediction modes can be additionally processed in the decoder.
[0173]
[0174] Third embodiment (template-based GPM segmentation)
[0175] According to one embodiment of the present invention, another method for efficiently transmitting a selected GPM from among a plurality of GPMs may be provided. According to one embodiment, a method may be applied in which a corresponding adjacent region among the restored regions around the current block (CU) is set as a template region, and a template matching process for the template region is used to derive a GPM partition structure. According to another embodiment, a method may be applied in which the order of an existing GPM partition list is rearranged through the template matching process, and an index value within the list based on the rearranged order is transmitted.
[0176] According to one embodiment of the present invention, it may be configured to first signal a plurality of intra-screen or inter-screen prediction information (which may include block vectors for intra-screen prediction or motion vector information for inter-screen prediction) used in a GPM in a bit-string syntax, update GPM segmentation information using the prediction information and a template region corresponding to a current block (CU), and then derive and / or signal information about an actual segmentation shape used in the current CU based on the updated GPM segmentation information.
[0177] FIG. 14 is a conceptual diagram illustrating an application of inter-screen and intra-screen prediction modes based on GPM according to an embodiment of the present invention. Referring to FIG. 14, a current block (CU) (1440) corresponds to an upper template (1410) and a left template (1420), and an example is shown in which a dividing line (1430) for the current CU (1440) is extended to the area of the template (in this example, the upper template (1410)). According to an embodiment, in FIG. 14, the templates (1410, 1420) can be newly divided and processed into a template area for a left dividing area of the current CU (1440) and a template area for a right dividing area based on the dividing line (1430).
[0178] According to one embodiment of the present invention, first, by signaling a plurality of pieces of prediction information through a bit string syntax, prediction information for a template region for a left partition area and prediction information for a template region for a right partition area can be recognized in a predetermined order through the signal information. Next, template matching can be performed based on each of the prediction pieces. That is, based on information such as a block vector and / or a motion vector included in each of the prediction pieces, each reference block can be obtained, and by obtaining a template region corresponding to each of the reference blocks, an error cost of a difference value with respect to a surrounding template region of the current CU can be obtained. Through this, the process including the template matching can be repeated equally for all partition modes, so that an index of a GPM partition form with the lowest cost can be finally selected and / or derived.
[0179] According to another embodiment of the present invention, the process of obtaining a difference cost through a template matching process between the surrounding template region of each reference block based on each prediction information and the surrounding template region of the current CU is repeated for all partition modes, and the index of the GPM partition list can be rearranged in the order of the lowest cost. The encoder can be configured to select a final partition shape index based on the list rearranged by a rate-distortion optimization (RDO) process and transmit it to the decoder.
[0180] According to the above-described implementation method, the number of bits when signaled as bit string syntax information can be saved by selecting or deriving a partition shape index that is likely to be selected as a GPM partition shape for the current CU through a template matching process, or by placing such a likely partition shape index at a higher priority in a list such as a list of GPM partition shapes and / or a list of GPM candidates described above.
[0181]
[0182] Example 4 (GPM Prediction Mode)
[0183] FIG. 15 is a conceptual diagram illustrating the application of inter- and intra-screen prediction modes based on GPM according to one embodiment of the present invention. According to one embodiment of the present invention, when GPM is applied to a current block (CU), multiple (e.g., two) PUs are generated within one CU, and a method of signaling syntax information so that inter-screen prediction and / or intra-screen prediction can be freely set and used for each PU may be applied. As shown in FIG. 15, multiple PUs generated by applying GPM to one CU each undergo an independent prediction process. At this time, according to one embodiment of the present invention, it is possible to determine whether to use intra-screen prediction or inter-screen prediction for each PU. Referring to FIG. 15, when one CU is divided into two PUs by a GPM, examples (1510, 1520, 1530) in which one PU performs intra-prediction while the other performs inter-prediction, and an example (1540) in which all PUs perform intra-prediction are shown. Each of the GPM examples (1510, 1520, 1530, 1540) shown in FIG. 15 has a reference sample area (1500) around it, and is divided into at least one inter-prediction area (1505) and an intra-prediction area (1506) by an arbitrary dividing line (1501).
[0184] According to one embodiment of the present invention, the prediction method for multiple PUs may not be limited to intra- and / or inter-prediction methods. For example, various intra- and / or inter-prediction modes, including intra-block copy (IBC), merge prediction mode, and adaptive motion vector prediction (AMVP), may be applied to each PU. For example, for two PUs, general intra-prediction and IBC modes, general intra-prediction and merge-based inter-prediction, IBC and merge / AMVP-based inter-prediction, and different IBC modes may be assigned and used, respectively.
[0185]
[0186] Example 4-1 (2-bit signal)
[0187] According to one embodiment of the present invention, as described above, the syntax information indicating which prediction method each of the plurality of (for example, two) PUs used may be binarized to be recorded in a bit string, and a fixed-bit binary signaling method, for example, a 2-bit signaling method, may be used. In one embodiment, if it is assumed that there are a total of four combinations of prediction methods and that each has a substantially equal probability of being selected, the fixed-bit signaling method as described above may be advantageous. Table 3 below illustrates a case of combining four prediction modes by fixed 2-bit signaling according to one embodiment of the present invention. According to one embodiment of the present invention, the criterion for which of the plurality of PUs corresponds to the first division or the second division may be determined by a raster scan method, for example, according to the order in which they proceed from left to right and / or top to bottom.
[0188] Index 1st partition prediction mode 2nd partition prediction mode Binarized code 1 Intra Intra 0 2 Intra Inter 0 1 3 Inter Intra 1 0 4 Inter Inter 1 1
[0189]
[0190] Example 4-2 (Restricting on-screen prediction mode by PU location)
[0191] According to one embodiment of the present invention, depending on the form of GPM block division, there may be cases where intra-screen prediction is typically disadvantageous for some PUs. For example, referring to FIG. 15, in some GPM divisions (1520, 1530), the second division area (1525, 1535) located at the lower right may be disadvantageous for application of intra-screen prediction. This is due to the characteristics of intra-screen prediction, which is configured to perform prediction based on the value of a reconstructed reference sample (1500) that has already completed the decoding process around the current block (CU), for example, located at the top or left, and thus may have the characteristic that the accuracy of prediction improves as the reference sample becomes closer. However, in the case of the second partition area (1525, 1535) exemplified above, there may not be an area adjacent to the reference sample of the adjacent block, and therefore, the accuracy of the intra-screen prediction may be expected to decrease, and it may be considered as an area more suitable for inter-screen prediction, and such consideration may help improve the compression performance.
[0192] According to one embodiment of the present invention, on the premise that intra-screen prediction is not performed for PUs of GPM partitioned areas that do not have adjacent areas with the reference sample, only one of the inter-screen prediction modes is selected and performed, and as a result, a bit-string syntax information signal for using either intra-screen or inter-screen prediction may not be transmitted. That is, in a specific GPM mode, intra-screen prediction may be restricted without any other conditions for some partitioned areas, and only for the unrestricted partitioned areas may it be configured to signal which mode of intra-screen prediction or inter-screen prediction was used.
[0193]
[0194] Encoder and decoder
[0195] It is self-evident that the method according to the present invention can be applied equally to both encoders and decoders. Additionally, as illustrated in FIG. 4 through the internal decoder (420) and the coding loop including it, this decoding process can be implemented identically within the encoder to predict the state of the decoder.
[0196] The encoding method according to the present invention described above can be implemented through an encoder as a device. The encoder as a device can be implemented in a form that maintains the conventional encoder structure exemplified through FIGS. 1 to 6 above or applies a predetermined change thereto. However, the form of implementation is not necessarily limited to that exemplified, and any form of encoder structure that can function as a video encoder should be considered an encoder established by the present invention as long as it implements the technical idea of the present invention.
[0197] In addition, the method for decoding the encoding result according to the present invention described above can be implemented through a decoder as a device. The decoder as a device can be implemented in a form that maintains the conventional decoder structure exemplified through FIGS. 1 to 6 above or applies a predetermined change thereto. However, the form of implementation is not necessarily limited to that exemplified, and any form of decoder structure that can function as a video decoder should be considered a decoder established according to the present invention as long as it implements the technical idea of the present invention.
[0198] A person skilled in the art will readily understand that a bit string encoded by the above-described method and device can be decoded by applying a method symmetrical and / or reverse to the encoding method. In one embodiment, when reading information for decoding from the encoded bit string, at least one variable length coded phrase included in the encoded bit string can be interpreted, and furthermore, in one embodiment, the variable length coding can be performed by an entropy coding method. The technical details and application method of implementing such a decoding procedure can be readily understood from the above-described encoding procedure.
[0199] The processor that may be included in the encoder and / or decoder described herein may mean one or more general purpose computers or special purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding.
[0200] Even if the processor is expressed singularly for ease of understanding, those skilled in the art will appreciate that the processor may include multiple processing elements and / or multiple types of processing elements. For example, a device according to one embodiment of the present invention may include multiple processors or one processor and one controller as the processor. Furthermore, the processor may be implemented using various processing configurations, such as a parallel processor or a multi-core processor.
[0201] The processor may be configured to execute an operating system (OS) and one or more software programs running on the operating system. Furthermore, the processor may access, store, manipulate, process, and generate data in response to the execution of the software.
[0202] The software may include a computer program, code, instructions, or a combination of one or more of these, and may be configured to control the processor to perform a desired operation and to issue commands to the processor, either independently or collectively. The software may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for interpretation by the processor or for providing commands or data to the processor. The software may also be distributed over networked computer systems, and stored or executed in a distributed manner.
[0203] The software may also be implemented in the form of program commands that can be executed through various computer means and recorded or stored in the memory. The memory may be a computer-readable recording medium, and program commands, data files, data structures, etc. may be recorded singly or in combination in the computer-readable recording medium. The program commands stored in the memory may be based on a command system specifically designed and configured for the embodiment of the present invention, or may follow a command system known and available to those skilled in the art of computer software, such as a command system exemplified by the assembly language, C, C++, Java, Python, etc. It should be understood that the command system and the program commands therefrom include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by the device and / or the processor according to an embodiment of the present invention using an interpreter or the like.
[0204] The computer-readable recording medium constituting the device according to one embodiment of the present invention, including the memory described in this specification, may include a temporary or volatile recording medium that is maintained only while the processor is operating, such as a processor cache, a RAM, a flash memory, or a relatively non-volatile or long-term recording medium, such as a magnetic media such as a hard disk, a floppy disk, and a magnetic tape, an optical media such as a CD-ROM, a DVD, a magneto-optical media such as a floptical disk, or a solid state memory, or may include a read-only recording medium, such as a ROM arranged on hardware, and further, the hardware itself configured to perform an operation equivalent to a series of program commands by a hard-wired structure by circuit wiring, and each step for performing the operation for implementing the embodiment of the present invention can be viewed as being recorded by the connection and arrangement of the hardware components, and thus the connection and arrangement method is the memory and It is obvious to a person skilled in the art that they can be considered equivalent.
[0205] The embodiments described above with respect to the processor and the memory are not mutually exclusive and may be selected or combined and implemented as needed. For example, a single hardware device may be configured to operate as a module composed of one or more of the software to perform the operations of an embodiment of the present invention, and vice versa. As another example, in the present specification, all or part of the operations assigned to a certain functional unit may be implemented by one or more of the software stored in a device according to an embodiment of the present invention (preferably, in any one of the recording media belonging to the category of the memory) and configured to be executed by the processor. In such a case, such a functional unit may be referred to as a functional unit "included" in the processor.
[0206]
[0207] Although the present invention has been described with reference to drawings and embodiments, as already mentioned above, it does not mean that the scope of protection of the present invention is limited to the drawings or embodiments presented above, and it will be understood that a person skilled in the relevant technical field can modify and change the present invention in various ways within a scope that does not depart from the spirit and scope of the present invention described in the claims of the present invention patent.
Claims
1. In a video decoding method based on geometric segmentation prediction, A step of reading at least one syntax information from an encoded bit string; A step of obtaining information about the use of GPM (geometric partitioning mode) of the current block currently being decrypted; A step of obtaining GPM index information for the current block; A step of decrypting the above GPM index information to obtain GPM segmentation shape information; A step of dividing the current block into a first region and a second region based on the above GPM division form information; A step of obtaining prediction mode information for each of the first and second regions; and A decoding method, comprising: a step of executing prediction decoding for each of the first region and the second region based on the prediction mode.
2. In paragraph 1, The step of decrypting the above GPM index information to obtain GPM segmentation shape information includes a step of executing a variable length decryption function, A decoding method, characterized in that the variable length decoding function includes at least one decoding method among Golomb-Rice encoding, truncated-binary encoding, truncated-rice encoding, and entropy encoding.
3. In paragraph 1, A decryption method, characterized in that the above GPM index information is derived from a list re-sorted according to the cost of template matching.
4. In paragraph 1, A decryption method, characterized in that the GPM index information is derived from a candidate list generated based on the GPM division shapes of blocks spatially adjacent to the current block.
5. In paragraph 1, A step of obtaining size information of a most-probable GPM list; A decryption method further comprising: a step of obtaining flag information indicating that the GPM index information is for one of the optimal GPM list and a remaining GPM list.
6. In paragraph 5, The step of obtaining the above GPM division type information is: A decryption method characterized in that, when the size of the optimal GPM list is 1 and flag information indicating that the GPM index information is for the optimal GPM list is obtained, the GPM partition shape information is obtained based on the optimal GPM list without decrypting the GPM index information.
7. In paragraph 5, A decoding method, characterized in that the above-mentioned optimal GPM list includes at least one of an index of a GPM division shape of surrounding blocks spatially adjacent to the current block, and another GPM division index having a division shape similar to the above-mentioned GPM division shape.
8. In paragraph 1, A decoding method, wherein the prediction mode includes at least one of intra-screen prediction and inter-screen prediction, and wherein use of the intra-screen prediction is restricted in at least one of the first region and the second region.
9. In paragraph 8, A decoding method, characterized in that the above prediction mode includes at least one of IBC (intra block copy), merge prediction mode, and AMVP (adaptive motion vector prediction).
10. In a video encoding method based on geometric segmentation prediction, A step for determining whether to use GPM (geometric partitioning mode) of the current block currently being encoded; A step of determining GPM division type information for the current block; A step of determining GPM index information by encoding the above GPM division form information; A step of dividing the current block into a first region and a second region based on the above GPM division form information; A step of determining prediction mode information for each of the first and second regions; A step of predicting encoding the current block based on the GPM division form and the prediction mode information, and generating a residual signal between the original signal and the prediction signal of the current block; and An encoding method, comprising: a step of encoding at least one of information on whether the GPM is used, information on the GPM index, information on the prediction mode for each of the first and second regions, and information on the residual signal as syntax information in a bit string.
11. In Article 10, The step of obtaining GPM index information by encoding the above GPM segmentation form information includes a step of executing a variable length encoding function, An encoding method, characterized in that the variable length encoding function includes at least one encoding method among Golomb-Rice encoding, truncated-binary encoding, truncated-rice encoding, and entropy encoding.
12. In paragraph 10, An encoding method, characterized in that the GPM division form is determined according to the cost of template matching, and the template matching is performed using a template area including a restoration area on the top and left of the current block.
13. In paragraph 12, The above template area is divided by the above dividing line, An encoding method, characterized in that the template matching is performed using different reference template areas in the first area and the second area.
14. In paragraph 10, An encoding method, characterized in that the GPM index information is derived from a candidate list generated based on the GPM division shapes of blocks spatially adjacent to the current block.
15. In paragraph 10, further comprising a step of obtaining size information of a most-probable GPM list; The step of determining GPM index information by encoding the above GPM division form information is characterized in that the GPM index information is determined based on index information when the GPM division form belongs to either the optimal GPM list or the remaining GPM list. An encoding method, characterized in that the step of encoding the bit string as syntax information further encodes signal information indicating whether the GPM index information belongs to one of the optimal GPM list and the residual GPM list into the bit string.
16. In paragraph 15, An encoding method, characterized in that the above-mentioned optimal GPM list includes at least one of an index of a GPM division shape of surrounding blocks spatially adjacent to the current block, and another GPM division index having a division shape similar to the above-mentioned GPM division shape.
17. In paragraph 10, An encoding method, wherein the prediction mode includes at least one of intra-screen prediction and inter-screen prediction, and wherein use of the intra-screen prediction is restricted in at least one of the first region and the second region.
18. In paragraph 17, An encoding method, characterized in that the above prediction mode includes at least one of IBC (intra block copy), merge prediction mode, and AMVP (adaptive motion vector prediction).
19. In paragraph 17, An encoding method characterized by encoding a combination of the prediction mode information for each of the first region and the second region as fixed-length information in the bit string.
20. In a video decoder device based on geometric segmentation prediction, A bit string parser unit configured to read at least one piece of syntax information from an encoded bit string; A GPM information processing unit configured to obtain information on the use of GPM (geometric partitioning mode) of a current block currently being decrypted from the above syntax information, obtain GPM index information, and decode the GPM index information to generate GPM partitioning form information; A block division unit configured to divide the current block into a first region and a second region based on the GPM division form information; A prediction mode information acquisition unit configured to acquire prediction mode information for each of the first and second regions; and A decoder device, comprising: a prediction decoding unit configured to perform prediction decoding for each of the first region and the second region based on the prediction mode.
Citation Information
Patent Citations
unhardened sediment Bar rice cake
KR1020230153752A
Wireless power transmission apparatus and method for registering at least one cooking device
KR1020230160676A
Dark count calibration apparatus and method for optical sensor
KR1020250080416A
Semiconductor memory devices
KR1020250081570A
Fuel cell stack structure and fuel cell system having the fuel cell stack structure
KR102680450B1