Method, apparatus, and program for video decoding
The method enhances video coding efficiency by decoding prediction information to reconstruct sub-partitions of a current block based on adjacent samples, addressing the challenges of redundancy reduction and bandwidth optimization in existing video coding technologies.
Patent Information
- Application Number
- JP2025030474
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-02-06
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-03
AI Technical Summary
Existing video coding technologies face challenges in efficiently encoding and decoding video data, particularly in reducing redundancy and optimizing bandwidth usage, especially for high-resolution videos.
The proposed solution involves an apparatus and method for video decoding that includes receiving and processing circuits to decode prediction information from an encoded video bitstream. This information is used to determine sub-partitions of a current block and reconstruct them based on adjacent samples outside the adjacent region of the sub-partitions, allowing for improved intra prediction modes and angle parameters.
This approach enhances video encoding and decoding efficiency by improving intra prediction techniques for sub-partitions, leading to better compression ratios and reduced bandwidth requirements without significant distortion.
Smart Images

Figure 2025084869000001_ABST
Abstract
Description
Technical Field
[0001] Incorporation by Reference This application claims the benefit of priority of U.S. Patent Application No. 16 / 783,388, filed Feb. 6, 2020, entitled “METHOD AND APPARATUS FOR VIDEO CODING,” which claims the benefit of priority of U.S. Provisional Application No. 62 / 803,231, filed Feb. 8, 2019, entitled “IMPROVED INTRA PREDICTION FOR INTRA SUB-PARTITIONS CODING MODE.” The entire disclosure of the above applications is hereby incorporated by reference in its entirety.
[0002] The present disclosure generally describes embodiments related to video coding.
Background Art
[0003] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. Aspects of the work of the inventors named herein that are not described in this background section may not be eligible as prior art at the time of filing and are neither expressly nor implicitly admitted as prior art to the present disclosure.
[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each picture having, for example, a spatial dimension of 1920×1080 luminance samples and associated chrominance samples. A series of pictures can have, for example, a fixed or variable picture rate (informally also called a frame rate) of 60 pictures per second or 60 Hz. Uncompressed video has specific bitrate requirements. For example, 1080p60 4:2:0 video (1920×1080 luminance sample resolution at a 60 Hz frame rate) with 8 bits per sample requires a bandwidth close to 1.5 Gbit / s. Storing such video for one hour requires more than 600 gigabytes of storage space.
[0005] One of the purposes of video encoding and decoding is to reduce the redundancy of the input video signal by compression. Compression helps to reduce the aforementioned bandwidth and / or memory space requirements, sometimes by more than two orders of magnitude. Both reversible compression and irreversible compression, as well as combinations thereof, can be used. Reversible compression refers to a technique where an exact copy of the original signal can be reconstructed from the compressed original signal. When using irreversible compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal is useful for the intended application. In the case of video, irreversible compression is widely adopted. The amount of allowable distortion varies depending on the application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that the higher the allowable / tolerable distortion, the higher the compression ratio can be.
[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy encoding.
[0007] Video codec technology can include a technique called intra coding. In intra coding, sample values are represented without reference to samples from previously reconstructed reference pictures or other data. In some video codecs, a picture is spatially divided into blocks of samples. If all blocks of samples are encoded in an intra mode, that picture can be an intra picture. Derivatives such as intra pictures and independent decoder refresh pictures can be used to reset the state of the decoder and thus can be used as the first picture of an encoded video bitstream and video session or as a still image. Samples of an intra block can be transformed, and the transform coefficients can be quantized before entropy coding. Intra prediction is a technique that minimizes sample values in the domain before transformation. In some cases, the smaller the DC value and AC coefficients after transformation, the fewer bits are required at a specific quantization step size to represent the block after entropy coding.
[0008] For example, conventional intra coding, such as that known from MPEG-2 generation coding technology, does not use intra prediction. However, some new video compression technologies include techniques that attempt, for example, from surrounding sample data and / or metadata obtained during the encoding / decoding of blocks of data that are spatially adjacent and precede in decoding order. Such techniques are hereinafter referred to as "intra prediction" techniques. It should be noted that in at least some cases, intra prediction uses only reference data from the current picture being reconstructed and not from reference pictures.
[0009] Intra prediction can have various forms. If two or more of such techniques can be used in a given video coding technique, the techniques used can be encoded in an intra prediction mode. In some cases, the mode can have sub - modes and / or parameters, which can be encoded individually or included in the mode codeword. The codeword used for a particular combination of mode / sub - mode / parameters can affect the improvement of encoding efficiency by intra prediction and also affects the entropy coding technique used to convert the codeword into a bitstream.
[0010] A particular mode of intra prediction was introduced in H.264, improved in H.265, and further improved in new coding techniques such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). The predictor block can be formed using adjacent sample values belonging to already available samples. The sample values of the adjacent samples are copied into the predictor block according to a direction. The reference to the direction at the time of use can be encoded in the bitstream or predicted by itself.
[0011] Referring to FIG. 1, shown at the lower right is a subset of 9 predictor directions known from 33 possible predictor directions in H.265 (corresponding to 33 angular modes of 35 intra modes). The point (101) where the arrows converge represents the sample to be predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples in the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples in the lower left of sample (101) at an angle of 22.5 degrees from the horizontal.
[0012] Continuing to refer to FIG. 1, a square block (104) of 4×4 samples is depicted in the upper left (indicated by the thick dashed line). The square block (104) includes 16 samples each labeled with an "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample of block (104) in both the Y and X dimensions. Since the block size is 4x4 samples, S44 is in the lower right. Additional reference samples according to a similar numbering scheme are further shown. The reference samples are labeled with an R for block (104), its Y position (e.g., row index), and X position (column index). In both H.264 and H.265, the predicted samples are adjacent to the block being reconstructed, and thus there is no need to use negative values.
[0013] Intra-picture prediction can function by copying reference sample values from adjacent samples according to the prediction direction of the signal. For example, assume that the encoded video bitstream includes signaling indicating a prediction direction that matches arrow (102) for this block. That is, the samples are predicted from one or more reference samples at a 45-degree angle from horizontal going up and to the right. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Next, sample S44 is predicted from reference sample R08.
[0014] In certain cases, in order to calculate the reference samples, especially when the direction is not evenly divisible by 45 degrees, the values of multiple reference samples can be combined, for example, by interpolation.
[0015] As video coding technology has evolved, the number of possible directions has increased. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions at the time of disclosure. Experiments are conducted to identify the most likely directions, and specific techniques of entropy coding are used to represent those likely directions with a few bits, accepting a specific penalty for less likely directions. Further, the direction itself may be predictable from adjacent directions used in adjacent, already decoded blocks.
[0016] FIG. 2 shows a schematic diagram (201) showing 65 intra prediction directions by JEM to illustrate the number of prediction directions increasing over time.
[0017] The mapping of intra prediction direction bits within an encoded video bitstream representing a direction can vary for each video coding technology and can range from a simple direct mapping to complex adaptive schemes including, for example, an intra prediction mode, a codeword, the most likely mode, from a prediction direction, and similar techniques. However, in all cases, there can be specific directions that are statistically less likely to occur in video content than certain other directions. Since the goal of video compression is redundancy reduction, these less likely directions are represented with more bits than the more likely directions in a properly functioning video coding technology. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0018] Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some examples, an apparatus for video decoding includes a receiving circuit and a processing circuit. The processing circuit decodes prediction information of a current block from an encoded video bitstream. In some embodiments, the processing circuit determines a first sub-partition and a second sub-partition of the current block based on the decoded prediction information, and then reconstructs the first sub-partition and the second sub-partition of the current block based on at least adjacent samples of the current block that are outside an adjacent region of at least one of the first sub-partition and the second sub-partition.
[0019] In one embodiment, the current block is vertically divided into at least a first sub-partition and a second sub-partition, the adjacent samples are the adjacent samples at the lower left of the current block, and are outside an adjacent region of at least one of the first sub-partition and the second sub-partition. In another embodiment, the current block is horizontally divided into at least a first sub-partition and a second sub-partition, the adjacent samples are the adjacent samples at the upper right of the current block, and are outside an adjacent region of at least one of the first sub-partition and the second sub-partition.
[0020] In some embodiments, the prediction information indicates a first intra prediction mode in a first set of intra prediction modes for square-shaped intra prediction. The processing circuit remaps the first intra prediction mode to a second intra prediction mode in a second set of intra prediction modes for non-square-shaped intra prediction based on the shape of the first sub-partition of the current block, and reconstructs samples of at least the first sub-partition according to the second intra prediction mode.
[0021] In one embodiment, when the aspect ratio of the first sub - partition is outside a specific range, the number of subsets of intra - prediction modes in the first set of intra - prediction modes that are replaced by the wide - angle intra - prediction mode in the second set of intra - prediction modes is fixed to a predefined number. In an example, the predefined number is one of 15 and 16.
[0022] In another embodiment, the processing circuit determines angle parameters related to the second intra - prediction mode based on a look - up table having an accuracy of 1 / 64.
[0023] In one example, when the aspect ratio of the first sub - partition is equal to 16 or 1 / 16, the number of intra - prediction modes replaced from the first set to the second set is 13. In another example, when the aspect ratio of the first sub - partition is equal to 32 or 1 / 32, the number of intra - prediction modes replaced from the first set to the second set is 14. In another example, when the aspect ratio of the first sub - partition is equal to 64 or 1 / 64, the number of intra - prediction modes replaced from the first set to the second set is 15.
[0024] In some embodiments, the processing circuit determines a partition direction based on the size information of the current block and decodes a signal from an encoded bitstream indicating several sub - partitions for partitioning the current block in the partition direction.
[0025] Aspects of the present disclosure also provide a non - transitory computer - readable medium that stores instructions which, when executed by a computer for video decoding, cause the computer to execute a method for video decoding.
[0026] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
Brief Description of the Drawings
[0027]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10A
Figure 10B
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Modes for Carrying Out the Invention
[0028] Figure 3 shows a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes, for example, a plurality of terminal devices that can communicate with each other via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) can encode video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to recover the video pictures, and display the video pictures according to the recovered video data. Unidirectional data transmission can be common in media serving applications and the like.
[0029] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data that may occur, for example, during a video conference. In bidirectional data transmission, in one example, each of the terminal devices (330) and (340) can encode video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other terminal device of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other terminal device of the terminal devices (330) and (340), decode the encoded video data to recover the video pictures, and display the video pictures on an accessible display device according to the recovered video data.
[0030] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) can be shown as servers, personal computers, and smartphones, but the principles of the present disclosure may not be so limited. Embodiments of the present disclosure find applications using laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (350) represents any number of networks that transfer encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (wiring) and / or wireless communication networks. The communication network (350) can exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of the network (350) may not be important for the operation of the present disclosure, unless otherwise described herein below.
[0031] FIG. 4 shows the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter can be equally applied to other video-enabled applications, including, for example, video conferencing, digital television, storage of compressed video on digital media such as CDs, DVDs, memory sticks, etc.
[0032] The streaming system can include a capture subsystem (413) that can include a video source (401), such as a digital camera, and creates, for example, a stream (402) of uncompressed video pictures. In one example, the stream (402) of video pictures includes samples taken by a digital camera. The stream (402) of video pictures, drawn with a thick line to emphasize the high data volume compared to the encoded video data (404) (or encoded video bitstream), can be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) can include hardware, software, or a combination thereof, and enables or implements aspects of the disclosed subject matter as will be described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), drawn as a thin line to emphasize the lower data volume compared to the stream (402) of video pictures, can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes an incoming copy (407) of the encoded video data and creates an outgoing stream of video pictures (411) that can be rendered on a display (412) (such as a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (such as a video bitstream) can be encoded according to a particular video encoding / compression standard specification. Examples of these standard specifications include ITU-T Recommendation H.265.In one example, the video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.
[0033] Note that electronic devices (420) and (430) can include other components (not shown). For example, electronic device (420) can include a video decoder (not shown), and electronic device (430) can also include a video encoder (not shown).
[0034] FIG. 5 shows a block diagram of a video decoder (510) according to an embodiment of the present disclosure. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (410) of the example of FIG. 4.
[0035] The receiver (531) can receive one or more encoded video sequences decoded by the video decoder (510), and in the same or another embodiment, is one encoded video sequence at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence can be received from the channel (501), and the channel (501) may be a hardware / software link to a storage device that stores the encoded video data. The receiver (531) can receive the encoded video data together with other data, such as encoded audio data and / or auxiliary data streams, which can be transferred to their respective using entities (not shown). The receiver (531) can separate the encoded video sequence from other data. To counter network jitter, the buffer memory (515) can be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter, "parser (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In others, it may be outside the video decoder (510) (not shown). In still others, outside the video decoder (510), there is, for example, a buffer memory (not shown) to counter network jitter, and inside the video decoder (510), there may be another buffer memory (515), for example, to accommodate the timing of playback. When the receiver (531) is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from an isochronous network, the buffer memory (515) may not be necessary or may be small. For use in a best-effort packet network such as the Internet, the buffer memory (515) may be required, can be relatively large, and advantageously can be of an adaptive size and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder (510).
[0036] Video decoder (510) can include a parser (520) for reconstructing symbols (521) from an encoded video sequence. The categories of those symbols include information used to manage the operation of video decoder (510), and potentially, information for controlling a rendering device (such as a display screen) (512) that is not an essential part of electronic device (530) but can be coupled to electronic device (530), as shown in FIG. 5. The control information for the rendering device may be in the form of supplementary enhancement information (SEI message) or a video user-ability information (VUI) parameter set fragment (not shown). Parser (520) can syntax analyze / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can follow video encoding techniques or standards, and can follow various principles including variable length encoding, Huffman encoding, arithmetic encoding with or without context sensitivity, etc. Parser (520) can extract a set of at least one subgroup parameter of a subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The subgroup can include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. Parser (520) can also extract from the encoded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.
[0037] Parser (520) can perform an entropy decoding / syntax analysis operation on the video sequence received from buffer memory (515) to create symbols (521).
[0038] The reconstruction of symbol (521) can include a plurality of different units depending on the type of the encoded video picture or a part thereof (e.g., inter and intra pictures, inter and intra blocks), and other factors. How each unit is involved can be controlled by subgroup control information parsed from the encoded video sequence by parser (520). Such a flow of subgroup control information between parser (520) and the plurality of following units is not shown for clarity.
[0039] Beyond the function blocks already described, video decoder (510) can be conceptually subdivided into several functional units as described below. In actual implementations operating under commercial constraints, many of these units interact closely with each other and can at least partially be integrated with each other. However, for the purpose of explaining the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate.
[0040] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives, as symbol (521) from parser (520), the quantized transform coefficients, and control information including the transform, block size, quantization coefficients, quantization scaling matrix, etc. to be used. The scaler / inverse transform unit (551) can output a block including sample values that can be input to aggregator (555).
[0041] In some cases, the output samples of the scaler / inverse transform (551) can be related to intra-coded blocks. That is, blocks that do not use prediction information from previously reconstructed pictures but can use prediction information from previously reconstructed parts of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses the surrounding already reconstructed information fetched from the current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, the partially reconstructed current picture and / or the fully reconstructed current picture. The aggregator (555) adds, in some cases, for each sample, the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).
[0042] In other cases, the output samples of the scaler / inverse transform unit (551) can be related to inter-coded and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (553) can access the reference picture memory (557) to fetch the samples used for prediction. After motion-compensating the samples fetched according to the symbols (521) related to the block, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case, called residual samples or residual signal or residual information) to generate output sample information. The address in the reference picture memory (557) where the motion compensation prediction unit (553) fetches the prediction samples can be controlled by, for example, the motion vectors available to the motion compensation prediction unit (553) in the form of symbols (521) that can have X, Y, and reference picture components. Motion compensation can also include interpolation of sample values fetched from the reference picture memory (557) when exact sub-sample motion vectors are used, and motion vector prediction mechanisms, etc.
[0043] The output samples of the aggregator (555) can be subject to various loop filtering techniques in the loop filter unit (556). The video compression technology is controlled by the parameters included in the encoded video sequence (also called the encoded video bitstream), and can include in-loop filter techniques made available to the loop filter unit (556) as symbols (521) from the parser (520), but can also respond to meta-information obtained during the decoding of the previous (in decoding order) part of the encoded picture or encoded video sequence, and previously reconstructed and loop-filtered sample values.
[0044] The output of the loop filter unit (556) can be not only output to the rendering device (512), but also a sample stream that can be stored in the reference picture memory (557) for use in future inter-picture prediction.
[0045] When a particular encoded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, when the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by a parser (520)), the current picture buffer (558) becomes part of the reference picture memory (557), and a new current picture buffer can be reallocated before starting the reconstruction of the next encoded picture.
[0046] The video decoder (510) can perform a decoding operation according to a predetermined video compression technology in a standard such as ITU-T Rec.H.265. The encoded video sequence can conform to the syntax specified by the video compression technology or standard being used, in the sense that the encoded video sequence conforms to both the syntax of the video compression technology or standard and the profile documented in the video compression technology or standard. Specifically, the profile can select specific tools from all the tools available in the video compression technology or standard as the only tools that can be used in that profile. Also, for compliance, it is necessary that the complexity of the encoded video sequence is within the range defined by the level of the video compression technology or standard. In some cases, the level may limit the maximum picture size, maximum frame rate, maximum reconstructed sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limitations set by the level may, in some cases, be further restricted by the specification of the hypothetical reference decoder (HRD) and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0047] In one embodiment, the receiver (531) can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data can be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can take the form of, for example, a temporal, spatial, or signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, and the like.
[0048] FIG. 6 shows a block diagram of a video encoder (603) according to an embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) of the example of FIG. 4.
[0049] The video encoder (603) can receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that can capture video images to be encoded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).
[0050] The video source (601) can provide a source video sequence to be encoded by a video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. Video data can be provided as a plurality of individual pictures that give motion when viewed in sequence. Each picture itself can be organized as a spatial array of pixels, and each pixel can contain one or more samples depending on the sampling structure, color space, etc. in use. A person skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0051] According to one embodiment, the video encoder (603) can encode pictures of the source video sequence in real time or under any other arbitrary time constraints required by the application and compress them into an encoded video sequence (643). Enforcing an appropriate encoding rate is one function of the controller (650). In some embodiments, the controller (650) controls other functional units and is functionally coupled to other functional units as described below. For clarity, the couplings are not shown. Parameters set by the controller (650) can include rate control related parameters (such as picture skip, quantizer, lambda value of rate distortion optimization techniques), picture size, picture group (GOP) layout, and maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.
[0052] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As an overly simplified explanation, in one example, the encoding loop can include a source coder (630) (e.g., responsible for creating symbols such as a symbol stream based on an input picture to be encoded and a reference picture), and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs symbols to create sample data in the same way as a (remote) decoder also does (since the compression between the symbols and the encoded video bitstream is lossless in the video compression techniques contemplated by the disclosed subject matter). The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream yields bit-exact results regardless of the location of the decoder (local or remote), the contents of the reference picture memory (634) are also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as the decoder "sees" as reference picture samples when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, e.g., due to channel errors) is also used in some related arts.
[0053] The operation of the "local" decoder (633) can be the same as that of a "remote" decoder such as the video decoder (510) already described in detail above in connection with FIG. 5. However, referring briefly to FIG. 5 as well, since symbols are available and the encoding / decoding of symbols into the encoded video sequence by the entropy encoder (645) and the parser (520) can be lossless, the entropy decoding part of the video decoder (510), including the buffer memory (515) and the parser (520), may not be fully implemented in the local decoder (633).
[0054] The observations that can be made at this point are that decoder technologies other than the syntax analysis / entropy decoding present in the decoder must necessarily exist in substantially the same functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder technologies can be omitted since it is the reverse of the decoder technologies described comprehensively. More detailed descriptions are necessary only in certain areas and are provided below.
[0055] During operation, in some examples, the source coder (630) can perform motion-compensated predictive coding that predictively codes an input picture by referring to one or more previously encoded pictures from a video sequence designated as a "reference picture". In this way, the encoding engine (632) encodes the difference between a pixel block of the input picture and a pixel block of a reference picture that can be selected as a predictive reference to the input picture.
[0056] The local video decoder (633) can decode the encoded video data of a picture that can be designated as a reference picture based on the symbols created by the source coder (630). The operation of the encoding engine (632) may advantageously be a lossy process. If the encoded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (633) can replicate the decoding process that can be performed by the video decoder for the reference picture and store the reconstructed reference picture in the reference picture cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference picture having common content as the reconstructed reference picture obtained by the remote video decoder (in the absence of transmission errors).
[0057] The predictor (635) can perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) can search the reference picture memory (634) for sample data (which can serve as an appropriate prediction criterion for the new picture, as a candidate reference pixel block), or for specific metadata such as the motion vectors and block shapes of the reference pictures. The predictor (635) can operate on a per-sample block and pixel block basis to find an appropriate prediction reference. Optionally, the input picture can have a prediction reference drawn from a plurality of reference pictures stored in the reference picture memory (634), as determined by the search results obtained by the predictor (635).
[0058] The controller (650) can manage the encoding operations of the source coder (630), including, for example, setting the parameters and subgroup parameters used to encode video data.
[0059] The outputs of all of the aforementioned functional units can undergo entropy encoding by the entropy coder (645). The entropy coder (645) converts the symbols generated by the various functional units into an encoded video sequence by reversibly compressing the symbols according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc.
[0060] The transmitter (640) can buffer the encoded video sequence created by the entropy coder (645) and prepare for transmission via a communication channel (660), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) can merge the encoded video data from the video coder (603) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0061] The controller (650) can manage the operation of the video encoder (603). During encoding, the controller (650) can assign a specific encoded picture type to each encoded picture, which can affect the encoding techniques applicable to each picture. For example, a picture is often assigned as one of the following picture types.
[0062] An intra picture (I picture) can be encoded and decoded without using other pictures in the sequence as a source of prediction. Some video codecs consider various types of intra pictures, such as independent decoder refresh (“IDR”) pictures, for example. Those skilled in the art know those variations of I pictures and their respective uses and characteristics.
[0063] A predicted picture (P picture) can be encoded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.
[0064] A bidirectional predicted picture (B picture) can be encoded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use metadata associated with three or more reference pictures for the reconstruction of a single block.
[0065] The source picture can typically be spatially divided into a plurality of sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and encoded block by block. The blocks can be encoded predictively by referring to other (already encoded) blocks as determined by the encoding assignment applied to each picture of the block. For example, blocks of an I picture can be encoded non-predictively, or they can be encoded predictively by referring to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be encoded predictively via spatial prediction or via temporal prediction that refers to one previously encoded reference picture. Blocks of a B picture can be encoded predictively via spatial prediction or via temporal prediction that refers to one or two previously encoded reference pictures.
[0066] The video encoder (603) can perform an encoding operation according to a predetermined video encoding technique or standard specification such as ITU-T Rec.H.265. In that operation, the video encoder (603) can perform various compression operations including a predictive encoding operation that utilizes the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technique or standard specification being used.
[0067] In one embodiment, the transmitter (640) can transmit additional data along with the encoded video. The source coder (630) can include such data as part of the encoded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.
[0068] Video can be captured as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation of a particular picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If a block of the current picture is similar to a reference block of a reference picture of the video that has been previously encoded and is still being buffered, the block of the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension that identifies the reference picture if multiple reference pictures are being used.
[0069] In some embodiments, bidirectional prediction techniques can be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures, such as a first reference picture and a second reference picture, both of which are earlier in the decoding order than the current picture in the video (however, they can be in the past and future respectively in the display order), are used. A block within the current picture can be encoded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.
[0070] Furthermore, merge mode techniques can be used for inter-picture prediction to improve the encoding efficiency.
[0071] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard specification, pictures of a series of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs of a picture are of the same size (such as 64x64 pixels, 32x32 pixels, 16x16 pixels, etc.). Generally, a CTU contains three coding tree blocks (CTBs). These are one luminance CTB and two chrominance CTBs. Each CTU can be recursively quad-tree divided into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one 64x64 pixel CU, four 32x32 pixel CUs, or sixteen 16x16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of the CU (e.g., inter-prediction type or intra-prediction type). A CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Generally, each PU contains a luminance prediction block (PB) and two chrominance PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luminance prediction block as an example of a prediction block, the prediction block contains a matrix of values (e.g., luminance values) for pixels such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.
[0072] FIG. 7 shows a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values in a current video picture within a series of video pictures and encode the processing block into an encoded picture that is part of an encoded video sequence. In one example, the video encoder (703) is used instead of the video encoder (403) of the example in FIG. 4.
[0073] In an example of HEVC, a video encoder (703) receives a matrix of sample values of a processing block, such as a prediction block of 8x8 samples. The video encoder (703) determines whether the processing block is optimally encoded using, for example, rate distortion optimization, in an intra mode, an inter mode, or a bi - directional prediction mode. If the processing block is encoded in the intra mode, the video encoder (703) can encode the processing block into an encoded picture using intra prediction techniques, and if the processing block is encoded in the inter mode or the bi - directional prediction mode, the video encoder (703) can encode the processing block into an encoded picture using inter prediction or bi - directional prediction techniques, respectively. In certain video encoding techniques, the merge mode is an inter - picture prediction sub - mode in which the motion vector is derived from one or more motion vector predictors without the benefit of the encoded motion vector components outside the predictor. In certain other video encoding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.
[0074] In the example of FIG. 7, the video encoder (703) includes an inter - encoder (730), an intra - encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general - purpose controller (721), and an entropy encoder (725) that are coupled together as shown in FIG. 7.
[0075] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks (e.g., blocks of a previous picture and a subsequent picture) in a reference picture, generate inter-prediction information (e.g., a description of redundant information by an inter-coding technique, a motion vector, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on the encoded video information.
[0076] The intra-encoder (722) is configured to receive samples of a current block (e.g., a processing block), in some cases compare the block with blocks already encoded in the same picture, generate quantized coefficients after transformation, and also generate, in some cases, intra-prediction information (e.g., intra-prediction direction information by one or more intra-coding techniques). In one example, the intra-encoder (722) also calculates an intra-prediction result (e.g., a predicted block) based on intra-prediction information and reference blocks within the same picture.
[0077] The general-purpose controller (721) is configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In one example, the general-purpose controller (721) determines the mode of a block and provides a control signal to the switch (726) based on the mode. For example, when the mode is the intra mode, the general-purpose controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723), controls the entropy encoder (725) to select the intra prediction information, includes the intra prediction information in the bitstream, and when the mode is the inter mode, the general-purpose controller (721) controls the switch (726) to select the inter prediction result used by the residual calculator (723), controls the entropy encoder (725) to select the inter prediction information, and includes the inter prediction information in the bitstream.
[0078] The residual calculator (723) is configured to calculate the difference (residual data or residual information) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) operates based on the residual data and is configured to encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate transform coefficients. Next, the transform coefficients are subject to quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture is buffered in a memory circuit (not shown) and can be used as a reference picture in some examples.
[0079] The entropy encoder (725) is configured to format the bitstream to include the encoded block. The entropy encoder (725) is configured to include various information according to an appropriate standard such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. Note that according to the disclosed subject matter, there is no residual information when encoding a block in either the merge submode of the inter mode or the bidirectional prediction mode.
[0080] FIG. 8 shows a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive an encoded picture that is part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In one example, the video decoder (810) is used in place of the video decoder (410) of the example of FIG. 4.
[0081] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) that are coupled together as shown in FIG. 8.
[0082] The entropy decoder (871) can be configured to reconstruct from the encoded picture specific symbols that represent syntax elements that make up the encoded picture. Such symbols can include, for example, the mode in which a block is encoded (e.g., intra mode, inter mode, bi - directional prediction mode, the latter two being merge sub - modes or another sub - mode, etc.), prediction information (e.g., intra prediction information or inter prediction information, etc.) that can identify specific samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880) respectively, and residual information in the form of, for example, quantized transform coefficients. In one example, when the prediction mode is an inter prediction mode or a bi - directional prediction mode, the inter prediction information is provided to the inter decoder (880), and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information can be subject to inverse quantization and is provided to the residual decoder (873).
[0083] The inter decoder (880) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.
[0084] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0085] The residual decoder (873) is configured to perform inverse quantization to extract non-quantized transform coefficients, and process the non-quantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require specific control information (including quantization parameter (QP)), and such information may be provided by the entropy decoder (871) (since this may be only a small amount of control information, the data path is not shown).
[0086] The reconstruction module (874) is configured to combine, in the spatial domain, the residual output by the residual decoder (873) and the prediction result (optionally output by the inter prediction module or the intra prediction module) to form a reconstructed block, which may be part of a reconstructed picture and thus part of a reconstructed video. Note that other appropriate operations, such as deblocking operations, can be performed to improve visual quality.
[0087] Note that the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using any suitable technology. In one embodiment, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603), and the video decoders (410), (510), and (810) can be implemented using one or more processors that execute software instructions.
[0088] Aspects of the present disclosure provide intra prediction improvement techniques for intra sub-partition (ISP).
[0089] Generally, a picture is partitioned into a plurality of blocks for encoding and decoding. In some examples, according to the HEVC standard, a picture can be divided into a plurality of coding tree units (CTUs). Further, a CTU can be partitioned into coding units (CUs) by using a quad tree (QT) structure shown as a coding tree in order to adapt to various local characteristics of the picture. Each CU corresponds to a leaf of the quad tree structure. The decision of whether to encode a picture area using, for example, inter-picture prediction (also called temporal prediction or inter prediction type), intra-picture prediction (also called spatial prediction or intra prediction type), etc. is made at the CU level. Each CU can be further partitioned into one, two, or four prediction units (PUs) according to a PU partition type having a root at the CU level. For each PU, the same prediction process is applied, and the relevant prediction information is sent to the decoder on a PU basis. After applying a prediction process based on the PU partition type to obtain prediction residual data (also called residual information), the CU can be partitioned into transform units (TUs) according to another QT structure having a root at the CU level. According to the HEVC standard, in some examples, the concept of partition includes CUs, PUs, and TUs. In some examples, a CU, a PU related to the CU, and a TU related to the CU may have different block sizes. Further, a CU or a TU needs to have a square shape in the QT structure, but a PU can have a square shape or a rectangular shape.
[0090] Note that in some embodiments, encoding / decoding is performed on blocks. For example, coding tree blocks (CTBs), coding blocks (CBs), prediction blocks (PBs), and transform blocks (TBs) can be used to specify, for example, a 2D sample array of one color component related to the corresponding CTU, CU, PU, and TU, respectively. For example, a CTU can include one luma CTB and two chroma CTBs, and a CU can include one luma CB and two chroma CBs.
[0091] In order to outperform HEVC in compression ability, many other partitioning structures have been proposed for the next-generation video coding standard beyond HEVC, namely, the so-called Versatile Video Coding (VVC). One of these proposed partitioning structures is called the QTBT structure that uses both a Quad Tree (QT) and a Binary Tree (BT). Compared with the QT structure of HEVC, the QTBT structure of VVC removes the separation between the concepts of CU, PU, and TU. In other words, the CU, the PU associated with the CU, and the TU associated with the CU can have the same block size in the QTBT structure of VVC. Furthermore, the QTBT structure supports more flexibility in the CU partition shape. The CU can have a square shape or a rectangular shape in the QTBT structure.
[0092] In some examples, the QTBT partitioning scheme defines certain parameters such as CTU size, MinQTSize, MaxBTsize, MaxBTDepth, MinBTSize. The CTU size is the root node size of the quad tree, which is the same concept as in HEVC. MinQTSize is the minimum allowable quad tree leaf node size. MaxBTSize is the maximum allowable binary tree root node size. MaxBTDepth is the maximum allowable binary tree depth. MinBTSize is the minimum allowable binary tree leaf node size.
[0093] In an example of the QTBT partitioning structure, the CTU size is set as 128×128 luma samples having chroma samples of two corresponding 64×64 blocks, MinQTSize is set as 16×16, MaxBTSize is set as 64×64, MinBTSize (both width and height) is set as 4×4, and MaxBTDepth is set as 4. Quad-tree partitioning is first applied to the CTU, and quad-tree leaf nodes are generated. The quad-tree leaf nodes can have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). When the leaf quad-tree node is 128×128, since the size exceeds MaxBTSize (i.e., 64×64), it is not further divided by the binary tree. Otherwise, the leaf quad-tree node can be further partitioned by the binary tree. Therefore, the quad-tree leaf node is also the root node of the binary tree, and the depth of the binary tree is 0. When the depth of the binary tree reaches MaxBTDepth (i.e., 4), further division is not considered. When the width of the binary tree node is equal to MinBTSize (i.e., 4), further horizontal division is not considered. Similarly, when the height of the binary tree node is equal to MinBTSize, further vertical division is not considered. The leaf nodes of the binary tree are further processed by prediction and transformation processing without being further partitioned. In JEM, the maximum CTU size is 256×256 luma samples.
[0094] FIG. 9 shows an example of a block (901) partitioned using a QTBT structure (900). In the example of FIG. 9, the QT partitions are represented by solid lines and the BT partitions are represented by dashed lines. Specifically, the QTBT structure (900) includes nodes corresponding to various blocks during partitioning within the block (901). If a node within the QTBT structure (900) has branches, that node is called a non-leaf node, and if a node within the QTBT structure (900) has no branches, that node is called a leaf node. Non-leaf nodes correspond to intermediate blocks that are further divided, and leaf nodes correspond to final blocks that are not further divided. If there are four branches in a non-leaf node, the QT partition divides the block corresponding to the non-leaf node into four small blocks of the same size. If there are two branches in a non-leaf node, the BT split divides the block corresponding to the non-leaf node into two small blocks of the same size. There are two types of BT splits: symmetric horizontal split and symmetric vertical split. In some examples, for each BT split node other than a leaf, a flag is signaled (e.g., in an encoded video bitstream) to indicate the split type (e.g., '0' for a symmetric horizontal split or '1' for a symmetric vertical split, generating two small rectangular blocks of the same size). In the case of a QT split node of the QTBT structure (900), since the QT partition divides the block corresponding to the QT split node both horizontally and vertically to generate four small blocks of the same size, there is no need to specify a split type for the QT partition.
[0095] Specifically, in the example of FIG. 9, the QTBT structure (900) includes a root node (950) corresponding to the block (901). The root node (950) has four branches that each generate a node (961)-(964), and thus the block (901) is divided by a QT into four blocks (911)-(914) of equal size. The nodes (961)-(964) each correspond to one of the four blocks (911)-(914).
[0096] Furthermore, node (961) has two branches that generate nodes (971) and (972), node (962) has two branches that generate nodes (973) and (974), and node (963) has four branches that generate nodes (975) to (978). Note that node (964) has no branches and is thus a leaf node. Also note that node (962) is a non-leaf BT split node with a split type of "1", and node (963) is a non-leaf BT split node with a split type of "0". Therefore, block (911) is split by a BT that is vertically split into two blocks (921) to (922), block (912) is split by a BT that is horizontally split into two blocks (923) to (924), and block (913) is split by a QT that is split into four blocks (925) to (928).
[0097] Similarly, node (971) has two branches that generate nodes (981) to (982), and node (975) has two branches that generate nodes (983 to (984) with a split type flag indicating a vertical split (e.g., split type "1"). Similarly, node (978) has two branches that generate nodes (985) to (986) and has a split type flag indicating a horizontal split (e.g., split type "0"). Node (984) has two branches that generate nodes (991) to (992) and has a split type flag indicating a horizontal split (e.g., split type "0"). Therefore, the right halves of the corresponding blocks (921), (928), and (925) are split into smaller blocks as shown in Figure 9. Next, nodes (981), (982), (972), (973), (974), (983), (991), (992), (976), (977), (985), and (986) are similar to node (964) which has no branches and are thus leaf nodes.
[0098] In the example of FIG. 9, leaf nodes (964), (981), (982), (972), (973), (974), (983), (991), (992), (976), (977), (985), and (986) are not further divided and are CUs used for prediction and conversion processing. In some examples such as VVC, a CU is used as a PU and a TU, respectively.
[0099] Furthermore, the QTBT scheme supports the flexibility that the luminance and chrominance have separate QTBT structures. In some examples, for P slices and B slices, the luminance CTB and chrominance CTB within one CTU share the same QTBT structure. However, for I slices, the luminance CTB is partitioned into CUs by the QTBT structure, and the chrominance CTB is partitioned into chrominance CUs by another QTBT structure. This means that in some examples, the CU of an I slice is composed of an encoding block of the luminance component or encoding blocks of two chrominance components, and the CU of a P or B slice is composed of encoding blocks of all three color components.
[0100] In some examples such as HEVC, in order to reduce the memory access of motion compensation, the inter prediction of small blocks is restricted. Therefore, bi - directional prediction is not supported for 4×8 and 8×4 blocks, and inter prediction is not supported for 4×4 blocks. In some examples such as QTBT implemented in JEM - 7.0, these restrictions are removed.
[0101] In addition to the above QTBT structure, another partitioning structure called the multi-type tree (MTT) structure is also used in VVC and can be more flexible than the QTBT structure. In the MTT structure, as shown in FIGS. 10A and 10B, in addition to the quad tree and binary tree, horizontal and vertical center-side triple tree (TT) partitions are introduced. The triple tree partition is also called triple tree partitioning, ternary tree (TT) partition, or ternary partition. FIG. 10A shows an example of a vertical center-side TT partition. For example, block (1010) is divided into three sub-blocks (1011) to (1013) vertically, and sub-block (1012) is located in the center of block (1010). FIG. 10B shows an example of a horizontal center-side TT partition. For example, block (1020) is divided into three sub-blocks (1021) to (1023) horizontally, and sub-block (1022) is located in the center of block (1020). Similar to the BT partition, in the TT partition, for example, a flag can be notified in the video bit stream from the encoder side to the decoder side to indicate the partition type (i.e., symmetric horizontal partition or symmetric vertical partition). In one example, "0" indicates a symmetric horizontal partition and "1" indicates a symmetric vertical partition.
[0102] According to one aspect of the present disclosure, the triple tree partition can complement the quad tree and binary tree partitions. For example, the triple tree partition can capture an object at the center of a block, while the quad tree and binary tree are always divided along the center of the block. Further, since the width and height of the proposed triple tree partition are powers of two, no additional conversion is required.
[0103] Theoretically, the complexity of tree traversal is TD, where T indicates the number of partition types and D indicates the depth of the tree. Thus, in some examples, a two-level tree is used to reduce the complexity.
[0104] FIG. 11 shows a diagram of exemplary intra prediction directions and intra prediction modes used in HEVC. There are a total of 35 intra prediction modes (modes 0 to 34) in HEVC. Modes 0 and 1 are non-directional modes, among which mode 0 is the planar mode and mode 1 is the DC mode. The DC mode uses the average of all samples. The planar mode uses the average of two linear predictions. Modes 2 to 34 are directional modes, among which mode 10 is the horizontal mode, mode 26 is the vertical mode, and modes 2, 18, and 34 are diagonal modes. In some examples, the intra prediction mode is signaled by three most probable modes (MPM) and the remaining 32 modes.
[0105] FIG. 12 shows exemplary intra prediction directions and intra prediction modes in some examples (e.g., VVC). There are a total of 95 intra prediction modes (modes -14 to 80), among which mode 18 is the horizontal mode, mode 50 is the vertical mode, and modes 2, 34, and 66 are diagonal modes. Modes -1 to -14 and modes 67 to 80 are called wide-angle intra prediction (WAIP) modes (also called wide-angle modes, wide-angle direction modes, etc.).
[0106] FIG. 13 shows Table 1 of the correspondence between intra prediction modes and angle parameters related thereto. In Table 1, predModeIntra indicates the intra prediction mode, and intraPredAngle indicates the intra prediction angle parameter of the corresponding intra prediction mode (e.g., a displacement parameter associated with the intra prediction angle related to the tangent value of the angle). In the example of FIG. 13, the accuracy of the intra prediction angle parameter is 1 / 32. In some examples, when the intra prediction mode has the corresponding value X in Table 1, the actual intraPredAngle parameter is X / 32. For example, in the case of mode 66, the corresponding value in Table 1 is 32, and the actual intraPredAngle parameter is 32 / 32.
[0107] Note that the conventional angular intra prediction direction is defined from 45° to -135° in the clockwise direction. In some embodiments, some conventional angular intra prediction modes are adaptively replaced with the wide-angle intra prediction mode for non-square blocks. The replaced mode is signaled using the original mode index, which is remapped to the index of the wide-angle mode after syntax analysis. The total number of intra prediction modes remains unchanged, i.e., 67, and the coding method of the intra mode does not change.
[0108] In an example, to support the prediction direction, an upper reference with a total width of 2W+1 can be generated, and a left reference with a total height of 2H+1 can be generated, where the width of the current block is W and the height of the current block is H.
[0109] In some examples, the number of modes replaced by the wide-angle direction mode depends on the aspect ratio of the block.
[0110] FIG. 14 shows Table 2 of the intra prediction modes replaced by the wide-angle direction mode based on the aspect ratio of the block. In an example, in the angular mode ang_mode of Table 2, when W / H>1, the angular mode ang_mode is mapped to (65+ang_mode), and when W / H<1, the angular mode ang_mode is mapped to (ang_mode-67).
[0111] According to one aspect of the present disclosure, an intra sub-partition (ISP) coding mode can be used. In the ISP coding mode, the luminance intra prediction block is divided into two or four sub-partitions in the vertical or horizontal direction according to the block size.
[0112] FIG. 15 shows Table 3 associating the number of sub - partitions with the block size. For example, when the block size is 4x4, no partitioning is performed on the block in the ISP encoding mode. When the block size is 4x8 or 8x4, the block is partitioned into two sub - partitions in the ISP encoding mode. For all other block sizes, the block is partitioned into four sub - partitions.
[0113] FIG. 16 shows an example of sub - partitioning of a block having a size of 4x8 or 8x4. In the example of horizontal partitioning, the block is partitioned into two equal sub - partitions each having a size of width×(height / 2). In the example of vertical partitioning, the block is partitioned into two equal sub - partitions each having a size of (width / 2)×height.
[0114] FIG. 17 shows another example of sub - partitioning of a block having a size other than 4x8, 8x4, and 4x4. In the example of horizontal partitioning, the block is partitioned into four equal sub - partitions each having a size of width×(height / 4). In the example of vertical partitioning, the block is partitioned into four equal sub - partitions each having a size of (width / 4)×height. In one example, all sub - partitions satisfy the condition of having at least 16 samples. For the chroma component, ISP is not applied.
[0115] In some examples, each sub - partition is regarded as a TU. For example, for each sub - partition, the decoder can entropy - decode the coefficients sent from the encoder to the decoder, and then the decoder can inverse - quantize and inverse - transform the coefficients to generate the residual of the sub - partition. Further, when the sub - partition is intra - predicted by the decoder, the decoder can add the residual together with the intra - prediction result to obtain the reconstructed samples of the sub - partition. Thus, the reconstructed samples of each sub - partition can be used to generate the prediction of the next sub - partition, and hereinafter, this process is repeated. In the example, all sub - partitions share the same intra - prediction mode.
[0116] In some examples, the ISP algorithm is tested only in the intra - prediction modes that are part of the MPM list. Therefore, when a block uses ISP, it can be inferred that the MPM flag is set. Further, when ISP is used for a particular block, the MPM list can be modified to exclude the DC mode and, in some examples, prioritize the horizontal intra - prediction mode of the horizontal split of ISP and the vertical intra - prediction mode of the vertical split of ISP.
[0117] In some examples, in ISP, since the transformation and reconstruction are performed individually for each sub - partition, each sub - partition can be regarded as a TU.
[0118] In the design of the related ISP algorithm, when the ISP mode is on, the mapping process of the wide - angle mode is executed at the CU level, and when the ISP mode is off, the wide - angle mode mapping process is executed at the TU level, but this is not a unified design.
[0119] Also, in the embodiments of the related ISP, when the current CU is split horizontally, the upper-right adjacent samples are marked as unavailable for use in the second, third, and fourth partitions. When the current CU is split vertically, the lower-left adjacent samples are marked as unavailable for use in the second, third, and fourth partitions, which may not be a desirable design.
[0120] Aspects of the present disclosure provide improved intra prediction techniques for sub-partition (ISP) coding modes. The proposed methods may be used separately or combined in any order. Further, each of the methods (or embodiments), encoders, and decoders may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0121] In the present disclosure, when the intra prediction mode is neither the planar mode nor the DC mode, the intra prediction mode generates prediction samples according to a given prediction direction, such as intra prediction mode 2-66 defined in the VVC draft 2, and the intra prediction mode is called the angular mode. When the intra prediction mode is not the directional intra prediction, for example, when the intra prediction mode is either the planar mode or the DC mode, the intra prediction mode is called the non-angular mode in the present disclosure. Each intra prediction mode is associated with a mode number (also called the intra prediction mode index). For example, in the VVC working draft, the intra prediction modes of planar, DC, horizontal, and vertical are associated with mode numbers 0, 1, 18, and 50, respectively.
[0122] In the present disclosure, the vertical prediction direction is assumed using a prediction angle v. Next, an intra prediction direction such as vertical is defined as an intra prediction direction associated with a prediction angle falling within the range of (v - thr, v + thr), where thr is a given threshold. Further, the horizontal prediction direction is assumed using a prediction angle h. An intra prediction direction such as horizontal is defined as an intra prediction direction associated with a prediction angle classified into (h - thr, h + thr), where thr is a given threshold.
[0123] In the present disclosure, a reference line index is used to refer to a reference line. A reference line adjacent to the current block, which is the reference line closest to the current block, is referred to using a reference line index of 0.
[0124] According to one aspect of the present disclosure, a wide-angle mapping process, i.e., a determination of whether a wide-angle intra prediction angle is applied, is performed at the TU level regardless of whether the current CU has a plurality of TUs (or partitions).
[0125] In one embodiment, a first set of intra prediction modes such as modes 0 to 66 of VCC is used for square-shaped blocks. If the block is non-square-shaped, a subset of the first set of intra prediction modes is replaced with wide-angle intra prediction modes. In some examples, the number of intra prediction modes replaced from the first set to the second set is a function of the aspect ratio of the block. However, if the aspect ratio is out of range, in the example, the number of intra prediction modes replaced from the first set to the second set is clipped to a predefined value. In some examples, if the aspect ratio of the block is greater than Thres1 or less than Thres2, the number of intra prediction modes to be replaced is set equal to K, so that the wide-angle mapping process functions even for blocks with the above aspect ratio. The aspect ratio can refer to either width / height or height / width, or the maximum (width / height, height / width), or a function of width / height and / or height / width.
[0126] In one embodiment, K is a positive integer. In an example, K is set equal to 15. In another example, K is set equal to 16.
[0127] In another embodiment, the replaced intra mode of a block whose aspect ratio is greater than Thres1 or less than Thres2 can be determined according to a predefined table.
[0128] FIG. 18 shows Table 4 used as an example for replacing the intra prediction mode with a wide-angle mode. The number of replaced intra prediction modes is 15.
[0129] FIG. 19 shows Table 5 used as another example for replacing the intra prediction mode with a wide-angle mode. The number of replaced intra prediction modes is 16.
[0130] In another embodiment, the upper threshold Thres1 is set equal to 16, and the lower threshold Thres2 is set equal to 1 / 16.
[0131] In another embodiment, the upper threshold Thres1 is set equal to 32, and the lower threshold Thres2 is set equal to 1 / 32.
[0132] In another embodiment, when the aspect ratio of the current block is greater than L1 (W / H > L1), the upper reference sample with a length of 2W + 1 and the left reference sample with a length of 4H + 1 are fetched, stored, and used in intra prediction. L1 is a positive integer. In an example, L1 is set equal to 16 or 32.
[0133] In another embodiment, when the aspect ratio of the current block is less than L2 (W / H < L2), the upper reference sample with a length of 4W + 1 and the left reference sample with a length of 2H + 1 are fetched, stored, and used in intra prediction. L2 is a positive number. In an example, L2 is set equal to 1 / 16 or 1 / 32.
[0134] In another embodiment, the accuracy of the intra prediction direction is changed to 1 / 64, and the intra prediction directions where tan(α) is equal to 1 / 64, 2 / 64, 4 / 64, 8 / 64, 16 / 64, 32 / 64, 64 / 64 are included. In some examples, the number of modes replaced with the wide-angle mode of blocks with an aspect ratio equal to 16 (1 / 16), 32 (1 / 32), 64 (1 / 64) is set to be equal to 13, 14, and 15, respectively. For example, the number of modes replaced with the wide-angle mode of blocks with an aspect ratio equal to 16 (1 / 16) is set to 13. For example, the number of modes replaced with the wide-angle mode of blocks with an aspect ratio equal to 32 (1 / 32) is set to 14. For example, the number of modes replaced with the wide-angle mode of blocks with an aspect ratio equal to 64 (1 / 64) is set to 15.
[0135] According to another aspect of the present disclosure, when the ISP mode is on, the adjacent samples of the first sub-partition can also be used for the second, third, and fourth partitions.
[0136] In one embodiment, when the current CU is vertically divided, the lower left adjacent sample used for intra prediction of the first sub-partition can also be used as the lower left adjacent sample for the second, third, and fourth partitions, and the lower left adjacent sample used for intra prediction of the first partition is shown in the following figure.
[0137] FIG. 20 shows an example of reference samples in the ISP mode. In the example of FIG. 20, the CU 2010 is vertically divided into a first sub-partition, a second sub-partition, a third sub-partition, and a fourth sub-partition. The adjacent samples of the CU 2010 include an upper adjacent sample 2020, an upper right adjacent sample 2030, a left adjacent sample 2040, and a lower left adjacent sample 2050.
[0138] In one example, in order to perform an intra prediction for the first sub - partition, the upper adjacent sample, the upper - right adjacent sample, the left adjacent sample, and the lower - left adjacent sample of the first sub - partition are available.
[0139] Furthermore, in the example, for performing an intra prediction for the second sub - partition, the lower - left adjacent sample of the second sub - partition has not been decoded and cannot be used. Similarly, the lower - left adjacent samples of the third and fourth sub - partitions cannot be used. In some embodiments, the lower - left adjacent sample 2050 of the first sub - partition is used as the lower - left adjacent sample of the second, third, and fourth sub - partitions.
[0140] In another embodiment, when the current CU is horizontally split, the upper - right adjacent sample used for the intra prediction of the first sub - partition can also be used as the upper - right adjacent sample of the second, third, and fourth sub - partitions. The following figure shows the upper - right adjacent sample used for the intra prediction of the first sub - partition.
[0141] FIG. 21 is a diagram showing an example of reference samples in the ISP mode. In the example of FIG. 21, CU2110 is horizontally split into a first sub - partition, a second sub - partition, a third sub - partition, and a fourth sub - partition. The adjacent samples of CU2110 include an upper adjacent sample 2120, an upper - right adjacent sample 2130, a left adjacent sample 2140, and a lower - left adjacent sample 2150.
[0142] In one example, in order to perform an intra prediction for the first sub - partition, the upper adjacent sample, the upper - right adjacent sample, the left adjacent sample, and the lower - left adjacent sample of the first sub - partition are available.
[0143] Furthermore, in the example, to intra-predict the second sub-partition, the adjacent sample to the upper right of the second sub-partition has not been decoded and cannot be used. Similarly, the adjacent samples to the upper right of the third and fourth sub-partitions cannot be used. In some embodiments, the upper right adjacent sample 2130 of the first sub-partition is used as the adjacent sample to the upper right of the second sub-partition, the third sub-partition, and the fourth sub-partition.
[0144] In another embodiment, when the current CU is horizontally split, for planar intra-prediction, the sample to the upper right of the first sub-partition is also shown as the sample to the upper right of the second, third, and fourth sub-partitions of the current CU.
[0145] In some examples, when the size of the current CU is N×4 or N×8, where N is any positive integer such as 2, 4, 8, 16, 32, 64, 128, etc., the current CU is horizontally split, and the sample to the right of the first sub-partition is also shown as the sample to the upper right of the second, third, and fourth sub-partitions of the current CU.
[0146] FIG. 22 is a diagram showing an example of reference samples in the ISP mode for planar mode prediction. In the example of FIG. 22, CU2210 is horizontally split into a first sub-partition, a second sub-partition, a third sub-partition, and a fourth sub-partition. The adjacent samples used for the planar mode prediction of CU2210 include an upper adjacent sample 2220, an upper right adjacent sample 2230, a left adjacent sample 2240, and a lower left adjacent sample 2250.
[0147] In one example, to intra-predict the first sub-partition in the planar mode, the upper adjacent sample, the upper right adjacent sample, the left adjacent sample, and the lower left adjacent sample of the first sub-partition are available.
[0148] Furthermore, in the example, in order to perform intra prediction on the second sub-partition in the planar mode, the adjacent sample to the upper right of the second sub-partition has not been decoded and cannot be utilized. Similarly, the adjacent sample to the upper right of the third sub-partition and the adjacent sample to the upper right of the fourth sub-partition cannot be utilized. In some embodiments, the samples 2230 adjacent to the upper right of the first sub-partition are used as the adjacent samples to the upper right of the second sub-partition, the third sub-partition, and the fourth sub-partition, respectively.
[0149] In some examples, when the size of the current CU is N×4 or N×8, where N is any positive integer such as 2, 4, 8, 16, 32, 64, 128, etc., the current CU is split horizontally, and the positions of the upper right samples used for the planar prediction of the second, third, and fourth partitions of the current CU may vary for each partition.
[0150] In another embodiment, when the current CU is split vertically, for planar intra prediction, the sample at the lower left of the first sub-partition is also shown as the sample at the lower left of the second, third, and fourth sub-partitions of the current CU.
[0151] In some examples, when the size of the current CU is 4×N or 8×N, where N is any positive integer such as 2, 4, 8, 16, 32, 64, 128, etc., the current CU is split vertically, and the sample to the left of the first sub-partition is also shown as the sample at the lower left of the second, third, and fourth sub-partitions of the current CU.
[0152] FIG. 23 is a diagram showing an example of a reference sample in the ISP mode for planar intra prediction. In the example of FIG. 23, CU 2310 is vertically divided into a first sub-partition, a second sub-partition, a third sub-partition, and a fourth sub-partition. The adjacent samples (in planar intra prediction) of CU 2310 include an upper adjacent sample 2320, an upper right adjacent sample 2330, a left adjacent sample 2340, and a lower left adjacent sample 2350.
[0153] In one example, in order to perform intra prediction on the first sub-partition in the planar mode, the upper adjacent sample, the upper right adjacent sample, the left adjacent sample, and the lower left adjacent sample of the first sub-partition are available.
[0154] Furthermore, in the example, in order to perform intra prediction on the second sub-partition in the planar mode, the lower left adjacent sample of the second sub-partition has not been decoded and cannot be used. Similarly, the lower left adjacent samples of the third sub-partition and the fourth sub-partition cannot be used. In some embodiments, the lower left adjacent sample 2350 of the first sub-partition is used as the lower left adjacent sample of the second sub-partition, the third sub-partition, and the fourth sub-partition, respectively.
[0155] In some examples, when the size of the current CU is 4×N or 8×N, where N is any positive integer such as 2, 4, 8, 16, 32, 64, 128, etc., the current CU is vertically divided. The positions of the lower left samples used for planar prediction of the second, third, and fourth partitions of the current CU may be different for each partition.
[0156] According to another aspect of the present disclosure, the number of sub-block partitions can depend on encoded information including, but not limited to, the block size of the current block and the intra prediction mode of adjacent blocks. Therefore, the number of sub-block partitions can be inferred and explicit signaling is not necessary.
[0157] In one embodiment, for an N×4 or N×8 CU, N is a positive integer such as 8, 16, 32, 64, 128, etc., the current CU can only be divided vertically, and one flag indicating whether the current CU is divided into two partitions or four partitions vertically is notified.
[0158] For example, when the block size is N×4, instead of partitioning it into four N×1 sub - partitions horizontally, the block is partitioned vertically into two N / 2×4 partitions. In one example, when the block size is N×8, instead of partitioning it into four N×2 sub - partitions horizontally, the block is partitioned vertically into two N / 2×8 partitions.
[0159] In another example, when the block size is N×8, instead of partitioning it into four N×2 sub - partitions horizontally, the block is partitioned horizontally into two N / 2×4 partitions.
[0160] In another embodiment, for a 4×N or 8×N CU, N is a positive integer such as 8, 16, 32, 64, 128, etc., the current CU can only be divided horizontally, and one flag indicating whether the current CU is divided into two partitions or four partitions horizontally is notified.
[0161] In one example, when the block size is 4×N, instead of partitioning it into four 1×N sub - partitions vertically, the block is partitioned horizontally into two 4×N / 2 partitions. In one example, when the block size is 8×N, instead of partitioning it into four 2×N sub - partitions vertically, the block is partitioned horizontally into two 8×N / 2 partitions.
[0162] In another example, when the block size is 8×N, instead of partitioning it into four 2×N sub - partitions vertically, the block is partitioned horizontally into two 4×N / 2 partitions.
[0163] According to another aspect of the present disclosure, for adjacent reference lines with the ISP mode disabled, adjacent reference lines with the ISP mode enabled, and non-adjacent reference lines, the same MPM list construction process is shared and the same candidate order is used. The planar mode and the DC mode are always inserted into the MPM lists at indices 0 and 1 in some examples.
[0164] In one embodiment, when the reference line index is notified as 0, two contexts are used for the first bin of the MPM index. When at least one of the adjacent blocks satisfies the following conditions: 1) the MPM flag is true, 2) the reference line index is 0, and 3) the MPM index is less than Th, the first context is used. Otherwise, the second context is used. Th is a positive integer such as 1, 2, 3, etc.
[0165] In another embodiment, when the reference line index is notified as 0, only one context is used for the first bin of the MPM index.
[0166] In another embodiment, when the reference line index is notified as 0, the ISP mode is turned off, and only one context is used for the second bin of the MPM index.
[0167] In another embodiment, if the above adjacent block exceeds the current CTU row, the above adjacent block is marked as unavailable for MPM index context derivation.
[0168] In another embodiment, when the reference line index is notified as 0, the first K bins of the MPM index depend on the MPM flag, and / or the MPM index, and / or the ISP flag, and / or the reference line index of its adjacent block. K is a positive integer such as 1 or 2.
[0169] In one example, when the adjacent reference line of index 0 is notified, two contexts are used for the first bin of the MPM index. If the ISP flag of the current block is on, the first context is used. Otherwise, the second context is used.
[0170] In another example, when the adjacent reference line of index 0 is notified, two contexts are used for the first bin of the MPM index. If at least one of the ISP flags of the current block and its adjacent block is on, the first context is used. Otherwise, the second context is used.
[0171] FIG. 24 shows a flowchart outlining a process (2400) according to an embodiment of the present disclosure. Since the process (2400) can be used for block reconstruction, it generates a predicted block of a block being reconstructed. In various embodiments, the process (2400) is executed by a processing circuit such as the processing circuits of terminal devices (310), (320), (330), and (340), the processing circuit that executes the functions of video encoder (403), the processing circuit that executes the functions of video decoder (410), the processing circuit that executes the functions of video decoder (510), and the processing circuit that executes the functions of video encoder (603). In some embodiments, the process (2400) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (2400). The process starts at (S2401) and proceeds to (S2410).
[0172] (In S2410), the prediction information of the current block is decoded from the encoded video bitstream. The prediction information indicates a first intra prediction mode in a first set of intra prediction modes for square-shaped intra prediction. In some examples, the prediction information also indicates the residual information of the sub-partition of the current block. For example, the prediction information includes the residual data of the TU.
[0173] In (S2420), the first intra prediction mode is remapped to a second intra prediction mode in a second set of intra prediction modes for non-square shape intra prediction based on the shape of the first sub-partition of the current block.
[0174] In (S2430), the samples of the sub-partition are reconstructed according to the second intra prediction mode and the residual information of the sub-partition. Thereafter, the process proceeds to (S2499) and ends.
[0175] FIG. 25 shows a flowchart outlining a process (2500) according to an embodiment of the present disclosure. Since the process (2500) can be used for reconstruction of a block, a predicted block of the block being reconstructed is generated. In various embodiments, the process (2500) is executed by a processing circuit such as the processing circuits of terminal devices (310), (320), (330) and (340), the processing circuit that executes the functions of video encoder (403), the processing circuit that executes the functions of video decoder (410), the processing circuit that executes the functions of video decoder (510), and the processing circuit that executes the functions of video encoder (603). In some embodiments, the process (2500) is implemented by software instructions, and thus when the processing circuit executes the software instructions, the processing circuit executes the process (2500). The process starts at (S2501) and proceeds to (S2510).
[0176] In (S2510), the prediction information of the current block is decoded from the encoded video bitstream.
[0177] In (S2520), the current block is partitioned into at least a first sub-partition and a second sub-partition based on the prediction information. For example, the current block can be partitioned as shown in FIGS. 16 and 17.
[0178] In (S2530), the samples of the first sub - partition and the second sub - partition are reconstructed based on at least the adjacent samples of the current block outside the adjacent region of at least one of the first sub - partition and the second sub - partition, as described with reference to FIGS. 20 - 23. Thereafter, the process proceeds to (S2599) and ends.
[0179] The above - described technology can be implemented as computer software using computer - readable instructions and can be physically stored on one or more computer - readable media. For example, FIG. 26 shows a computer system (2600) suitable for implementing a particular embodiment of the disclosed subject matter.
[0180] The computer software can be coded using any suitable machine code or computer language, which is subject to mechanisms such as assembly, compilation, linking, etc., to create code that includes instructions that can be executed directly, or through interpretation, micro - code execution, etc., by one or more central processing units (CPUs), graphics processing units (GPUs), etc.
[0181] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet - of - Things devices, etc.
[0182] The components shown in FIG. 26 for the computer system (2600) are exemplary in nature and are not intended to suggest any limitation regarding the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. Also, the configuration of the components should not be construed as having dependencies or requirements related to any one or combination of the components shown in the exemplary embodiment of the computer system (2600).
[0183] The computer system (2600) can include a specific human interface input device. Such a human interface input device can respond to input by one or more human users, for example, through tactile input (such as keystrokes, swipes, movement of a data glove, etc.), voice input (such as voice, clapping, etc.), visual input (such as gestures, etc.), and olfactory input (not shown). The human interface device can also be used to capture specific media such as voice (such as speech, music, environmental sounds, etc.), images (such as scanned images, photographic images taken by a still image camera, etc.), video (such as 2D video, 3D video including stereoscopic video, etc.), which are not necessarily directly related to conscious human input.
[0184] The input human interface device can include one or more of a keyboard (2601), a mouse (2602), a trackpad (2603), a touch screen (2610), a data glove (not shown), a joystick (2605), a microphone (2606), a scanner (2607), and a camera (2608) (only one of each is depicted).
[0185] The computer system (2600) can also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen (2610), a data glove (not shown), or a joystick (2605), although there can also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers (2609), headphones (not shown), etc.), visual output devices (such as screens (2610) including CRT screens, LCD screens, plasma screens, OLED screens, which can output three-dimensional or more output by means such as two-dimensional visual output or stereoscopic output, regardless of whether they have a touch screen input function and regardless of whether they have a tactile feedback function, virtual reality glasses (not shown), holographic displays, smoke tanks (not shown)), and can also include printers (not shown).
[0186] The computer system (2600) can also include optical media such as a CD / DVD ROM / RW (2620) with a medium (2621) such as a CD / DVD, a thumb drive (2622), a removable hard drive or solid state drive (2623), legacy magnetic media such as tapes and floppy disks (not shown), special ROM / ASIC / PLD-based devices such as security dongles (not shown), etc., and human-accessible storage devices and their associated media.
[0187] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0188] The computer system (2600) can also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, terrestrial broadcast TV, etc., vehicle and industrial including CANBus. A particular network typically requires an external network interface adapter connected to a particular general-purpose data port or peripheral bus (2649) (e.g., the USB port of the computer system (2600)), while others are generally integrated into the core of the computer system (2600) by connecting to the system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2600) can communicate with other entities. Such communication can be unidirectional, receive only (e.g., TV broadcast), transmit only unidirectional (e.g., from a CANbus to a particular CANbus device), or bidirectional, e.g., to other computer systems using local or wide area digital networks. As described above, specific protocols and protocol stacks can be used for each of these networks and network interfaces.
[0189] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core (2640) of the computer system (2600).
[0190] The core (2640) can include one or more central processing units (CPUs) (2641), a graphics processing unit (GPU) (2642), special programmable processing units in the form of field programmable gate arrays (FPGAs) (2643), and hardware accelerators (2644) for specific tasks, etc. These devices can be connected via a system bus (2648) together with read-only memory (ROM) (2645), random access memory (2646), internal mass storage devices such as internal hard drives and SSDs that are not accessible to users (2647). In some computer systems, it is possible to access the system bus (2648) in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be directly connected to the system bus (2648) of the core or via a peripheral bus (2649). Architectures of peripheral buses include PCI, USB, etc.
[0191] The CPU (2641), GPU (2642), FPGA (2643), and accelerator (2644) can execute specific instructions that can together constitute the aforementioned computer code. That computer code can be stored in the ROM (2645) or RAM (2646). Migration data can also be stored in the RAM (2646), while persistent data can be stored, for example, in the internal mass storage device (2647). Fast storage and reading to / from any of the memory devices can be made possible by using cache memory that is closely related to one or more CPUs (2641), GPUs (2642), the mass storage device (2647), ROM (2645), RAM (2646), etc.
[0192] A computer-readable medium can have computer code thereon for performing operations implemented on various computers. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure, or they can be of the kind well-known and available to those having skill in the computer software arts.
[0193] As an example, and not by way of limitation, a computer system (2600) having an architecture, specifically a core (2640), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media can be the user-accessible mass storage device introduced above, as well as media associated with specific storage devices of the core (2640) of a non-transitory nature such as an internal core mass storage device (2647) or ROM (2645). The software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (2640). The computer-readable media can include one or more memory devices or chips, depending on specific requirements. The software can cause the core (2640) and specifically the processors (including a CPU, GPU, FPGA, etc.) therein to define data structures stored in the RAM (2646) and modify such data structures according to processes defined by the software, thereby executing specific processes or specific parts of specific processes described herein. Additionally, or alternatively, the computer system can provide functionality as a result of logic hardwired or otherwise embodied in a circuit (e.g., an accelerator (2644)), which can operate in place of or in conjunction with software to execute specific processes or specific parts of specific processes described herein. Optionally, references to software can include logic, and vice versa. References to computer-readable media can optionally include circuits (such as integrated circuits (ICs)) that store software for execution, circuits that embody logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software. Appendix A: Abbreviations JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplemental Enhancement Information VUI: Video User Information GOP: Group of Pictures TU: Transform Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coding Tree Block PB: Prediction Block HRD: Hypothetical Reference Decoder SNR: Signal-to-Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: Cathode Ray Tube LCD: Liquid Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Versatile Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field-Programmable Gate Array SSD: Solid State Drive IC: Integrated Circuit CU: Coding Unit
[0194] Although this disclosure describes several exemplary embodiments, there are changes, substitutions, and various alternative equivalents within the scope of this disclosure. Therefore, it will be understood that those skilled in the art can devise numerous systems and methods that embody the principles of the disclosure and are thus within its spirit and scope, even though not explicitly shown or described herein.
Explanation of Signs
[0195] 101 Sample 102 Arrow 103 Arrow 104 Block 201 Schematic Diagram 300 Communication System 310 Terminal Device 320 Terminal Device 330 Terminal Device 350 Network 400 Communication System 401 Video Source 402 Video Picture Stream 403 Video Encoder 404 Encoded Video Data 405 Streaming Server 406 Client Subsystem 407 Incoming Copy 410 Video Decoder 411 Video Picture 412 Display 413 Capture Subsystem 420 Electronic Device 430 Electronic Device 501 Channel 510 Video Decoder 512 Rendering Device 515 Buffer Memory 520 Entropy Decoder / Parser 521 Symbol 530 Electronic Device 531 Receiver 551 Scaler / Inverse Conversion Unit 552 Intra-Picture Prediction Unit 553 Motion Compensation Prediction Unit 555 Aggregator 556 Loop Filter Unit 557 Reference Picture Memory 558 Picture Buffer 601 Video Source 603 Video Encoder 620 Electronic Device 630 Source Coder 632 Encoding Engine 633 Local Decoder 634 Reference Picture Memory 635 Predictor 640 Transmitter 643 Encoded Video Sequence 645 Entropy Coder 550 Controller 660 Communication Channel 703 Video Encoder 721 General-Purpose Controller 722 Intra Encoder 723 Residual Calculator 724 Residual Encoder 725 Entropy Encoder 726 Switch 728 Residual Decoder 730 Inter Encoder 810 Video Decoder 871 Entropy Decoder 872 Intra Decoder 873 Residual Decoder 874 Reconstruction Module 880 Inter Decoder 900 QTBT Structure 901 Block 911 Block 912 Block 913 Block 914 Block 921 Block 922 Block 923 Block 924 Block 925 Block 926 Block 927 Block 928 Block 950 Root Node 961 Node 962 Node 963 Node 964 Node 971 Node 972 Node 973 Node 974 Node 975 Node 976 Node 977 Node 978 Node 981 Node 982 Node 983 Node 984 Node 985 Node 986 Node 991 Node 992 Node 1010 Block 1011 Sub - block 1012 Sub - block 1013 Sub - block 1020 Block 1021 Sub - block 1022 Sub - block 1023 Sub - block 2020 Upper Adjacent Sample 2030 Upper - Right Adjacent Sample 2040 Left Adjacent Sample 2050 Lower - Left Adjacent Sample 2120 Upper Adjacent Sample 2130 Upper - Right Adjacent Sample 2140 Left Adjacent Sample 2150 Lower - Left Adjacent Sample 2220 Upper Adjacent Sample 2230 Upper-right adjacent sample 2240 Left adjacent sample 2250 Lower-left adjacent sample 2320 Upper adjacent sample 2330 Upper-right adjacent sample 2340 Left adjacent sample 2350 Lower-left adjacent sample 2400 Process 2500 Process 2600 Computer system 2601 Keyboard 2602 Mouse 2603 Track pad 2605 Joystick 2606 Microphone 2607 Scanner 2608 Camera 2609 Speaker 2610 Touch screen 2621 Media 2622 Thumb drive 2623 Solid state drive 2640 Core 2643 Field programmable gate array (FPGA) 2644 Accelerator 2645 Read only memory (ROM) 2646 Random access memory 2647 Mass storage device 2648 System bus 2649 Peripheral bus
Claims
[Claim 1] decoding, by a processor, prediction information for the current block from the encoded video bitstream; determining, by the processor, a first sub-partition and a second sub-partition of the current block based on the decoded prediction information; reconstructing, by the processor, the first subpartition and the second subpartition of the current block based on at least adjacent samples of the current block that are outside adjacent regions of at least one of the first subpartition and the second subpartition; 23. A method for video decoding in a decoder, comprising:
Citation Information
Patent Citations
Method and apparatus for prediction based on block shape
KR1020180107762A
Video and image coding with wide-angle intra prediction
WO2018127624A1