Improvement of Intra-Mode Propagation
By employing template matching prediction and a universal intra mode map to refine intra mode propagation, the method addresses inefficiencies in determining MPM lists, enhancing video coding efficiency and compression ratios.
Patent Information
- Application Number
- JP2023558771
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-09-14
- Filing Date
- 2022-09-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-09-16
AI Technical Summary
Existing video coding technologies face inefficiencies in intra prediction modes, particularly in determining the most probable mode (MPM) lists, which can lead to increased bit usage for less likely prediction directions, thereby reducing compression efficiency.
The proposed method involves improving intra mode propagation by using template matching prediction (TMP) and derived modes for chroma, along with a universal intra mode map to enhance the accuracy of MPM lists, reducing bit usage for less likely directions.
This approach enhances video coding efficiency by optimizing intra prediction modes, reducing bit usage for less likely directions and improving compression ratios.
Smart Images

Figure 0007704886000001 
Figure 0007704886000002 
Figure 0007704886000003
Abstract
Description
Technical Field
[0001] [Incorporation by Reference] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 254,526, filed Oct. 11, 2021, entitled “Intra Mode Propagation,” and to U.S. Patent Application No. 17 / 945,032, filed Sep. 14, 2022, entitled “IMPROVEMENT ON INTRA MODE PROPAGATION.” The disclosure of the prior application is hereby incorporated by reference in its entirety.
[0002] [Technical Field] The present disclosure generally describes embodiments related to video coding.
Background Art
[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. The research of the inventors, as currently named, is not admitted as prior art for the present disclosure, either expressly or implicitly, to the same extent as any aspect of the description that may not otherwise be considered prior art at the time of filing as long as that research is not described in this background art section.
[0004] Video coding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each picture having, for example, spatial dimensions of 1920×1080 luminance samples and associated chrominance samples. A series of pictures can have, for example, a fixed or variable picture rate of 60 pictures per second or 60 Hz (informally also known as the frame rate). Uncompressed video has specific bitrate requirements. For example, 8-bit / sample 1080p60 4:2:0 video (1920×1080 luminance sample resolution at 60 Hz frame rate) requires a bandwidth close to 1.5 Gbit / second. Such video requires more than 600 gigabytes of storage space per hour.
[0005] One purpose of video coding and decoding can be to reduce redundancy in the input video signal by compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements, sometimes by more than two orders of magnitude. Both lossless compression and lossy compression as well as combinations thereof can be employed. Lossless compression refers to techniques that can reconstruct an exact copy of the original signal from the compressed signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal can be small enough to make the reconstructed signal useful for its intended application. In the case of video, lossy compression is widely employed. The amount of allowable distortion depends on the application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect the fact that a higher allowable / tolerable distortion can result in a higher compression ratio.
[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy coding.
[0007] Video coding techniques can include techniques known as intra coding. In intra coding, sample values are represented without reference to samples from previously reconstructed reference pictures or other data. In some video coders, a picture is spatially subdivided into blocks of samples. If all blocks of samples are coded in an intra mode, that picture can be an intra picture. Those derivatives such as intra pictures and independent decoder refresh pictures can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and video session or as a still image. Samples of an intra block can be subjected to a transform and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique to minimize sample values in a pre-transform region. In some cases, the smaller the post-transform DC value and the smaller the AC coefficients, the fewer bits are required at a given quantization step size to represent the block after entropy coding.
[0008] For example, conventional intra coding as known from MPEG-2 generation coding techniques does not use intra prediction. However, some newer video compression techniques include techniques that attempt, for example, from encoded and / or decoded blocks of data that are spatially adjacent and precede in decoding order and / or from surrounding sample data and / or metadata obtained during decoding. Such techniques are hereinafter referred to as "intra prediction" techniques. Note that in at least some cases, intra prediction uses only reference data from the current picture being reconstructed and not from a reference picture.
[0009] There can be many different forms of intra prediction. If two or more of such techniques can be used in a given video coding technology, the technique in use can be coded in an intra prediction mode. In certain cases, the mode can have sub - modes and / or parameters, which can be coded individually or can be included in the mode codeword. Which codeword to use for a given mode, sub - mode, and / or parameter combination can affect the coding efficiency gain by intra prediction and can also affect the entropy coding technique used to convert the codeword into the bitstream.
[0010] Certain modes of intra prediction were introduced in H.264, improved in H.265, and further improved in more recent coding technologies such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). The predictor block can be formed using adjacent sample values belonging to already available samples. The sample values of the adjacent samples are copied to the predictor block according to a direction. The reference to the direction in use can be coded in the bitstream or can itself be predicted.
[0011] Referring to Figure 1, in the lower right, a subset of 9 predictor directions known from the 33 possible predictor directions of H.265 (corresponding to 33 angular modes out of 35 intra modes) is shown. The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples in the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples in the lower left of sample (101) at an angle of 22.5 degrees from the horizontal.
[0012] Still referring to FIG. 1, in the upper left, a square block (104) of 4×4 samples (indicated by the thick dashed line) is shown. The square block (104) contains 16 samples each labeled with an "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Since the block is 4×4 samples in size, S44 is in the lower right. Reference samples following a similar numbering scheme are further shown. The reference samples are labeled with an "R" and its Y position (e.g., row index) and X position (column index) relative to the block (104). In both H.264 and H.265, the predicted samples are adjacent to the block being reconstructed and thus there is no need to use negative values.
[0013] Intra-picture prediction can function by copying the reference sample value from adjacent samples as assigned by the signaled prediction direction. For example, assume that the coded video bitstream contains signaling indicating a prediction direction that matches the arrow (102) for this block, i.e., the samples are predicted from one or more prediction samples in the upper right at an angle of 45 degrees from the horizontal. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.
[0014] In certain cases, especially when the direction is not evenly divisible by 45 degrees, the values of multiple reference samples can be combined, for example, by interpolation, to calculate the reference sample.
[0015] The number of possible directions has been increasing as video coding technology has evolved. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions at the time of disclosure. Experiments have been conducted to identify the most likely directions, and using specific techniques in entropy coding, those likely directions are represented with a few bits, while accepting a specific penalty for the less likely directions. Further, the direction itself may be predicted from adjacent directions used in adjacent, already decoded blocks.
[0016] FIG. 2 shows a schematic diagram (201) showing 65 intra prediction directions by JEM to illustrate the prediction directions increasing over time.
[0017] The mapping of intra prediction direction bits in the coded video bitstream representing the direction can vary for each video coding technology and can involve, for example, a simple direct mapping of the prediction direction to the intra prediction mode, to the codeword, to a complex adaptive scheme with the most probable mode, and similar techniques. However, in all cases, there may be specific directions that are statistically less likely to occur in the video content than certain other directions. Since the goal of video compression is redundancy reduction, those less likely directions are represented by more bits than the more likely directions in a well - functioning video coding technology. SUMMARY OF THE INVENTION
[0018] Aspects of the present disclosure provide methods and apparatus for video encoding / decoding. In some examples, an apparatus for video decoding includes a receiving circuit and a processing circuit.
[0019] According to one aspect of the present disclosure, a method for video decoding executed in a video decoder is provided. In this method, (i) the current block of a picture and (ii) the coded information of the reconstructed area of the picture can be received from a coded video bitstream. A matching area of the current block in the reconstructed area of the picture can be determined. A first corresponding position of the current block can be determined within the current block. The first corresponding position can include a first coordinate value on a first axis and a second coordinate value on a second axis with respect to a first reference point on the current block. The first axis can be perpendicular to the second axis. A corresponding block of the current block can be determined in the matching area. The corresponding block can include a second corresponding position having a first coordinate value on the first axis and a second coordinate value on the second axis with respect to a second reference point on the matching area. The intra prediction mode of the current block can be determined based on the corresponding block.
[0020] To determine the matching area, a plurality of candidate matching areas can be searched in the reconstructed area of the picture. A respective cost value between the template area of the current block and the respective template area of each of the plurality of candidate matching areas can be determined. The matching area can be determined as the candidate matching area having the minimum cost value among the plurality of candidate matching areas. The template area of the current block can include a first area adjacent to the left side of the current block and a second area adjacent to the upper side of the current block. The respective template area of each of the plurality of candidate matching areas can include a first area adjacent to the left side of each one of the plurality of candidate matching areas and a second area adjacent to the upper side of each one of the plurality of candidate matching areas.
[0021] In some embodiments, the first corresponding position of the current block can be predefined or signaled in the coded information.
[0022] In some embodiments, the first corresponding position may be the center of the current block.
[0023] In some embodiments, the intra prediction mode of the corresponding block may be determined as the intra prediction mode of the current block.
[0024] In some embodiments, the intra prediction mode of the current block may be determined based on a universal intra mode map. The universal intra mode map can divide the matching area into a plurality of sub-areas. Each of the plurality of sub-areas may be associated with a respective intra prediction mode. The intra prediction mode of the current block may be the intra prediction mode associated with the sub-area that includes the second corresponding position among the plurality of sub-areas.
[0025] According to another aspect of the present disclosure, a method of video coding executed in a video decoder may be provided. In this method, (i) a chroma coding unit (CU) and (ii) coding information of a luma area may be received from a coded video bitstream. A corresponding position of the chroma CU may be determined within the chroma CU. The corresponding position may include a first coordinate value on a first axis and a second coordinate value on a second axis. The first axis may be perpendicular to the second axis. A collocated luma CU of the chroma CU may be determined in the luma area. The collocated luma CU may include the corresponding position. The intra prediction mode of the chroma CU may be determined based on the collocated luma CU.
[0026] In one embodiment, in response to a collocated luma CU being intra-coded based on an intra prediction mode, the intra prediction mode of the collocated luma CU can be determined as the intra prediction mode of the chroma CU. In other embodiments, in response to a collocated luma CU not being intra-coded, the propagation intra mode of the collocated luma CU can be determined as the intra prediction mode of the chroma CU. The propagation intra mode of the collocated luma CU can be obtained based on the intra prediction mode of an adjacent luma CU of the collocated luma CU.
[0027] In some embodiments, the corresponding position of the chroma CU can be predefined or signaled in the coding information.
[0028] In some embodiments, the intra prediction mode of the chroma CU can be determined based on a universal intra mode map. The universal intra mode map can divide a luma area into a plurality of sub-areas. Each of the plurality of sub-areas can be assigned a respective intra prediction mode. The intra prediction mode of the chroma CU can be the intra prediction mode associated with the sub-area that includes the corresponding position among the plurality of sub-areas.
[0029] In another aspect of the present disclosure, an apparatus is provided. The apparatus includes a processing circuit. The processing circuit can be configured to execute any of the methods for video coding.
[0030] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video coding, cause the computer to execute any of the methods for video coding.
Brief Description of the Drawings
[0031] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12A
Figure 12B
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
[0032] FIG. 3 shows a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional data transmission. For example, the terminal device (310) can code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The coded video data can be transmitted in the form of one or more coded video bitstreams. The terminal device (320) can receive the coded video data from the network (350), decode the coded video data to restore the video pictures, and display the video pictures according to the restored video data. Unidirectional data transmission can be common in media serving applications and the like.
[0033] In another example, the communication system (300) includes, for example, a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data that may occur during a video conference. For bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) may code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) may also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), may decode the encoded video data to restore the video picture, and may display the video picture on a display device accessible according to the restored video data.
[0034] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) may be shown as a server, a personal computer, and a smartphone, but the principles of the present disclosure may not be so limited. Embodiments of the present disclosure are applied to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (350) represents any number of networks that transmit encoded video data among the terminal devices (310), (320), (330), and (340), including, for example, wireline (wired) and / or wireless communication networks. The communication network (350) may exchange data on a circuit-switched channel and / or a packet-switched channel. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (350) may not be important for the operation of the present disclosure, unless otherwise described herein below.
[0035] FIG. 4 shows an arrangement of a video encoder and a video decoder in a streaming environment as an example of an application example for the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications including, for example, storage of compressed video to digital media including videoconferencing, digital TV, CDs, DVDs, memory sticks, and the like.
[0036] A streaming system may include a capture subsystem (413) that can include, for example, a video source (401) that creates a stream (402) of uncompressed video pictures, such as a digital camera. In one example, the stream (402) of video pictures includes samples taken by a digital camera. The stream (402) of video pictures is drawn in bold to emphasize the high data volume compared to the encoded video data (404) (or encoded video bitstream) and can be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoded video data (404) (or encoded video bitstream (404)) is drawn in thin lines to emphasize the lower data volume compared to the stream (402) of video pictures and can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes an incoming copy (407) of the encoded video data and creates an outgoing stream (411) of video pictures that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to a particular video coding / compression standard. Examples of such standards include ITU-T Recommendation H.265.In one example, the video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.
[0037] Note that the electronic devices (420) and (430) can include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and similarly the electronic device (430) can include a video encoder (not shown).
[0038] FIG. 5 shows a block diagram of a video decoder (510) according to an embodiment of the present disclosure. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (410) of the example of FIG. 4.
[0039] The receiver (531) receives one or more coded video sequences to be decoded by the video decoder (510), and in the same or another embodiment may receive one coded video sequence at a time, where the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequences may be received from a channel (501) that may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data along with other data that may be transferred to respective using entities (not shown), such as coded audio data and / or auxiliary data streams. The receiver (531) may separate the coded video sequences from the other data. To handle network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter, “parser (520)”). In certain applications, the buffer memory (515) is part of the video decoder (510). In other applications, it may be external to the video decoder (510) (not shown). In yet other applications, for example, to handle network jitter, a buffer memory (not shown) may be provided external to the video decoder (510), and in addition, another buffer memory (515) may be provided inside the video decoder (510), for example, to handle playback timing. When the receiver (531) is receiving data from a storage / transfer device with sufficient bandwidth and controllability or from a synchronous network, the buffer memory (515) may not be required or may be small. For use on a best-effort packet network such as the Internet, the buffer memory (515) may be required, can be relatively large, advantageously can be of an adaptive size, and can be at least partially implemented within an operating system or similar element (not shown) external to the video decoder (510).
[0040] Video decoder (510) may include a parser (520) for reconstructing symbols (521) from the encoded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (510) and, optionally, information for controlling a rendering device (such as a display screen) (512) that is not an integral part of the electronic device (530) but can be coupled to the electronic device (530), as shown in FIG. 5. The control information for the rendering device(s) may be in the form of supplementary enhancement information (SEI message) or a video user utility information (VUI) parameter set fragment (not shown). The parser (520) may perform syntax analysis / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence can follow video coding techniques or standards and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract a set of subgroup parameters regarding at least one of the subgroups of pixels within the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroups can include picture groups (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The parser (520) may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the encoded video sequence.
[0041] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).
[0042] The reconstruction of symbol (521) may involve multiple different units depending on the type of the coded video picture or a portion thereof (such as inter and intra pictures, inter and intra blocks, etc.) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the coded video sequence by parser (520). Such a flow of subgroup control information between parser (520) and the following multiple units is not shown for clarity.
[0043] In addition to the functional blocks already described, video decoder (510) may be conceptually subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0044] The first unit is a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives, as symbols (521), the quantized transform coefficients and control information including which transform to use, block size, quantization coefficient, quantization scaling matrix, etc. from parser (520). The scaler / inverse transform unit (551) can output a block including sample values that can be input to aggregator (555).
[0045] In some cases, the output samples of the scaler / inverse transform unit (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from previously reconstructed pictures but can use prediction information from the previously reconstructed part of the current picture. Such prediction information may be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses the surrounding already reconstructed information fetched from the current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, the partially reconstructed current picture and / or the fully reconstructed current picture. The aggregator (555) may, in some cases, add, for each sample, the prediction information generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).
[0046] In other cases, the output samples of the scaler / inverse transform unit (551) may relate to inter-coded and, in some cases, motion-compensated blocks. In such cases, the motion-compensation prediction unit (553) can access the reference picture memory (557) to fetch the samples used for prediction. According to the symbol (521) related to the block, after motion-compensating the fetched samples, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case, called the residual samples or residual signal) to generate the output sample information. The address in the reference picture memory (557) where the motion-compensation prediction unit (553) fetches the prediction samples can be controlled by the motion vectors available to the motion-compensation prediction unit (553) in the form of, for example, a symbol (521) having X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory (557) when an accurate sub-sample motion vector is used, a motion vector prediction mechanism, etc.
[0047] The output sample of the aggregator (555) can undergo various loop filtering techniques in the loop filter unit (556). The video compression technology is controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), and can include in-loop filter techniques made available to the loop filter unit (556) as symbols (521) from the parser (520), but can also respond to meta information obtained during the decoding of the previous part (in decoding order) of the encoded picture or encoded video sequence, and can also respond to previously reconstructed and loop-filtered sample values.
[0048] The output of the loop filter unit (556) can be an output to the render device (512) and can be a sample stream that can be stored in the reference picture memory (557) for use in future inter-picture prediction.
[0049] When a specific encoded picture is completely reconstructed, it can be used as a reference picture for future prediction. For example, when the encoded picture corresponding to the current picture is completely reconstructed and the encoded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before starting the reconstruction of the next encoded picture.
[0050] The video decoder (510) may perform a decoding operation according to a predetermined video compression technique in a standard such as ITU-T Rec. H.265. The encoded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that the encoded video sequence conforms to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, a profile can select specific tools from all the tools available in the video compression technique or standard as the only tools available for use under that profile. Also, for compliance, the complexity of the encoded video sequence needs to be within the range defined by the level of the video compression technique or standard. In some cases, the level restricts, for example, the maximum picture size, the maximum frame rate, the maximum reconstructed sample rate (measured in, for example, megasamples per second), the maximum reference picture size, etc. The restrictions set by the level may, in some cases, be further restricted through the hypothetical reference decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.
[0051] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence(s). The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0052] FIG. 6 shows a block diagram of a video encoder (603) according to an embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) of the example of FIG. 4.
[0053] The video encoder (603) may receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that can capture video images (s) to be coded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).
[0054] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit,...), any color space (e.g., BT.601 Y CrCB, RGB,...), and any suitable sampling structure (e.g., Y CrCB 4:2:0, Y CrCB 4:4:4). In a media supply system, the video source (601) can be a storage device that stores previously prepared video. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual pictures that give motion when viewed in succession. The pictures themselves can be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. One of ordinary skill in the art can easily understand the relationship between pixels and samples. In the following description, samples are focused on.
[0055] According to one embodiment, the video encoder (603) can code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate coding speed is one function of the controller (650). In some embodiments, the controller (650) controls other functional units and is functionally coupled to other functional units as described below. The couplings are not shown for clarity. Parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, …), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.
[0056] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop may include a source coder (630) (which is responsible for creating symbols such as a symbol stream based on, for example, an input picture to be coded and reference picture(s)) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in the same way as a (remote) decoder also does (since any compression between the symbols and the coded video bitstream is reversible in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream results in a bit-exact result regardless of the location of the decoder (local or remote), the content in the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as the decoder "sees" as reference picture samples when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, for example, due to channel errors) is also used in some related technologies.
[0057] The operation of the "local" decoder (633) can be the same as that of a "remote" decoder such as the video decoder (510) already described in detail above in connection with FIG. 5. However, referring briefly to FIG. 5 as well, since symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy coder (645) and the parser (520) can be reversible, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633).
[0058] The observations that can be made at this point are that any decoder technique, except for syntax analysis / entropy decoding that exists in the decoder, must also exist in the corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder techniques can be omitted since it is the reverse of the decoder techniques described comprehensively. More detailed descriptions are necessary only in certain areas and are provided below.
[0059] During operation, in some examples, the source coder (630) may perform motion-compensated predictive coding that predictively codes an input picture by referring to one or more previously coded pictures from a video sequence designated as a "reference picture". In this way, the coding engine (632) codes the difference between a pixel block of the input picture and a pixel block of a reference picture (s) that can be selected as a predictive reference (s) to the input picture.
[0060] The local video decoder (633) may decode the coded video data of a picture that can be designated as a reference picture based on the symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be an irreversible process. When the coded video data can be decoded in a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that can be performed by the video decoder for the reference picture and store the reconstructed reference picture in the reference picture memory (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture that has the same content as the reconstructed reference picture obtained by a remote video decoder (in the absence of transmission errors).
[0061] Predictor (635) can perform predictive search on the coding engine (632). That is, for a picture to be coded, predictor (635) can search the reference picture memory (634) for sample data (as candidate reference pixel blocks), or specific metadata such as reference picture motion vectors and block shapes, which can function as appropriate predictive references for the new picture. Predictor (635) can operate on a sample block-by-pixel block basis to find an appropriate predictive reference. In some cases, the input picture can have a predictive reference drawn from a plurality of reference pictures stored in the reference picture memory (634) as determined by the search result obtained by predictor (635).
[0062] Controller (650) can manage the coding operations of source coder (630), including, for example, setting parameters and subgroup parameters used for encoding video data.
[0063] The outputs of all the aforementioned functional units can undergo entropy coding in entropy coder (645). Entropy coder (645) converts the symbols generated by various functional units into a coded video sequence by reversibly compressing the symbols according to techniques such as Huffman coding, variable length coding, and arithmetic coding.
[0064] Transmitter (640) can buffer the coded video sequence(s) created by entropy coder (645) for transmission via communication channel (660), which can be a hardware / software link to a storage device where the coded video data will be stored. Transmitter (640) can merge the coded video data from video encoder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0065] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign to each coded picture a specific coded picture type that may affect the coding technique applicable to that picture. For example, a picture may often be assigned as one of the following picture types:
[0066] An intra picture (I picture) may be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, for example, including an instantaneous decoder refresh (「IDR」) picture. Those skilled in the art are aware of these variations of I pictures and their respective uses and characteristics.
[0067] A predictive picture (P picture) may be coded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.
[0068] A bidirectional predictive picture (B picture) may be coded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple-predictive pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0069] The source picture is generally spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be coded block by block. The blocks can be coded predictively by referring to other (already coded) blocks as determined by the coding assignment applied to each picture of the block. For example, blocks of an I picture may be coded non-predictively or may be coded predictively by referring to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be coded predictively via spatial prediction or via temporal prediction by referring to one previously coded reference picture. Blocks of a B picture can be coded predictively via spatial prediction or via temporal prediction by referring to one or two previously coded reference pictures.
[0070] The video encoder (603) can perform a coding operation according to a predetermined video coding technology or standard such as ITU-T Rec. H.265. In that operation, the video encoder (603) can perform various compression operations including a predictive coding operation that utilizes temporal redundancy and spatial redundancy in the input video sequence. Thus, the coded video data can conform to the syntax specified by the video coding technology or standard being used.
[0071] In one embodiment, the transmitter (640) can transmit additional data together with the coded video. The source coder (630) can include such data as part of the coded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.
[0072] Video can be captured as a plurality of source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation in a given picture, and inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a reference picture that has been previously coded and is still buffered in the video, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to the reference block in the reference picture and can have a third dimension that identifies the reference picture when multiple reference pictures are being used.
[0073] In some embodiments, a dual prediction technique can be used in inter-picture prediction. According to the dual prediction technique, two reference pictures such as a first reference picture and a second reference picture that are both earlier than the current picture in the video in decoding order (however, the display order can be past and future respectively) are used. A block in the current picture can be coded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.
[0074] Furthermore, to enhance coding efficiency, a merge mode technique can be used in inter-picture prediction.
[0075] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), namely, one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type for the CU, such as an inter prediction type or an intra prediction type. The CU is divided into one or more prediction units (PUs) according to temporal predictability and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation during coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels.
[0076] FIG. 7 shows a diagram of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) receives a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures and is configured to encode the processing block into an encoded picture that is part of an encoded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) in the example of FIG. 4.
[0077] In an example of HEVC, a video encoder (703) receives a matrix of sample values for a processing block, such as a prediction block of 8×8 samples. The video encoder (703) determines whether the processing block is best coded using an intra mode, an inter mode, or a bi-prediction mode, for example, using rate distortion optimization. If the processing block is to be coded in the intra mode, the video encoder (703) may encode the processing block into the coded picture using an intra prediction technique, and if the processing block is to be coded in the inter mode or the bi-prediction mode, the video encoder (703) may encode the processing block into the coded picture using an inter prediction or a bi-prediction technique, respectively. In certain video coding techniques, the merge mode may be an inter-picture prediction sub-mode in which a motion vector is derived from one or more motion vector predictors without benefit of coded motion vector components external to the predictor. In certain other video coding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.
[0078] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a master controller (721), and an entropy encoder (725) that are coupled to each other as shown in FIG. 7.
[0079] The inter encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks (e.g., blocks in a previous picture and a subsequent picture) in a reference picture, generate inter prediction information (e.g., a description of redundant information by an inter coding technique, a motion vector, merge mode information), and calculate an inter prediction result (e.g., a predicted block) based on the inter prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on coded video information.
[0080] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), optionally compare the block with already coded blocks in the same picture, generate quantized coefficients after transformation, and optionally also generate intra prediction information (e.g., intra prediction direction information by one or more intra coding techniques). In one example, the intra encoder (722) also calculates an intra prediction result (e.g., a predicted block) based on intra prediction information and reference blocks in the same picture.
[0081] The overall controller (721) is configured to determine overall control data and control other components of the video encoder (703) based on the overall control data. In one example, the overall controller (721) determines the mode of a block and provides a control signal to the switch (726) based on this mode. For example, when the mode is the intra mode, the overall controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723), controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream, and when the mode is the inter mode, the overall controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723), controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.
[0082] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) operates based on the residual data and is configured to encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate transform coefficients. The transform coefficients then undergo quantization processing, and the quantized transform coefficients are obtained. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse transformation and generate the decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. In some examples, the decoded block is appropriately processed to generate a decoded picture, and the decoded picture can be buffered in a memory circuit (not shown) and used as a reference picture.
[0083] The entropy encoder (725) is configured to format the bitstream to include the encoded block. The entropy encoder (725) is configured to include various information in accordance with an appropriate standard such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. Note that according to the disclosed subject matter, there is no residual information when coding a block in either the merge submode of the inter mode or the bi-prediction mode.
[0084] FIG. 8 shows a diagram of a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive an encoded picture that is part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In one example, the video decoder (810) is used in place of the video decoder (410) of the example of FIG. 4.
[0085] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) that are coupled to each other as shown in FIG. 8.
[0086] The entropy decoder (871) may be configured to reconstruct from the encoded picture specific symbols that represent the syntax elements that make up the encoded picture. Such symbols can include, for example, the mode in which a block is coded (e.g., intra mode, inter mode, bi-prediction mode, the latter two in the merge sub-mode, or another sub-mode, etc.), prediction information (e.g., intra prediction information or inter prediction information, etc.) that can identify specific samples or metadata used by the intra decoder (872) or the inter decoder (880) for prediction, and residual information in the form of, for example, quantized transform coefficients. In one example, when the prediction mode is inter mode or bi-prediction mode, inter prediction information is provided to the inter decoder (880), and when the prediction type is intra prediction type, intra prediction information is provided to the intra decoder (872). The residual information can undergo inverse quantization and is provided to the residual decoder (873).
[0087] The inter decoder (880) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.
[0088] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0089] The residual decoder (873) is configured to perform inverse quantization to extract the inverse quantized transform coefficients, and process the inverse quantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (since it includes quantization parameter (QP)), and that information may be provided by the entropy decoder (871) (since this may only be low volume control information, the data path is not shown).
[0090] The reconstruction module (874) is configured to combine, in the spatial domain, the residual output by the residual decoder (873) and the prediction result (optionally output by an inter or intra prediction module) to form a reconstructed block that may be part of the reconstructed picture, and the reconstructed picture may be part of the reconstructed video. Note that other appropriate operations, such as a deblocking operation, can be performed to enhance visual quality.
[0091] Note that the video encoders (403), (603) and (703), and the video decoders (410), (510) and (810) may be implemented using any suitable technique. In one embodiment, the video encoders (403), (603) and (703), and the video decoders (410), (510) and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603) and (603), and the video decoders (410), (510) and (810) may be implemented using one or more processors that execute software instructions.
[0092] This disclosure includes improvements to the most probable mode (MPM) list construction.
[0093] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, these two standards organizations jointly formed the JVET (Joint Video Exploration Team) to explore the possibility of developing the next video coding standard beyond HEVC. In April 2018, the JVET officially started the standardization process for the next-generation video coding beyond HEVC. The new standard is named Versatile Video Coding (VVC), and the JVET was renamed the Joint Video Expert Team. In July 2020, H.266 / VVC version 1 was completed. In January 2021, an ad hoc group was established to investigate extended compression beyond VVC capabilities.
[0094] To form the MPM list, in one example, a general MPM list having 22 entries can first be constructed. The first 6 entries in the general MPM list can be included in the primary MPM (PMPM) list, and the remaining entries can form the secondary MPM (SMPM) list. The first entry in the general MPM list can be the planar mode. The remaining entries in the general MPM list can include (i) the intra modes of the left (L), above (A), bottom-left (BL), top-right (AR), and top-left (AL) adjacent blocks, (ii) the directional modes with additional offsets from the first two available directional modes of the adjacent blocks, and (iii) the default mode. The positions of the L, A, BL, AR, and AL adjacent blocks can be shown in FIG. 9.
[0095] When the height of the CU block (e.g., (902) in FIG. 9) is greater than or equal to the width of the CU block, the order of the adjacent blocks can be A, L, BL, AR, and AL; otherwise, the order of the adjacent blocks can be L, A, BL, AR, and AL. Since the index of the entry can be coded in truncated binary, the order of the MPM entries may be important. The latter of the truncated binary may require more bits for coding.
[0096] When the adjacent blocks AL, A, and AR are in a different coding tree unit (CTU) from the current CU (e.g., (902)), the adjacent blocks AL, A, and AR may be considered unavailable due to the line buffer limit. Therefore, the intra-mode of the unavailable adjacent CUs may not be inserted into the MPM list.
[0097] The propagation intra-mode can also be applied to the MPM list construction. For example, in VVC, when the adjacent CU is an inter-coded CU, the intra-mode of the adjacent CU can be considered as the planar mode and inserted into the MPM list.
[0098] The intra mode of an intra-coded CU can be stored in memory in units of 4×4 pixel samples. For an intra-coded CU to which decoder-side intra mode derivation (DIMD) is applied, the decoder-side derived intra mode with the highest gradient histogram (HoG) can be stored in memory. For an intra-coded CU to which template-based intra mode derivation (TIMD) is applied, the decoder-side derived intra mode with the minimum SATD (sum of absolute transformed difference) cost can be stored. For an intra-coded CU to which block-based differential pulse code modulation (BDPCM) is applied, the signaled BDPCM direction can be stored. For an intra-coded CU to which matrix-based intra prediction (MIP) or template matching prediction (TMP) is applied, the planar mode can be stored. For other intra-coded CUs, the intra mode derived from the MPM list or non-MPM list can be stored.
[0099] To increase the accuracy of the MPM list, when an adjacent block is inter-coded, the propagated intra prediction mode can be derived using the motion vector and the reference picture. For an inter-coded CU, the intra mode can be propagated from the referenced area to the inter-coded CU.
[0100] Intra-template matching prediction (also called Intra TMP) is a special intra prediction mode that copies the best predicted block from the reconstructed part of the current frame where the L-shaped template matches the current template (e.g., the template of the current block). For a predefined search range, the encoder searches for the template that is most similar to the current template in the reconstructed part of the current frame and uses the corresponding block as the predicted block, where the most similar template is associated with the corresponding block and the current template is associated with the current block. Then, the encoder signals the use of the Intra TMP mode so that the same prediction operation can be performed on the decoder side. The matching block (or corresponding block) (1002) is shown in FIG. 10 and can function as a matching area for the current CU (1004).
[0101] As shown in FIG. 10, the prediction signal can be generated by matching the L-shaped causal neighbor (or L-shaped template) of the current block (1004) with another block in a predefined search area. Exemplary predefined search areas can include R1 (the current CTU), R2 (the upper left CTU), R3 (the upper CTU), and R4 (the left CTU). The sum of absolute differences (SAD) can be used as the cost function.
[0102] Within each search area, the decoder can search for the template of the block with the minimum SAD with respect to the current template (e.g., the template of the current block (1004)) and use the block with the minimum SAD as the corresponding block of the current block. The corresponding block can further function as the predicted block for the current block (e.g., (1004)).
[0103] The dimensions of all search regions (e.g., SearchRange_w, SearchRange_h) can be set in proportion to the block dimensions of the current block (e.g., BlkW, BlkH). Thus, a certain number of SAD comparisons can be obtained at each pixel. For example, the dimensions of the search region (or search range) can be defined in Equations (1) and (2) as follows: SearchRange_w = a * BlkW Equation (1) SearchRange_h = a * BlkH Equation (2) Here, "a" is a constant that controls the trade-off between gain and the complexity of the search process. In one example, "a" is equal to 5.
[0104] Intra TMP can be enabled for CUs of a specific size. For example, Intra TMP can be enabled for CUs with a width and height size of 64 or less. The maximum CU size for Intra TMP is configurable. The Intra TMP mode can be signaled. For example, the Intra TMP mode can be signaled at the CU level through a dedicated flag.
[0105] In the derived mode (DM) for chroma, a chroma CU can have a collocated luma CU. The intra mode of the collocated luma CU of the chroma CU can be used as the intra mode of the chroma CU.
[0106] When the position (x, y) specifies the top-left sample of the chroma CU for the current picture and the chroma CU includes a width of W and a height of H, the location of the collocated luma CU of the chroma CU can be (x + W / 2, y + H / 2). If the collocated luma CU is not intra-coded, a default value (e.g., DC or planar) can be used as the intra mode of the collocated luma CU.
[0107] If the collocated luma CU is one of specific modes such as coded intra-block copy (IBC) or palette coding (PLT) coding, the intra-mode of the collocated luma CU may be regarded as the DC mode. Otherwise, the intra-mode of the collocated luma CU may be the intra-mode of the CU including the location (x + W / 2, y + H / 2).
[0108] FIG. 11 shows an exemplary derivation mode for chroma including a chroma CU (1104) and a luma area (1102) associated with the chroma CU (1104). As shown in FIG. 11, the chroma CU (1104) can include a corresponding position (1106). The corresponding position (1106) can include a first coordinate value on a first axis (e.g., the X axis) and a second coordinate value on a second axis (e.g., the Y axis) with respect to a reference point on the chroma CU (1104). For example, the reference point can be the lower left corner of the chroma CU (1104). The first axis can be perpendicular to the second axis. In some embodiments, the corresponding position (1106) can have a location of (W / 2, H / 2) with respect to the lower left corner of the chroma CU (1104), where W is the width of the chroma CU (1104) and H is the height of the chroma CU (1104). Thus, the corresponding position (1106) can be the center of the chroma CU (1104). FIG. 11 is only an example, and the corresponding position (1106) can be at any position of the chroma CU (1104), and the reference point can also be at any position of the chroma CU (1104). The collocated luma CU of the chroma CU (1104) can be a luma CU (1108) including the corresponding position (1106) in the luma area (1102).
[0109] Using a universal intra mode map, the intra mode can be memorized in units of samples. Any intra mode can be memorized, including the signaled intra mode, the decoder-derived intra mode, the default intra mode, or the propagated intra mode. The unit of samples can be implicitly predefined or explicitly signaled. For example, the encoder and decoder can implicitly predefine 4×4 pixels as a unit or explicitly signal 2×2 or 8×8 in the bitstream. Note that the universal intra mode map can spread across a CU or CTU. Furthermore, not only can intra or inter CUs memorize the intra mode, but all CUs can memorize the intra mode.
[0110] An example of a partial universal intra mode map is shown in FIGS. 12A and 12B, where each square can represent a 4×4 unit of samples.
[0111] In one embodiment, the universal intra mode map can initially be empty. Furthermore, the universal intra mode map can memorize or otherwise include the signaled intra mode, the decoder-derived intra mode, and / or the propagated intra mode, which are obtained during the decoding process.
[0112] In another embodiment, the default intra mode can be used to initially initialize the universal intra mode map. Further, the default intra mode can be replaced by the signaled intra mode, the decoder-derived intra mode, or the propagated intra mode obtained during the decoding process. For example, as shown in FIG. 12A, the default intra mode 0 (or planar) can be stored in all units of the universal intra mode map (1202). Thereafter, the default intra mode 0 can be replaced with another intra mode. For example, the other intra mode can be the signaled intra mode, the decoder-derived intra mode, or the propagated intra mode.
[0113] In yet another embodiment, the universal intra mode map can initially be empty. Further, some of the empty (or blank) units of the universal intra mode map can store the signaled intra mode, the decoder-derived intra mode, and / or the propagated intra mode achieved during the decoding process. After the CU is decoded, the remaining blank units can be filled with a default intra mode such as intra mode 0 (or planar).
[0114] As shown in FIG. 12B, the universal intra mode map (1204) can initially be empty, and some of the empty units can be filled with intra modes during the decoding process. After the CU is decoded, the blank units (e.g., (1206) and (1208)) can be filled with a default intra mode such as intra mode 0 (or planar).
[0115] The propagated intra mode can include the propagated intra mode for non-intra-coded CUs, thereby diversifying the possible intra modes that non-intra-coded CUs can have. In the present disclosure, improvements to intra mode propagation are provided, including improvements to propagation for the TMP mode and the derived mode of chroma.
[0116] In one embodiment, the intra mode of the matching area can be propagated to the current block for the TMP mode. The matching area of the current block (or current CU) can be the matched predictor found in the search area using an L-shaped template. Templates having other shapes can also be applied. The search area can be the reconstructed area. The search area and the current block can be included in the same frame or the same picture. In some embodiments, the L-shaped templates of the matching area and the current block can have the minimum cost value (e.g., SAD) within the search area. The L-shaped template of the matching area can include adjacent pixels such as adjacent to the left side and the upper side of the matching area. The L-shaped template of the current CU can include adjacent pixels such as adjacent to the left side and the upper side of the current CU.
[0117] The propagation intra mode can be derived using the corresponding position of the current CU (e.g., correspondPosition(a, b)). The corresponding position can be any position within the current CU. For example, the corresponding position can be located at the center of the current CU. The corresponding position can include a first coordinate value (e.g., a) on a first axis (e.g., the X-axis) and a second coordinate value (e.g., b) on a second axis (e.g., the Y-axis) with respect to a reference point such as the lower left corner of the current block, and the first axis can be perpendicular to the second axis.
[0118] The corresponding position is predefined and can be implicitly agreed upon on both the encoder and the decoder. Alternatively, the corresponding position can be signaled explicitly or implicitly in the bitstream within or outside the band.
[0119] In one example, a CU including a corresponding position (e.g., correspondPosition(a,b)) within a matching area can be the corresponding CU of the current block regardless of how the corresponding CU is coded. Therefore, the corresponding position within the matching area can include a first coordinate value (e.g., a) on a first axis (e.g., the X axis) with respect to a reference point such as the lower left corner of the matching area, and a second coordinate value (e.g., b) on a second axis (e.g., the Y axis).
[0120] In the present disclosure, any type or combination of the intra modes of the corresponding CU (e.g., normal intra mode, propagation intra mode, default intra mode, etc.) can be propagated to the current CU based on the TMP mode. The intra mode propagated to the current block can be further stored and propagated to the next CU.
[0121] FIG. 13 shows an exemplary corresponding CU (1308) of the current block (1304). As shown in FIG. 13, the corresponding CU (1308) can be included in the matching area (1306). The matching area (1306) can be a reconstructed area within the picture (1302). The current block 1304 can also be included in the picture (1302). The matching area (1306) can have a template such as an L-shaped template (1316). As shown in FIG. 13, the L-shaped template (1316) can include areas adjacent to the left side and the upper side of the matching area. The current CU (1304) can have a template such as an L-shaped template (1318). The L-shaped template (1318) can include areas adjacent to the left side and the upper side of the current CU. The L-shaped template of the matching area and the L-shaped template of the current block can have a minimum cost value (e.g., SAD) in the search area of the picture (1302). FIG. 13 is only an example. The template (1316) and the template (1318) can each include areas having other shapes adjacent to the matching area (1306) and the current CU (1304), respectively.
[0122] In one example, the current CU (1304) has a corresponding position (1312) for any location of the current CU (1304), and can have a location of (a, b), where a is the first coordinate value on the first axis (e.g., X) with respect to a reference point on the current CU (1304), and b is the second coordinate value on the second axis (e.g., Y). The reference point can be located at any position of the current CU (1304). For example, the reference point can be the lower left corner (1314) of the current CU (1304). The corresponding position can be located at any position of the current CU (1304). For example, the corresponding position (1312) of the current CU (1304) can be located at the center of the current CU (1304). Therefore, the corresponding position (1312) can have a location of (W / 2, H / 2) with respect to the lower left corner of the current CU (1304), where W and H are the width and height of the current CU (1304), respectively. The CU (e.g., (1308)) including the corresponding position (1310) in the matching area (1306) can be denoted as the corresponding CU (1308) of the current CU (1304). The corresponding position (1310) of the matching area (1306) can also have a location of (a, b) on the first axis and the second axis with respect to a reference point (e.g., the lower left corner (1320)) on the matching area (1306). Therefore, when the corresponding position (1312) has a location of (W / 2, H / 2) with respect to the left corner of the current CU (1304), the corresponding position (1310) can have a location of (W / 2, H / 2) with respect to the lower left corner of the matching area (1306).
[0123] In another example, when the corresponding position (1312) has a location of (W / 2, H / 2) with respect to a reference point on the current CU (1304), such as the left corner (1314) of the current CU (1304), the corresponding position (1310) can have a location of (W' / 2, H' / 2) with respect to a reference point on the matching area (1306), such as the lower left corner (1320) of the matching area (1306), where W' and H' are the width and height of the matching area (1306), respectively.
[0124] In yet another example, the corresponding position (1312) can have a location of (W / a, H / b) with respect to a reference point such as the left corner (1314) of the current CU (1304), and the corresponding position (1310) can have a location of (W’ / a, H’ / b) with respect to a reference point on the matching area (1306) such as the lower left corner (1320) of the matching area (1306), where a and b are positive integers. In some embodiments, W is not equal to W’ and H is not equal to H’. In some embodiments, W may be equal to W’ and H may be equal to H’. Thus, the size of the matching area (1306) is equal to the size of the current CU (1304).
[0125] The intra mode of the corresponding CU (1308) can be used as the intra mode of the current CU (1304) based on the TMP mode. The intra mode is further stored in units of 4×4 pixel samples and can be propagated to the next CU.
[0126] In another embodiment, the intra mode of the current CU can be achieved by accessing an entry (or the assigned intra mode) at the corresponding position in the universal intra mode map. The intra mode achieved through the universal intra mode map can be further stored and propagated to the next CU.
[0127] For example, as shown in FIG. 14, a picture (1402) can include a current CU (1404) and a matching area (1406) of the current block (1404). The current CU (1404) can include a corresponding position (1414) and a corresponding CU (1408) within the matching area (1406). The corresponding CU (1408) can include a corresponding position (1412). The corresponding position (1414) can be located at any location of the current CU (1404). The corresponding position (1414) can have a location of (a, b), where a is a first coordinate value on a first axis (e.g., X) with respect to a reference point on the current CU (1404), and b is a second coordinate value on a second axis (e.g., Y). The reference point on the current CU (1404) can be located at any location of the current CU (1404), such as the lower left corner (1461) of the current CU (1404).
[0128] In one example, the corresponding position (1414) can be located at the center of the current CU (1404). Therefore, the corresponding position (1414) can have a location of (W / 2, H / 2), where W and H are the width and height of the current CU (1404), respectively. The corresponding position (1412) can also have a location of (a, b) with respect to a reference point on the matching area (1406). The reference point on the matching area (1406) can be located at any location of the matching area (1406), such as the lower left corner (1418) of the matching area (1406). Therefore, when the corresponding position (1414) has a location of (W / 2, H / 2) with respect to the lower left corner (1416) of the current CU (1404), the corresponding position (1412) can have a location of (W / 2, H / 2) with respect to the lower left corner (1418) of the matching area (1406).
[0129] In another example, the corresponding position (1414) can have a location of (W / a, H / b) with respect to the left corner (1416) of the current CU (1404), and the corresponding position (1412) can have a location of (W’ / a, H’ / b) with respect to the lower left corner (1418) of the matching area (1406). W and H are the width and height of the current CU (1404), respectively. W’ and H’ are the width and height of the matching area (1406), respectively. In some embodiments, W is not equal to W’, and H is not equal to H’. In some embodiments, W may be equal to W’, and H may be equal to H’.
[0130] Furthermore, the universal map (1410) can be assigned to the matching area (1406). The universal map (1410) can divide the matching area (1406) into a plurality of sub-areas. A specific intra-mode can be assigned to each of the plurality of sub-areas. Therefore, the intra-mode (e.g., intra-mode (51)) assigned to the sub-area including the corresponding position (1412) can be propagated to the current CU (1404).
[0131] In the present disclosure, the current chroma CU can have a collocated luma CU. The collocated luma CU can be associated with the current chroma CU. For example, both the collocated luma CU and the luma CU can have corresponding positions (e.g., correspondPositionDM(a,b)) and can be included in a coding tree unit (CTU). In some embodiments, the current chroma CU can be aligned with the collocated luma CU. According to the chroma DM mode, the intra mode of the collocated luma CU of the current chroma CU can be propagated to the current chroma CU. When the chroma DM mode is used, the collocated luma CU can be coded using an intra prediction mode, and other prediction modes such as intra block copy (IBC), inter mode, and PLT. Accordingly, the corresponding position (e.g., correspondPositionDM(a,b)) is applied to derive the propagated intra mode from the collocated luma CU to the current chroma CU.
[0132] The corresponding position is predefined and can be implicitly agreed upon on both the encoder and the decoder. Alternatively, the collocated position (or corresponding position) can be explicitly or implicitly signaled in the bitstream, either in-band or out-of-band. Further, the corresponding position can be located at any location of the chroma CU and can include the location (a,b), where a is the first coordinate value on the first axis (e.g., X) with respect to the reference point of the current chroma CU, such as the lower left corner of the current chroma CU, and b is the second coordinate value on the second axis (e.g., Y).
[0133] In one embodiment, according to the chroma DM mode, any type of intra mode associated with the collocated luma CU, such as a normal intra mode, a propagated intra mode, a default intra mode, etc., can be propagated to the current chroma CU.
[0134] Furthermore, when the collocated luma CU is coded or inter-coded in a specific mode such as IBC, instead of using the default intra-mode value of the collocated luma CU, the propagated intra-mode of the collocated luma CU can be propagated to the current chroma CU. The propagated intra-mode of the collocated luma CU can be obtained based on the intra-mode of the adjacent luma CU of the collocated luma CU. In such a case, the current chroma CU can have a value other than the default intra-mode value from the IBC or inter-coded collocated luma CU.
[0135] In another embodiment, the intra-mode of the current chroma CU can be achieved by accessing the entry (or intra-mode) of the universal intra-mode map assigned to the collocated luma CU according to the position of the collocated luma CU in the universal intra-mode map. The universal intra-mode map can divide the collocated luma CU into a plurality of sub-areas. Each of the plurality of sub-areas can be assigned a respective intra-mode according to the universal intra-mode map. The intra-mode assigned to the sub-area including the corresponding position can be propagated to the current chroma CU.
[0136] FIG. 15 shows a flowchart outlining a first exemplary decoding process (1500) according to some embodiments of the present disclosure. FIG. 16 shows a flowchart outlining a second exemplary decoding process (1600) according to some embodiments of the present disclosure. FIG. 17 shows a flowchart outlining a first exemplary encoding process (1700) according to some embodiments of the present disclosure. FIG. 18 shows a flowchart outlining a second exemplary encoding process (1800) according to some embodiments of the present disclosure. The proposed processes may be used separately or combined in any order. Further, each of the processes (or embodiments), encoders, and decoders may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.
[0137] In embodiments, any operations of the processes (e.g., (1500), (1600), (1700), and (1800)) may be combined or arranged in any amount or order as desired. In embodiments, two or more of the operations of the processes (e.g., (1500), (1600), (1700), and (1800)) may be executed in parallel.
[0138] Processes (e.g., (1500), (1600), (1700), and (1800)) can be used in block reconstruction and / or encoding to generate prediction blocks for blocks being reconstructed. In various embodiments, processes (e.g., (1500), (1600), (1700), and (1800)) are executed by processing circuits in terminal devices (210), (220), (230), and (240), processing circuits that execute the functions of video encoder (403), processing circuits that execute the functions of video decoder (410), processing circuits that execute the functions of video decoder (510), processing circuits that execute the functions of video encoder (603), etc. In some embodiments, processes (e.g., (1500), (1600), (1700), and (1800)) are implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes processes (e.g., (1500), (1600), (1700), and (1800)).
[0139] As shown in FIG. 15, process (1500) starts from (S1501) and can proceed to (S1510). In (S1510), (i) the current block of the picture and (ii) the coding information of the reconstructed area of the picture can be received from the coded video bitstream.
[0140] (S1520), the matching area of the current block in the reconstructed area of the picture can be determined.
[0141] (S1530), the first corresponding position of the current block can be determined within the current block. The first corresponding position can include a first coordinate value on a first axis and a second coordinate value on a second axis with respect to a first reference point on the current block. The first axis can be perpendicular to the second axis.
[0142] In (S1540), the corresponding block of the current block can be determined in the matching area. The corresponding block can include a second corresponding position having a first coordinate value on a first axis and a second coordinate value on a second axis with respect to a second reference point on the matching area.
[0143] In (S1550), the intra prediction mode of the current block can be determined based on the corresponding block.
[0144] To determine the matching area, a plurality of candidate matching areas can be searched in the reconstruction area of the picture. A respective cost value between the template area of the current block and the respective template area of each of the plurality of candidate matching areas can be determined. The matching area can be determined as the candidate matching area having the minimum cost value among the plurality of candidate matching areas. The template area of the current block can include a first area adjacent to the left side of the current block and a second area adjacent to the upper side of the current block. The respective template area of each of the plurality of candidate matching areas can include a first area adjacent to the left side of each one of the plurality of candidate matching areas and a second area adjacent to the upper side of each one of the plurality of candidate matching areas.
[0145] In some embodiments, the first corresponding position of the current block can be predefined or signaled in the coded information.
[0146] In some embodiments, the first corresponding position can be the center of the current block.
[0147] In some embodiments, the intra prediction mode of the corresponding block can be determined as the intra prediction mode of the current block.
[0148] In some embodiments, the intra prediction mode of the current block may be determined based on a universal intra mode map. The universal intra mode map may divide the matching area into a plurality of sub-areas. Each of the plurality of sub-areas may be associated with a respective intra prediction mode. The intra prediction mode of the current block may be the intra prediction mode associated with the sub-area that includes the second corresponding position among the plurality of sub-areas.
[0149] As shown in FIG. 16, process (1600) may start from (S1601) and proceed to (S1610). In (S1610), (i) a chroma coding unit (CU) and (ii) coding information of a luma area may be received from a coded video bitstream.
[0150] In (S1620), a corresponding position of the chroma CU may be determined within the chroma CU. The corresponding position may include a first coordinate value on a first axis and a second coordinate value on a second axis. The first axis may be perpendicular to the second axis.
[0151] In (S1630), a collocated luma CU of the chroma CU may be determined in the luma area. The collocated luma CU may include the corresponding position.
[0152] In (S1640), the intra prediction mode of the chroma CU may be determined based on the collocated luma CU.
[0153] In one embodiment, in response to the collocated luma CU being intra-coded based on an intra prediction mode, the intra prediction mode of the collocated luma CU may be determined as the intra prediction mode of the chroma CU. In other embodiments, in response to the collocated luma CU not being intra-coded, the propagation intra mode of the collocated luma CU may be determined as the intra prediction mode of the chroma CU. The propagation intra mode of the collocated luma CU may be obtained based on the intra prediction mode of an adjacent luma CU of the collocated luma CU.
[0154] In some embodiments, the corresponding position of the chroma CU may be predefined or signaled in the coding information.
[0155] In some embodiments, the intra prediction mode of the chroma CU may be determined based on a universal intra mode map. The universal intra mode map may divide the luma area into a plurality of sub-areas. Each of the plurality of sub-areas may be assigned a respective intra prediction mode. The intra prediction mode of the chroma CU may be the intra prediction mode associated with the sub-area including the corresponding position among the plurality of sub-areas.
[0156] As shown in FIG. 17, the process (1700) starts from (S1701) and can proceed to (S1710). In (S1710), the matching area of the current block of the picture may be determined from the reconstructed area of the picture.
[0157] In (S1720), the first corresponding position of the current block may be determined within the current block. The first corresponding position may include a first coordinate value on a first axis and a second coordinate value on a second axis with respect to a first reference point on the current block, and the first axis may be perpendicular to the second axis.
[0158] In (S1730), the corresponding block of the current block can be determined from the matching area, and the corresponding block can include a second corresponding position having a first coordinate value on a first axis and a second coordinate value on a second axis with respect to a second reference point on the matching area.
[0159] In (S1740), the intra prediction mode of the current block can be determined based on the corresponding block.
[0160] In (S1750), intra prediction can be performed on the current block based on the determined intra prediction mode.
[0161] As shown in FIG. 18, the process (1800) starts from (S1801) and can proceed to (S1810). In (S1810), the corresponding position of the chroma CU can be determined within the chroma CU. The corresponding position can include a first coordinate value on a first axis and a second coordinate value on a second axis, and the first axis can be perpendicular to the second axis.
[0162] In (S1820), the collocated luma CU of the chroma CU can be determined in the luma area, and the collocated luma CU can include the corresponding position.
[0163] In (S1830), the intra prediction mode of the chroma CU can be determined based on the collocated luma CU.
[0164] In (S1840), intra prediction can be performed on the chroma CU based on the determined intra prediction mode.
[0165] In (S1850), a coded video bitstream can be generated to include (i) the chroma CU and (ii) the coding information of the luma area.
[0166] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 19 shows a computer system (1900) suitable for implementing certain embodiments of the disclosed subject matter.
[0167] Computer software can be coded using any suitable machine code or computer language that can follow assembly, compilation, linking, or similar mechanisms to create code that includes instructions that can be executed directly, or through interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), and the like.
[0168] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet-of-things devices, and the like.
[0169] The components shown in FIG. 19 with respect to computer system (1900) are exemplary in nature and are not intended to suggest any limitation with respect to the use or functionality of the computer software implementing embodiments of the present disclosure. Also, the configuration of the components should not be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiments of computer system (1900).
[0170] The computer system (1900) may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users through, for example, tactile input (keystrokes, swipes, movements of a data glove, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), and olfactory input (not shown). The human interface device may also be used to capture certain media that is not necessarily directly related to conscious human input, such as audio (speech, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), video (2D video, 3D video including stereoscopic video, etc.).
[0171] The input human interface device may include one or more of a keyboard (1901), a mouse (1902), a trackpad (1903), a touch screen (1910), a data glove (not shown), a joystick (1905), a microphone (1906), a scanner (1907), and a camera (1908) (only one of each is shown).
[0172] The computer system (1900) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen (1910), a data glove (not shown), or a joystick (1905), although there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers (1909), headphones (not shown), etc.), visual output devices (screens (1910) including CRT screens, LCD screens, plasma screens, OLED screens, each of which may or may not have a touch screen input function, each of which may or may not have a tactile feedback function, and some of which may be capable of outputting 2D visual output or output beyond 3D by means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and may include a printer (not shown).
[0173] The computer system (1900) may also include a memory device accessible by humans and optical media including a CD / DVD ROM / RW (1920) having a CD / DVD or similar medium (1921), a thumb drive (1922), a removable hard drive or solid state drive (1923), legacy magnetic media such as tapes and floppy (registered trademark) disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), etc., and their associated media.
[0174] One skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include a transmission medium, a carrier wave, or other transient signals.
[0175] The computer system (1900) can also include an interface (1954) to one or more communication networks (1955). The network can be, for example, wireless, wireline, optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet (registered trademark), wireless LAN, cellular networks including GSM (registered trademark), 3G, 4G, 5G, LTE, etc., TV wireline or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicle and industrial including CANBus, etc. A particular network generally requires an external network interface adapter attached to a particular general-purpose data port or peripheral bus (1949) (e.g., a USB port of the computer system (1900)), and others are generally integrated into the core of the computer system (1900) by attaching to the system bus as described below (e.g., an Ethernet (registered trademark) interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1900) can communicate with other entities. Such communication can be unidirectional receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., CANbus to a particular CANbus device), or bidirectional to other computer systems using, for example, local or wide area digital networks. Specific protocols and protocol stacks can be used for each of these networks and network interfaces as described above.
[0176] The aforementioned human interface device, human-accessible memory device, and network interface can be attached to the core (1940) of the computer system (1900).
[0177] The core (1940) can include one or more central processing units (CPUs) (1941), a graphics processing unit (GPU) (1942), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (1943), a hardware accelerator (1944) for specific tasks, a graphics adapter (1950), etc. These devices can be connected through a system bus (1948) together with a read-only memory (ROM) (1945), a random access memory (1946), an internal hard drive that is not accessible to the user, an internal mass storage device such as an SSD (1947). In some computer systems, the system bus (1948) can be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be connected directly to the system bus (1948) of the core or through a peripheral bus (1949). In one example, a screen (1910) can be connected to a graphics adapter (1950). Architectures of peripheral buses include PCI, USB, etc.
[0178] The CPU (1941), GPU (1942), FPGA (1943), and accelerator (1944) can execute specific instructions that can together constitute the aforementioned computer code. That computer code can be stored in the ROM (1945) or RAM (1946). Transient data can also be stored in the RAM (1946), and persistent data can be stored, for example, in the internal mass storage device (1947). Fast storage and retrieval to / from any of the memory devices can be enabled by the use of cache memory, which can be closely associated with one or more CPUs (1941), GPUs (1942), mass storage devices (1947), ROM (1945), RAM (1946), etc.
[0179] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or they may be of the kind well-known and available to those of ordinary skill in the computer software arts.
[0180] As an example, rather than a limitation, a computer system (1900) having an architecture, specifically a core (1940), can provide functionality as a result of software embodied on one or more tangible computer-readable media being executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.). Such computer-readable media can be associated with media of a user-accessible mass storage device as introduced above, as well as specific storage devices of the core (1940) that are non-transitory in nature, such as a core internal mass storage device (1947) or a ROM (1945). The software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (1940). The computer-readable media can include one or more memory devices or chips according to specific needs. The software can cause the core (1940) and specifically the processors (including CPUs, GPUs, FPGAs, etc.) therein to define data structures stored in a RAM (1946) and modify such data structures according to processes defined by the software, thereby executing specific processes or specific portions of specific processes described herein. Additionally or alternatively, the computer system can provide functionality as a result of logic hardwired or otherwise embodied within a circuit (e.g., an accelerator (1944)), which can operate instead of or in conjunction with the software to execute specific processes or specific portions of specific processes described herein. References to software can, as necessary, include logic, and vice versa. References to computer-readable media can, as necessary, include circuits (such as integrated circuits (ICs)) that store software for execution, circuits that embody logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software. [Appendix A: Acronyms] JEM: joint exploration model VVC: Versatile Video Coding (General-Purpose Video Coding) BMS: Benchmark Set (Benchmark Set) MV: Motion Vector (Motion Vector) HEVC: High Efficiency Video Coding (High-Efficiency Video Coding) SEI: Supplementary Enhancement Information (Supplementary Enhancement Information) VUI: Video Usability Information (Video Usability Information) GOPs: Groups of Pictures (Picture Groups) TUs: Transform Units (Transform Units) PUs: Prediction Units (Prediction Units) CTUs: Coding Tree Units (Coding Tree Units) CTB: Coding Tree Blocks (Coding Tree Blocks) PB: Prediction Blocks (Prediction Blocks) HRD: Hypothetical Reference Decoder (Hypothetical Reference Decoder) SNR: Signal Noise Ratio (Signal-to-Noise Ratio) CPUs: Central Processing Units (Central Processing Units) GPUs: Graphics Processing Units (Graphics Processing Units) CRT: Cathode Ray Tube (Cathode Ray Tube) LCD: Liquid-Crystal Display (Liquid Crystal Display) OLED: Organic Light-Emitting Diode (Organic Light-Emitting Diode) CD: Compact Disc (Compact Disc) DVD: Digital Video Disc (Digital Video Disk) ROM: Read-Only Memory (Read-Only Memory) RAM: Random Access Memory (Random Access Memory) ASIC: Application-Specific Integrated Circuit (Application-Specific Integrated Circuit) PLD: Programmable Logic Device (Programmable Logic Device) LAN: Local Area Network (Local Area Network) GSM: Global System for Mobile communications (Global System for Mobile Communications) LTE: Long-Term Evolution (Long-Term Evolution) CANBus: Controller Area Network Bus (Controller Area Network Bus) USB: Universal Serial Bus (Universal Serial Bus) PCI: Peripheral Component Interconnect (Peripheral Component Interconnect) FPGA: Field Programmable Gate Areas (Field Programmable Gate Areas) SSD: solid-state drive (Solid-State Drive) IC: Integrated Circuit (Integrated Circuit) CU: Coding Unit (Coding Unit)
[0181] In this disclosure, several exemplary embodiments have been described, but there are modifications, substitutions, and various alternative equivalents that are included within the scope of this disclosure. Thus, those skilled in the art will understand that, although not explicitly illustrated or described herein, many systems and methods that embody the principles of this disclosure and are thus within the spirit and scope of this disclosure can be devised.
Claims
1. A method for video decoding executed in a video decoder, comprising: (i) receiving coded information of a current block of a picture and (ii) a reconstruction area of the picture from a coded video bitstream; determining a matching area of the current block in the reconstruction area of the picture; searching for a plurality of candidate matching areas in the reconstruction area of the picture; determining respective cost values between a template area of the current block and respective template areas of each of the plurality of candidate matching areas; determining the matching area as a candidate matching area having a minimum cost value among the plurality of candidate matching areas; including the steps of; determining a first corresponding position of the current block within the current block, the first corresponding position including a first coordinate value on a first axis and a second coordinate value on a second axis with respect to a first reference point on the current block, the first axis being perpendicular to the second axis; determining a corresponding block of the current block in the matching area, the corresponding block including a second corresponding position having the first coordinate value on the first axis and the second coordinate value on the second axis with respect to a second reference point on the matching area; determining an intra prediction mode of the current block based on the corresponding block and a universal intra mode map, the universal intra mode map dividing the matching area into a plurality of sub-areas, each of the plurality of sub-areas being associated with a respective intra prediction mode, the intra prediction mode of the current block being the intra prediction mode associated with the sub-area including the second corresponding position among the plurality of sub-areas; A method including the above steps.
2. The template area of the current block includes a first area adjacent to the left side of the current block and a second area adjacent to the upper side of the current block. For each template area of each of the plurality of candidate matching areas, each template area of each of the plurality of candidate matching areas includes a first area adjacent to the left of each one of the plurality of candidate matching areas and a second area adjacent to the upper side of each one of the plurality of candidate matching areas. The method according to claim 1.
3. The method according to claim 1, wherein the first corresponding position of the current block is predefined or signaled in the coded information.
4. The method according to claim 1, wherein the first corresponding position is the center of the current block.
5. The step of determining the intra prediction mode of the current block includes further the step of determining the intra prediction mode of the corresponding block as the intra prediction mode of the current block The method according to claim 1.
6. An apparatus comprising a processing circuit, wherein the processing circuit is configured to execute the method according to any one of claims 1 to 5.