Video encoding and decoding method, device, system and computer-readable medium
By comparing the nominal angles of the current block and adjacent blocks, using the cumulative density function CDF signaling and mapping table, the problem of low angle difference processing efficiency in intra prediction is solved, and the video encoding and decoding efficiency and quality are improved.
Patent Information
- Application Number
- CN202180006188.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-07-30
- Filing Date
- 2021-08-18
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-08-18
AI Technical Summary
Existing video encoding and decoding technologies are difficult to efficiently handle angle differences in intra prediction, resulting in inefficient encoding efficiency.
By obtaining the nominal angles of the current block and adjacent blocks, comparing the absolute difference to determine the allowed incremental angle, using the cumulative density function CDF signaling and mapping table to represent the incremental angle, achieving efficient intra prediction.
It improves the efficiency and quality of video encoding and decoding, reduces the amount of data, and is suitable for a variety of video transmission and storage applications.
Smart Images

Figure CN114641988B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Application No. 63 / 067,791, filed on August 19, 2020, and U.S. Application No. 17 / 390,256, filed on July 30, 2021, the contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] The embodiments of the present disclosure relate to a set of advanced video encoding technologies, and more particularly, to a method, apparatus, system, and computer-readable medium for video encoding and decoding. Background Art
[0004] AOMedia Video 1 (AV1) is an open video codec designed for video transmission over the internet. It is the successor to VP9, developed by the Alliance for Open Media (AOMedia), a consortium of semiconductor companies, video-on-demand providers, video content producers, software developers, and web browser vendors, founded in 2015. Many components of the AV1 project stem from previous research work by alliance members. Independent contributors began working on experimental technology platforms several years ago: Xiph / Mozilla's Daala released code in 2010, Google's experimental VP9 evolution project, VP10, was announced on September 12, 2014, and Cisco's Thor was announced on August 11, 2015. Building on the VP9 base code, AV1 incorporates additional technologies, several of which were developed in these experimental forms. The first version of the AV1 reference codec (version 0.1.0) was released on April 7, 2016. The Alliance announced the release of the AV1 bitstream specification, along with a software-based reference encoder and decoder, on March 28, 2018. On June 25, 2018, a validated version 1.0.0 of the specification was released. On January 8, 2019, the "AV1 Bitstream and Decoding Process Specification," a validated version 1.0.0 with Errata 1 to that specification, was released. The AV1 bitstream specification includes the reference video codec. The "AV1 Bitstream and Decoding Process Specification (Version 1.0.0, Errata 1)" of the Alliance for Open Media (January 8, 2019) is incorporated herein by reference in its entirety. Summary of the Invention
[0005] According to one or more embodiments, a method for video encoding and decoding is provided, the method being performed by at least one processor. The method includes: receiving an encoded picture; and decoding the encoded picture. The decoding includes: obtaining a nominal angle of a current block of the encoded picture, the nominal angle of the current block being used for intra-frame prediction; obtaining a nominal angle of at least one neighboring block of the current block, the nominal angle of the at least one neighboring block being used for intra-frame prediction; determining, based on a comparison between the nominal angle of the current block and the nominal angle of the at least one neighboring block, whether to signal all allowed incremental angles of the nominal angle of the current block or only a subset of the allowed incremental angles; based on the determination, signaling all allowed incremental angles or the subset of allowed incremental angles; and predicting the current block based on the signaling of all allowed incremental angles or the subset of allowed incremental angles.
[0006] According to an embodiment, the determining comprises determining to signal all of the allowed incremental angles based on an absolute difference between a first value corresponding to the nominal angle of the current block and a second value corresponding to the nominal angle of the at least one neighboring block being less than or equal to a threshold.
[0007] According to an embodiment, the threshold is 2.
[0008] According to an embodiment, the determining comprises determining to signal only the subset of the allowed incremental angles based on an absolute difference between a first value corresponding to the nominal angle of the current block and a second value corresponding to the nominal angle of the at least one neighboring block being greater than a threshold.
[0009] According to an embodiment, the number of allowed incremental angles in the subset is determined based on the absolute difference.
[0010] According to an embodiment, said comparison is performed between said nominal angle of said current block and said nominal angles of only a predetermined number of neighboring blocks.
[0011] According to an embodiment, said signaling all said allowed incremental angles or said subset of said allowed incremental angles comprises: based on said determining, signaling at least one of said allowed incremental angles by using a cumulative density function (CDF).
[0012] According to an embodiment, the decoding further comprises: signaling an index; and identifying the incremental angle of the current block by using a mapping table, wherein the index is mapped to the incremental angle of the current block in the mapping table.
[0013] According to an embodiment, the decoding further comprises: signaling an index; and identifying an incremental angle of the current block by using a mapping table, wherein in the mapping table, the index is mapped to the incremental angle of the current block and the incremental angle of the at least one neighboring block.
[0014] According to an embodiment, the current block is a luma block.
[0015] According to one or more embodiments, a system for video encoding and decoding is provided, comprising: at least one memory configured to store computer program code; and at least one processor configured to access the computer program code and operate according to instructions of the computer program code to perform a video encoding and decoding method according to an embodiment of the present application.
[0016] According to one or more embodiments, a video encoding and decoding apparatus is provided, comprising: a decoding module for decoding a received encoded picture, the decoding module comprising: a first acquisition module for acquiring a nominal angle of a current block of the encoded picture, wherein the nominal angle of the current block is used for intra-frame prediction; a second acquisition module for acquiring a nominal angle of at least one neighboring block of the current block, wherein the nominal angle of the at least one neighboring block is used for intra-frame prediction; a determination module for determining, based on a comparison between the nominal angle of the current block and the nominal angle of the at least one neighboring block, whether to signal all allowed incremental angles of the nominal angle of the current block or only a subset of the allowed incremental angles; a signaling module for signaling all allowed incremental angles or the subset of the allowed incremental angles based on the determination; and a prediction module for predicting the current block based on the signaling of all allowed incremental angles or the subset of the allowed incremental angles.
[0017] According to an embodiment, the determination module is configured to cause at least one processor to: determine to signal all the allowed incremental angles based on an absolute difference between a first value and a second value being less than or equal to a threshold, wherein the first value corresponds to the nominal angle of the current block and the second value corresponds to the nominal angle of the at least one neighboring block.
[0018] According to an embodiment, the threshold is 2.
[0019] According to an embodiment, the determination module is further used to: determine to signal only the subset of the allowed incremental angles based on an absolute difference between a first value and a second value being greater than a threshold, wherein the first value corresponds to the nominal angle of the current block and the second value corresponds to the nominal angle of the at least one neighboring block.
[0020] According to an embodiment, the number of allowed incremental angles in the subset is determined based on the absolute difference.
[0021] According to an embodiment, said comparison is performed between said nominal angle of said current block and said nominal angles of only a predetermined number of neighboring blocks.
[0022] According to an embodiment, the signaling module is configured to signal at least one of the allowed incremental angles by using a cumulative density function (CDF) based on the determination.
[0023] According to an embodiment, the decoding module further includes: an index signaling module that represents an index with a signal; and an identification module that identifies the incremental angle of the current block by using a mapping table, wherein the index is mapped to the incremental angle of the current block in the mapping table.
[0024] According to an embodiment, the decoding module further includes: index signaling, which represents the index with a signal; and an identification module, which identifies the incremental angle of the current block by using a mapping table, wherein in the mapping table, the index is mapped to the incremental angle of the current block and the incremental angle of the at least one adjacent block.
[0025] According to one or more embodiments, a non-transitory computer-readable medium is provided, storing computer instructions, wherein when executed by at least one processor, the computer instructions are configured to cause the at least one processor to perform a video encoding and decoding method according to an embodiment of the present application.
[0026] BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Other features, properties, and various advantages of the disclosed subject matter will become further apparent from the following detailed description and accompanying drawings, in which:
[0028] Figure 1 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.
[0029] Figure 2 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.
[0030] Figure 3 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment.
[0031] Figure 4 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment.
[0032] Figure 5A FIG. 1 is a diagram of a first example partition structure of VP9.
[0033] Figure 5BFIG. 4 is a diagram of a second example partition structure of VP9. FIG.
[0034] Figure 5C FIG. 4 is a diagram illustrating a third exemplary partition structure of VP9.
[0035] Figure 5D FIG. 4 is a diagram illustrating a fourth example partition structure of VP9.
[0036] Figure 6A Schematic diagram of the first example partition structure of AV1.
[0037] Figure 6B Schematic diagram of a second example partition structure for AV1.
[0038] Figure 6C Schematic diagram of the third example partition structure of AV1.
[0039] Figure 6D Schematic diagram of the fourth example partition structure of AV1.
[0040] Figure 6E Schematic diagram of the fifth example partition structure of AV1.
[0041] Figure 6F Schematic diagram of the sixth example partition structure of AV1.
[0042] Figure 6G Schematic diagram of the seventh example partition structure of AV1.
[0043] Figure 6H Schematic diagram of the eighth example partition structure of AV1.
[0044] Figure 6I This is a diagram of the ninth example partition structure of AV1.
[0045] Figure 6J Schematic diagram of the tenth example partition structure of AV1.
[0046] Figure 7A A diagram showing types of vertical binary splits in a multi-type tree structure.
[0047] Figure 7B Diagram showing types of horizontal binary splits in a multi-type tree structure.
[0048] Figure 7C A diagram showing types of vertical ternary splits in a multi-type tree structure.
[0049] Figure 7D Diagram showing types of horizontal ternary splits in a multi-type tree structure.
[0050] Figure 8An example diagram shows a CTU divided into multiple CUs, where the CTU has a quadtree and nested multi-type tree coding block structure.
[0051] Figure 9 A diagram showing eight nominal angles in AV1.
[0052] Figure 10 A diagram showing the current block and samples.
[0053] Figure 11 is an example diagram of positions of neighboring blocks of a current block according to an embodiment of the present disclosure.
[0054] Figure 12 is a schematic diagram of a decoder according to an embodiment of the present disclosure.
[0055] Figure 13 is a schematic diagram of a computer system suitable for implementing embodiments of the present disclosure. DETAILED DESCRIPTION
[0056] Figure 1 A simplified block diagram of a communication system (100) according to an embodiment of the present disclosure is shown. The system (100) may include at least two terminals, such as a first terminal (110) and a second terminal (120), interconnected via a network (150). For unidirectional transmission of data, the first terminal (110) may encode video data at a local location for transmission to the second terminal (120) via the network (150). The second terminal (120) may receive the encoded video data from the first terminal (110) from the network (150), decode the encoded video data, and display the recovered video data. Unidirectional data transmission is common in applications such as media services.
[0057] Figure 1 A second pair of terminals, namely a third terminal (130) and a fourth terminal (140), is shown. The third terminal (130) and the fourth terminal (140) are provided to support bidirectional transmission of encoded video, such as may occur during a video conference. For bidirectional transmission of data, each of the third terminal (130) and the fourth terminal (140) can encode video data captured at a local location for transmission to the other terminal via a network (150). Each of the third terminal (130) and the fourth terminal (140) can also receive encoded video data sent by the other terminal, can decode the encoded video data, and can display the recovered video data on a local display device.
[0058] exist Figure 1In the embodiment of the present invention, the first terminal (110), the second terminal (120), the third terminal (130), and the fourth terminal (140) may be servers, personal computers, smartphones, and / or any other type of terminal. For example, the first terminal (110), the second terminal (120), the third terminal (130), and the fourth terminal (140) may be laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (150) represents any number of networks that transmit encoded video data between the first terminal (110), the second terminal (120), the third terminal (130), and the fourth terminal (140), including, for example, wired and / or wireless communication networks. The communication network (150) may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunication networks, local area networks, wide area networks, and / or the Internet. For the purposes of this application, unless otherwise explained below, the architecture and topology of the network (150) may be irrelevant to the operations disclosed herein.
[0059] As an example, Figure 2 The video encoder and video decoder are shown in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0060] like Figure 2 As shown, a streaming system (200) may include an acquisition subsystem (213), which may include a video source (201) and an encoder (203). The video source (201) may be, for example, a digital camera and may be configured to create an uncompressed video sample stream (202). Compared to an encoded video stream, the uncompressed video sample stream (202) may provide a high data volume and may be processed by an encoder (203) coupled to the camera (201). The encoder (203) may include hardware, software, or a combination thereof to implement or implement various aspects of the disclosed subject matter as described in more detail below. Compared to the sample stream, the encoded video stream (204) may include a lower data volume and may be stored on a streaming server (205) for future use. One or more streaming clients (206) may access the streaming server (205) to retrieve a video stream (209), which may be a copy of the encoded video stream (204).
[0061] In an embodiment, the streaming server (205) can also function as a media-aware network element (MANE). For example, the streaming server (205) can be configured to tailor the encoded video stream (204) to provide different streams to one or more streaming clients (206). In an embodiment, the MANE can be provided separately from the streaming server (205) in the streaming system (200).
[0062] The streaming client (206) may include a video decoder (210) and a display (212). The video decoder (210) may, for example, decode a video stream (209). The video stream (209) is a copy of the received encoded video stream (204). The video decoder (210) also creates an output video sample stream (211). The video sample stream (211) may be presented on a display (212) or another rendering device (not depicted). In some streaming systems, the video streams (204, 209) may be encoded according to a certain video coding / compression standard. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. A video coding standard under development is informally referred to as Versatile Video Coding (VVC). Embodiments of the present disclosure may be used in the context of VVC.
[0063] Figure 3 A schematic functional block diagram of a video decoder (210) connected to a display (212) according to one embodiment of the present disclosure is shown.
[0064] The video decoder (210) may include a channel (312), a receiver (310), a buffer memory (315), an entropy decoder / parser (320), a scaler / inverse transform unit (351), an intra-frame prediction unit (352), a motion compensated prediction unit (353), an aggregator (355), a loop filter unit (356), a reference picture memory (357), and a current picture memory (358). In at least one embodiment, the video decoder (210) may include an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The video decoder (210) may also be partially or fully embodied in software. The software runs on one or more CPUs with associated memory.
[0065] In this or other embodiments, a receiver (310) may receive one or more encoded video sequences to be decoded by a video decoder (210); one encoded video sequence is received at a time, wherein each encoded video sequence is decoded independently of the other encoded video sequences. The encoded video sequence may be received from a channel (312), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not shown). The receiver (310) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (315) may be coupled between the receiver (310) and the entropy decoder / parser (320) (hereinafter referred to as the "parser"). However, when the receiver (310) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (315) may not be used, or the buffer memory (315) may be made smaller. For use on a best-effort packet network such as the Internet, a buffer memory (315) may also be required, which may be relatively large and may have an adaptive size.
[0066] The video decoder (210) may include a parser (320) to reconstruct symbols (321) from an entropy-coded video sequence. The types of symbols may include, for example, information for managing the operation of the video decoder (210) and potentially information for controlling a display device (e.g., display screen 212) that may be coupled to a display such as a video screen. Figure 2The decoder shown. The control information for the display device may be, for example, a parameter set fragment (not shown) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (320) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may be performed according to a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. The parser (320) may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), and the like. The parser (320) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.
[0067] The parser (320) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (315), thereby creating symbols (321).
[0068] Depending on the type of coded video picture or portion of a coded video picture (e.g., inter-frame and intra-frame pictures, inter-frame blocks and intra-frame blocks) and other factors, the reconstruction of the symbol (321) may involve multiple different units. Which units are involved and in what manner can be controlled by subgroup control information parsed from the coded video sequence by the parser (320). For the sake of brevity, the flow of such subgroup control information between the parser (320) and the multiple units below is not described.
[0069] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually broken down into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and may be integrated with each other. However, for the purposes of describing the disclosed subject matter, the conceptual breakdown into the following functional units is appropriate.
[0070] One unit is a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives quantized transform coefficients as symbols (321) from the parser (320) along with control information, including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (351) may output a block comprising sample values, which may be input to an aggregator (355).
[0071] In some cases, the output samples of the scaler / inverse transform unit (351) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) uses reconstructed information extracted from the current (partially reconstructed) picture in the current picture memory (358) to generate a block of the same size and shape as the block being reconstructed. In some cases, the aggregator (355) adds the prediction information generated by the intra-prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) on a per-sample basis.
[0072] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to inter-frame coded and potentially motion compensated blocks. In this case, the motion compensated prediction unit (353) may access the reference picture memory (357) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (321), these samples may be added to the output of the scaler / inverse transform unit (351) (in this case referred to as residual samples or residual signal) by the aggregator (355) to generate output sample information. The motion compensated prediction unit (353) may be controlled by a motion vector to obtain the predicted samples from the address in the reference picture memory (357). The motion vector is provided to the motion compensated prediction unit (353) in the form of the symbols (321). The symbols (321) may, for example, include X, Y and reference picture components. Motion compensation may also include interpolation of sample values extracted from the reference picture memory (357) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.
[0073] The output samples of the aggregator (355) may be used by various loop filtering techniques in a loop filter unit (356). The video compression techniques may include in-loop filtering techniques that are controlled by parameters included in the coded video stream and made available to the loop filter unit (356) as symbols (321) from the parser (320). However, in other embodiments, the video compression techniques may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of a coded picture or coded video sequence, as well as to previously reconstructed and loop-filtered sample values.
[0074] The output of the loop filter unit (356) may be a sample stream that may be output to a display device such as a display (212) and stored in a reference picture memory (357) for subsequent inter-picture prediction.
[0075] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (320)), the current picture may become part of the reference picture memory (357) and new current picture memory may be reallocated before starting reconstruction of a subsequent coded picture.
[0076] The video decoder (210) may perform decoding operations according to a predetermined video compression technique as documented in a standard such as ITU-T H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard used, in the sense that the coded video sequence adheres to the syntax of the video compression technique or standard, as specified in the video compression technique document or standard, and specifically in a profile within the video compression technique document or standard. To conform to some video compression techniques or standards, the complexity of the coded video sequence is also required to be within a range defined by a hierarchy of the video compression technique or standard. In some cases, the hierarchy limits a maximum picture size, a maximum frame rate, a maximum reconstruction sampling rate (measured in, for example, megasamples per second), a maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further defined by a Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.
[0077] In an embodiment, a receiver (310) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.
[0078] Figure 4 It is a schematic functional block diagram of a video encoder (203) related to a video source (201) according to an embodiment disclosed in the present application.
[0079] The video encoder (203) may, for example, include an encoder as a source encoder (430), a coding engine (432), a (local) decoder (433), a reference picture memory (434), a predictor (435), a transmitter (440), an entropy encoder (445), a controller (450) and a channel (460).
[0080] The video encoder (203) may receive video samples from a video source (201) (not part of the encoder) that may capture video images to be encoded by the video encoder (203).
[0081] The video source (201) may provide a source video sequence in the form of a stream of digital video samples to be encoded by the video encoder (203), wherein the stream of digital video samples may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (e.g., Y CrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (201) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (201) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed sequentially. The pictures themselves may be constructed as a spatial array of pixels, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be readily understood by those skilled in the art. The following description focuses on samples.
[0082] According to an embodiment, the video encoder (203) may encode and compress pictures of a source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller (450). The controller (450) controls other functional units as described below and is functionally coupled to these units. For the sake of brevity, the coupling is not shown in the figure. The parameters set by the controller (550) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The relationship between pixels and samples can be easily understood by those skilled in the art. The following description focuses on samples.
[0083] Some video encoders operate in a coding loop (which is readily recognized by those skilled in the art as a coding loop). As a simplified description, the coding loop may include an encoding portion of a source encoder (430) (e.g., responsible for creating symbols based on the input picture to be encoded and a reference picture) and a (local) decoder (433) embedded in the video encoder (203). The decoder (433) reconstructs the symbols to create sample data. When any compression between the symbols and the encoded video stream is lossless in certain video compression techniques, the (remote) decoder also creates this sample data. The reconstructed sample stream (sample data) is input to a reference picture memory (434). Because the decoding of the symbol stream produces bit-accurate results that are independent of the decoder's location (local or remote), the contents of the reference picture memory are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same sample values that the decoder will "see" when using prediction during decoding. Those skilled in the art are aware of this basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, for example due to channel errors).
[0084] The operation of the "local" decoder (433) can be combined with the operation of Figure 3 The "remote" decoder described in detail is identical. However, when symbols are available and the entropy encoder (445) and parser (320) are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210), including the channel (312), receiver (310), buffer memory (315), and parser (320), may not be fully implemented in the local decoder (433).
[0085] At this point, it can be observed that any decoder technology other than parsing / entropy decoding present in the decoder also needs to be present in a substantially identical functional form in the corresponding encoder. For this reason, this application focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is mutually inverse to the decoder technology described comprehensively. A more detailed description is only required in certain areas and is provided below.
[0086] As part of its operation, the source encoder (430) may perform motion-compensated predictive coding. The motion-compensated predictive coding predictively encodes an input frame with reference to one or more previously encoded frames in a video sequence designated as "reference pictures." In this manner, the encoding engine (432) encodes the differences between pixel blocks of an input frame and pixel blocks of a reference frame that may be selected as a prediction reference for the input frame.
[0087] The local video decoder (433) may decode the encoded video data of a frame that may be designated as a reference frame based on the symbols created by the source encoder (430). The operation of the encoding engine (432) may be a lossy process. When the encoded video data is available at the video decoder ( Figure 4 When decoded at a remote location (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (433) replicates the decoding process that the video decoder may perform on the reference frame and may cause the reconstructed reference frame to be stored in the reference picture memory (434). In this way, the video encoder (203) may locally store a copy of the reconstructed reference frame that has common content (absent transmission errors) with the reconstructed reference picture that will be obtained by the remote video decoder.
[0088] The predictor (435) may perform a prediction search for the encoding engine (432). That is, for a new frame to be encoded, the predictor (435) may search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable prediction references for the new frame. The predictor (435) may operate on a pixel-by-pixel-block basis based on sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (435), it may be determined that the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (434).
[0089] The controller (450) may manage encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding video data.
[0090] The outputs of all of the above functional units may be entropy encoded in an entropy encoder (445). The entropy encoder performs lossless compression on the symbols generated by the various functional units using techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.
[0091] The transmitter (440) can buffer the encoded video sequence created by the entropy encoder (445) in preparation for transmission over a channel (460), which can be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (440) can combine the encoded video data from the source encoder (430) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0092] The controller (450) can manage the operation of the video encoder (203). During encoding, the controller (450) can assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture can generally be assigned as an intra picture (I picture), a predictive picture (P picture), or a bidirectional predictive picture (B picture).
[0093] An intra picture (I picture) can be a picture that can be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and features.
[0094] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict sample values for each block.
[0095] Bidirectionally predictive pictures (B pictures) can be encoded and decoded using intra prediction or inter prediction, which uses up to two motion vectors and reference indices to predict sample values for each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.
[0096] A source picture is typically spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples) and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, determined according to the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively coded, or they may be predictively coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. Pixel blocks of a P picture may be predictively coded using spatial prediction with reference to one previously coded reference picture or using temporal prediction. Blocks of a B picture may be predictively coded using spatial prediction with reference to one or two previously coded reference pictures or using temporal prediction.
[0097] The video encoder (203) may perform encoding operations according to a predetermined video coding technique or standard, such as Recommendation ITU-T H.265. In operation, the video encoder (203) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to a syntax specified by the video coding technique or standard used.
[0098] In an embodiment, the transmitter (440) may transmit additional data along with the encoded video. The source encoder (430) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, redundant pictures and slices, and other forms of redundant data, SEI messages, VUI parameter set fragments, and the like.
[0099] [Coding block partitioning in VP9 and AV1]
[0100] refer to 5A to 5D Partition structure (502) to (508), VP9 uses 4 partition trees starting from 64×64 level down to 4×4 level, with some additional restrictions on block 8×8. It should be noted that Figure 5D The partitioning denoted as R in [1] involves a recursion that repeats the same partitioning tree at a lower scale until the lowest 4×4 level is reached.
[0101] refer to Figures 6A to 6J AV1 not only expands the partition tree to 10 structures, but also increases the maximum size (called super block in VP9 / AV1 terminology) to start from 128×128. Note that AV1 includes 4:1 / 1:4 rectangular partitions that do not exist in VP9. Figures 6C to 6FThe partition type shown with 3 sub-partitions is called a "T-type" partition. Any rectangular partition will not be further subdivided. In addition to the coding block size, the coding tree depth can be defined to indicate the partition depth from the root node. Specifically, the coding tree depth of the root node (e.g., 128×128) is set to 0, and after the tree block is further split once, the coding tree depth is increased by 1.
[0102] Instead of enforcing fixed transform unit sizes as in VP9, AV1 allows luma coding blocks to be partitioned into transform units of multiple sizes, which can be represented by recursive partitioning down to level 2. To include AV1's extended coding block partitioning, square, 2:1 / 1:2, and 4:1 / 1:4 transform sizes from 4×4 to 64×64 are supported. For chroma blocks, only the largest possible transform unit is allowed.
[0103] [Block partitioning in HEVC]
[0104] In HEVC, the coding tree unit (CTU) can be partitioned into coding units (CU) to adapt to various local characteristics by using a quadtree (QT) structure represented as a coding tree. At the CU level, it can be decided whether to use inter-picture (temporal) or intra-picture (spatial) prediction to encode the picture area. Depending on the PU partition type, each CU is further partitioned into one, two or four prediction units (PUs). Within a PU, the same prediction process can be applied, and relevant information is sent to the decoder on a PU basis. After obtaining the residual block by applying the prediction process based on the PU partition type, the CU can be partitioned into transform units (TUs) according to another quadtree structure (such as the coding tree of the CU). One of the key features of the HEVC structure is that it has multiple partition concepts including CU, PU and TU. In HEVC, a CU or TU can have only a square shape, while a PU can have a square or rectangular shape for an inter-frame prediction block. In HEVC, a coding block can be further partitioned into four square sub-blocks, and a transform is performed on each sub-block (i.e., TU). Each TU may be further recursively partitioned (using quadtree partitioning) into smaller TUs, which may be referred to as a residual quadtree (RQT).
[0105] At picture boundaries, HEVC employs implicit quadtree partitioning so that a block will remain quadtree partitioned until the size fits the picture boundary.
[0106] [Quadtree with nested multi-type tree coding block structure in VVC]
[0107] In VVC, a quadtree of nested multi-type trees using binary and ternary partitioning structures replaces the concept of multiple partition unit types. That is, VVC does not include separation of the concepts of CU, PU, and TU unless the size of the CU is too large for the maximum transform length, and VVC supports more flexibility in the shape of CU partitions. In the coding tree structure, the CU can have a square or rectangular shape. The coding tree unit (CTU) is first partitioned by a quadtree (also known as a quadtree) structure. Then, the quadtree leaf nodes can be further partitioned by a multi-type tree structure. 7A to 7D As shown in Figures (532), (534), (536) and (538), there are four types of segmentation in the multi-type tree structure: Figure 7A The vertical binary split (SPLIT BT VER) shown in the figure, such as Figure 7B The horizontal binary split (SPLIT BT HOR) shown in the figure, Figure 7C The vertical three-way split (SPLIT TT VER) shown in the figure and Figure 7D The illustrated horizontal ternary split (SPLIT_TT_HOR). The multi-type tree leaf nodes may be referred to as coding units (CUs), and this partitioning may be used for prediction and transform processing without any further partitioning unless the CU is too large for the maximum transform length. This means that in most cases, CUs, PUs, and TUs have the same block size in a quadtree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is less than the width or height of the color components of the CU. Figure 8 An example of block partitioning of a CTU is shown in FIG. Figure 8 A CTU (540) is shown partitioned into multiple CUs with a quadtree and nested multi-type tree coding block structure, where bold edges represent quadtree partitions and dashed edges represent multi-type tree partitions. The quadtree with nested multi-type tree partitions provides a content-adaptive coding tree structure consisting of CUs.
[0108] In VVC, the maximum supported luma transform size is 64 × 64, and the maximum supported chroma transform size is 32 × 32. When the width or height of the CB is larger than the maximum transform width or maximum transform height, the CB can be automatically split in the horizontal and / or vertical direction to meet the transform size restrictions in that direction.
[0109] In VTM7, the coding tree scheme supports the ability to have separate block tree structures for luma and chroma. For P and B slices, the luma CTB and chroma CTB in one CTU may have to share the same coding tree structure. However, for I slices, luma and chroma can have separate block tree structures. When separate block tree mode is applied, the luma CTB is partitioned into CUs through one coding tree structure, and the chroma CTB is partitioned into chroma CUs through another coding tree structure. This means that a CU in an I slice can consist of coding blocks of the luma component or coding blocks of two chroma components, and a CU in a P or B slice can consist of coding blocks of all three color components, unless the video is monochrome.
[0110] [Directional intra prediction in AV1]
[0111] VP9 supports eight directional modes corresponding to angles from 45 degrees to 207 degrees. In order to exploit a wider variety of spatial redundancies in directional textures, in AV1, directional intra modes are extended to a set of angles with finer granularity. The original eight angles are slightly changed and set as nominal angles, and these eight nominal angles are named V_PRED (542), H PRED (543), D45 PRED (544), D135 PRED (545), D113 PRED (546), D157 PRED (547), D203 PRED (548), and D67_PRED (549). These eight nominal angles are used in Figure 9 , relative to the current block (541). The nominal angles may also be referred to as the basic directional modes supported by the codec standard. For each nominal angle, there are seven more refined angles, so AV1 has a total of 56 directional angles. The predicted angle is represented by the nominal intra angle plus an incremental angle that is -3 to 3 times the 3 degree step size. In order to implement the directional prediction modes in AV1 in a common way, all 56 directional intra prediction modes in AV1 are implemented using a unified directional predictor that projects each pixel to a reference sub-pixel position and interpolates the reference pixels through a 2-tap bilinear filter.
[0112] [Non-directional smooth intra predictor in AV1]
[0113] In AV1, there are five non-directional smooth intra prediction modes, namely DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H. For DC prediction, the average of the left neighboring samples and the upper neighboring samples is used as the predictor for the block to be predicted. For the PAETH predictor, the top reference sample, the left reference sample, and the upper left reference sample are first retrieved, and then the value closest to (top + left - upper left) is set as the predictor for the pixel to be predicted. Figure 10 The positions of the top sample (554), left sample (556), and top-left sample (558) of the current pixel (552) in the current block (550) are shown. For SMOOTH, SMOOTH_V, and SMOOTH_H modes, the current block (550) is predicted using quadratic interpolation in the vertical or horizontal direction, or an average of these two directions.
[0114] [Chromaticity predicted from luminance]
[0115] For chroma components, in addition to the 56 directional modes and 5 non-directional modes, Chroma from Luma (CfL) is a chroma-only intra prediction mode that models chroma pixels as linear functions of coincident reconstructed luma pixels. CfL prediction can be expressed as shown in Equation 1 below:
[0116] CfL(α)=α×LAC+DC (Equation 1)
[0117] Where LAC represents the AC contribution of the luma component, α represents the parameters of the linear model, and DC represents the DC contribution of the chroma component. Specifically, the reconstructed luma pixels are subsampled to the chroma resolution, and then the average value is subtracted to form the AC contribution. In order to approximate the chroma AC component from the AC contribution, the decoder does not need to calculate the scaling parameter as in some background technologies. Instead, AV1 CfL determines the parameter α based on the original chroma pixels and signals the parameter α in the codestream. This reduces the decoder complexity and produces more accurate predictions. For the DC contribution of the chroma component, it can be calculated using the intra-frame DC mode, which is sufficient for most chroma content and has a mature and fast implementation.
[0118] For signaling chroma intra prediction modes, eight nominal directional modes, five non-directional modes, and the CfL mode can first be signaled. The context for signaling these modes can depend on the corresponding luma mode at the top left position of the current block. Then, if the current chroma mode is a directional mode, an additional flag can be signaled to indicate the delta angle relative to the nominal angle.
[0119] In AV1, for each directional nominal pattern, there are seven incremental angles, and all seven incremental angles are signaled / resolved regardless of the orientation of adjacent nominal patterns, which is not optimal.
[0120] Embodiments of the present disclosure may address the above-mentioned problems and / or other problems.
[0121] The embodiments of the present disclosure can be used alone or in any combination. In addition, each embodiment (e.g., method, encoder, and decoder) can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-volatile computer-readable medium.
[0122] In this disclosure, when it is mentioned that one directional intra prediction mode is close to another directional intra prediction mode, it means that the absolute difference in prediction angle between the two modes is within a given threshold T. In one example, T can be set to 1 or 2.
[0123] According to one or more embodiments, rather than signaling / resolving the same number of incremental angles for each nominal angle, the number of incremental angles signaled / resolved may depend on the direction of the current nominal angle and / or adjacent nominal angles.
[0124] According to an embodiment, if the nominal angle of the current block is the same as (or close to) the nominal angle of at least one neighboring block, all allowed incremental angles for the current block (e.g., 7 incremental angles) are signaled / parsed. Otherwise, only a subset of allowed incremental angles (e.g., 3 incremental angles) is signaled / parsed for the nominal angle. The allowed incremental angles are incremental angles that, when combined with the nominal angle, correspond to prediction angles supported by the codec standard. In one embodiment, the subset of allowed incremental angles is predefined and fixed. In one example, the subset of allowed incremental angles is {-2, 0, 2}. In another example, the subset of allowed incremental angles is {-1, 0, 1}. In another embodiment, the subset of allowed incremental angles depends on the absolute difference in prediction angle between the nominal angle of the current block and the nominal angle of the neighboring block. In one example, the size of the subset becomes smaller as the absolute difference in prediction angle between the nominal angle of the current block and the nominal angle of the neighboring block increases. In one example, when the nominal angle of the current block is equal to the nominal angle of one of the neighboring blocks, all incremental angles (e.g., 7 incremental angles) are signaled / resolved. Otherwise, when the nominal angle of the current block is adjacent to the nominal angle of one of the neighboring blocks, five incremental angles, such as {-2, -1, 0, 1, 2}, are signaled / resolved. Otherwise, only three incremental angles, such as {-2, 0, 2}, are signaled / resolved.
[0125] According to an embodiment, the positions of neighboring blocks are fixed and predefined, and up to N neighboring blocks are employed, where N is a positive integer, such as 7. Figure 11An example of the positions of neighboring blocks of the current block (600) is shown. In an embodiment, only a subset of the neighboring blocks at seven positions may be used. In one example, only the neighboring blocks at positions A, B, C, D, and F may be used. In another example, only the neighboring blocks at positions A, B, C, D, E, and F may be used. In another example, only the neighboring blocks at positions B, C, D, and F may be used.
[0126] According to one or more embodiments, a context or cumulative density function (CDF) is used to signal the incremental angles of a luma block, and the context or cumulative density function (CDF) may be based on the nominal angles of neighboring blocks. According to an embodiment, the context (or CDF) used to signal the incremental angles of a luma block depends on the nominal angle of the current luma block and the nominal angles of neighboring luma blocks. In an embodiment, when the nominal angle of the luma block is equal to one of the nominal angles of the neighboring luma blocks (or is close to a given threshold), one or more contexts (or CDFs) are used to signal the incremental angles of the luma block. Otherwise, another one or more different contexts (or CDFs) may be used to signal the incremental angles of the chroma blocks. In an embodiment, a reference may be implemented Figure 11 The location of the neighboring block in question.
[0127] According to one or more embodiments, rather than directly signaling the delta angle for a current luma block, the delta angle may be mapped to an associated index, and the index may be signaled. According to an embodiment, the delta angle may be mapped to the associated index according to a predefined mapping table. According to an embodiment, when the nominal angle of the luma block is equal to one of the nominal angles of a neighboring luma block (or close to a given threshold), the delta angle may be mapped to the index, and the index may be signaled. According to an embodiment, the mapping from delta angle to associated index may depend on the delta angle of the current luma block and the delta angle of the neighboring luma blocks.
[0128] Two examples of predefined mapping tables are shown below in Tables 1 and 2, where the values in the first column represent the delta angles of adjacent luminance blocks, the values in the first row represent the associated indexes of the delta angles of the current luminance block, and the values in the remaining entries represent the delta angles of the current luminance block. As an example with reference to Table 1, when the delta angle of the adjacent luminance block is -9 and the delta angle of the current luminance block is -9, the associated index is 0. As another example with reference to Table 1, when the delta angle of the adjacent luminance block is -3 and the delta angle of the current luminance block is -9, the associated index is 3.
[0129] Table 1: Example of a mapping table #1
[0130]
[0131]
[0132] Table 2: Example of mapping table #2
[0133] 0 1 2 3 4 5 6 -9 -9 -6 -3 0 3 6 9 -6 -6 -3 -9 0 3 6 9 -3 -3 0 -6 3 -9 6 9 0 0 3 -3 6 -6 9 -9 3 3 6 0 9 -3 -6 -9 6 6 9 3 0 -3 -6 -9 9 9 6 3 0 -3 -6 -9
[0134] According to an embodiment, at least one processor and a memory storing computer program instructions may be provided. The computer program instructions, when executed by the at least one processor, may implement an encoder or a decoder and may perform any number of functions described in the present disclosure. For example, referring to Figure 12 At least one processor may implement a decoder (800). The decoder (800) may include, for example, a decoding module (810), which may decode an encoded picture received (e.g., from an encoder). The decoding module (810) may include, for example, a first acquisition module (820), a second acquisition module (830), a determination module (840), a signaling module (850), an index signaling module (860), an identification module (870), and / or a prediction module (880).
[0135] According to an embodiment of the present disclosure, the first acquisition module (820) can acquire the nominal angle of the current block of the encoded picture, and the nominal angle of the current block is used for intra-frame prediction.
[0136] According to an embodiment of the present disclosure, the second acquisition module (830) can acquire the nominal angle of at least one neighboring block of the current block, and the nominal angle of the at least one neighboring block is used for intra-frame prediction.
[0137] According to an embodiment of the present disclosure, the determination module (840) can determine whether to signal all allowed incremental angles of the nominal angle of the current block or only signal a subset of the allowed incremental angles based on a comparison between the nominal angle of the current block and the nominal angle of the at least one adjacent block.
[0138] According to an embodiment of the present disclosure, the signaling module (850) may signal all of the allowed incremental angles or the subset of the allowed incremental angles based on the determination.
[0139] According to an embodiment of the present disclosure, the index signaling module (860) can represent the index with a signal.
[0140] According to an embodiment of the present disclosure, the identification module (870) may identify the incremental angle of the current block by using a mapping table, wherein the index is mapped to the incremental angle of the current block in the mapping table.
[0141] According to an embodiment of the present disclosure, the prediction module (880) may predict the current block.
[0142] According to an embodiment, as understood by those of ordinary skill in the art based on the above description, encoder-side processing corresponding to the above processing may be implemented by an encoding code for encoding a picture.
[0143] The technology of the above-mentioned embodiments of the present disclosure can be implemented as computer software through computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 13 A computer system (900) is shown that is suitable for implementing certain embodiments of the disclosed subject matter.
[0144] The computer software may be encoded in any suitable machine code or computer language, and may be assembled, compiled, linked, or similar mechanisms to create a code comprising instructions, which may be directly executed by a computer central processing unit (CPU), graphics processing unit (GPU), or the like, or executed through decoding, microcode, or the like.
[0145] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablets, servers, smartphones, gaming devices, IoT devices, and the like.
[0146] Figure 13 The components shown for the computer system (900) are exemplary in nature and are not intended to limit the scope of use or functionality of computer software implementing embodiments of the present application. Nor should the configuration of the components be interpreted as having any dependency or requirement on any one or combination of components shown in the exemplary embodiment of the computer system (900).
[0147] The computer system (900) may include certain human-computer interface input devices. Such human-computer interface input devices may respond to input from one or more human users through tactile input (e.g., keyboard input, sliding, data glove movement), audio input (e.g., sound, applause), visual input (e.g., gestures), and olfactory input (not shown). The human-computer interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0148] The human-machine interface input device may include one or more of the following (only one of which is depicted): a keyboard (901), a mouse (902), a touchpad (903), a touch screen (910), a data glove (not shown), a joystick (905), a microphone (906), a scanner (907), and a camera (908).
[0149] The computer system (900) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen (910), a data glove (not shown), or a joystick (905), but there may also be tactile feedback devices that are not used as input devices). For example, these human-computer interface output devices may be audio output devices (e.g., speakers (909), headphones (not shown)), visual output devices (e.g., screens (910) including cathode ray tube screens, liquid crystal screens, plasma screens, organic light emitting diode screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which may output two-dimensional visual output or output of more than three dimensions through means such as stereoscopic image output; virtual reality glasses (not shown), holographic displays, and cigarette boxes (not shown)) and printers (not shown).
[0150] The computer system (900) may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable optical disks (CD / DVD ROM / RW) (920) with CD / DVD or similar media (921), thumb drives (922), removable hard drives or solid-state drives (923), traditional magnetic media such as magnetic tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security software dongles (not shown), and the like.
[0151] Those skilled in the art will also understand that the term "computer-readable media" used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0152] The computer system (900) may also include an interface to one or more communication networks. For example, the network may be wireless, wired, or optical. The network may also be a local area network, a wide area network, a metropolitan area network, an in-vehicle network, an industrial network, a real-time network, a delay-tolerant network, and the like. Networks also include local area networks such as Ethernet, wireless local area networks, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), television wired or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), in-vehicle and industrial networks (including CANBus), and the like. Some networks typically require an external network interface adapter for connecting to some universal data port or peripheral bus (949) (e.g., a USB port of the computer system (900)); other systems are typically integrated into the core of the computer system (900) by connecting to a system bus as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smartphone computer system). By using any of these networks, the computer system (900) can communicate with other entities. The communication can be one-way, receiving only (e.g., wireless television), one-way, sending only (e.g., CAN bus to certain CAN bus devices), or two-way, such as to other computer systems via a local or wide area digital network. The communication can include communication with a cloud computing environment (955). Each of the networks and network interfaces described above can use certain protocols and protocol stacks.
[0153] The aforementioned human-machine interface devices, human-accessible storage devices, and network interfaces (954) may be connected to the core (940) of the computer system (900).
[0154] The core (940) may include one or more central processing units (CPUs) (941), graphics processing units (GPUs) (942), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (943), hardware accelerators for specific tasks (944), and the like. These devices, as well as read-only memory (ROM) (945), random access memory (946), internal mass storage (e.g., internal non-user accessible hard disk drives, solid-state drives, etc.) (947), and the like, may be connected via a system bus (948). In some computer systems, the system bus (948) may be accessed in the form of one or more physical plugs so that it can be expanded with additional central processing units, graphics processing units, and the like. Peripheral devices may be attached directly to the core's system bus (948) or connected via a peripheral bus (949). Peripheral bus architectures include PCI (Peripheral Controller Interface), USB (Universal Serial Bus), and the like. A graphics adapter (950) may be included in the core (940).
[0155] The CPU (941), GPU (942), FPGA (943), and accelerator (944) can execute certain instructions, which, when combined, can constitute the aforementioned computer code. The computer code can be stored in ROM (945) or RAM (946). Transient data can also be stored in RAM (946), while permanent data can be stored, for example, in internal mass storage (947). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (941), GPUs (942), mass storage (947), ROM (945), RAM (946), etc.
[0156] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be specially designed and constructed for the purposes of this application, or may be medium and code well known and available to those skilled in the art of computer software.
[0157] As an example and not a limitation, a computer system having architecture (900), in particular core (940), can provide the function of executing software contained in one or more tangible computer-readable media as a processor (including CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be a medium associated with the above-mentioned user-accessible large-capacity memory, as well as a specific memory of the core (940) having non-volatile properties, such as core internal large-capacity memory (947) or ROM (945). Software for implementing various embodiments of the present application can be stored in such a device and executed by the core (940). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can enable the core (940), in particular the processor therein (including CPU, GPU, FPGA, etc.) to perform a specific process or a specific part of a specific process described herein, including defining a data structure stored in RAM (946) and modifying such a data structure according to a process defined by the software. Additionally or alternatively, the computer system may provide functionality hardwired in logic or otherwise contained in circuitry (e.g., accelerator (944)) that may operate in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic and vice versa. Where appropriate, references to computer-readable media may include circuitry (e.g., an integrated circuit (IC)) storing the executing software, circuitry containing the executing logic, or both. The present application includes any suitable combination of hardware and software.
[0158] Although this application has described a number of non-limiting exemplary embodiments, various modifications, permutations, and equivalent substitutions of the embodiments are within the scope of this application. It should be understood that those skilled in the art will be able to design various systems and methods that, although not explicitly shown or described herein, embody the principles of this application and are therefore within the spirit and scope of this application.
Claims
1. A video decoding method, characterized in that: The method comprises: receiving an encoded picture; and Decoding the encoded picture, the decoding comprising: Obtaining a nominal angle of a current block of the encoded picture, wherein the nominal angle of the current block is a basic directional mode used for intra prediction of the current block; Obtaining a nominal angle of at least one neighboring block of the current block, where the nominal angle of the at least one neighboring block is a basic directional mode used for intra prediction of the at least one neighboring block; determining, based on a comparison between the nominal angle of the current block and the nominal angle of the at least one neighboring block, whether to signal all allowed incremental angles of the nominal angle of the current block or only a subset of the allowed incremental angles; based on said determining, signaling all of said allowed incremental angles or signaling said subset of said allowed incremental angles; and The current block is predicted based on the signaling of all or the subset of the allowed delta angles.
2. The method according to claim 1, wherein The determining includes signaling all of the allowed incremental angles based on an absolute difference between a first value and a second value being less than or equal to a threshold, wherein the first value corresponds to the nominal angle of the current block and the second value corresponds to the nominal angle of the at least one neighboring block.
3. The method according to claim 2, wherein: The threshold is 2.
4. The method according to any one of claims 1 to 3, wherein The determining includes signaling only the subset of the allowed incremental angles based on an absolute difference between a first value corresponding to the nominal angle of the current block and a second value corresponding to the nominal angle of the at least one neighboring block being greater than a threshold.
5. The method according to claim 4, wherein The number of allowed incremental angles in the subset is determined based on the absolute difference.
6. The method according to any one of claims 1 to 3, wherein The comparison is performed between the nominal angle of the current block and the nominal angles of only a predetermined number of neighboring blocks.
7. The method according to any one of claims 1 to 3, wherein Said signaling all or said subset of said allowed incremental angles comprises: based on said determining, signaling at least one of said allowed incremental angles by using a cumulative density function (CDF).
8. The method according to any one of claims 1 to 3, wherein the decoding further comprises: Use signals to represent indexes; as well as The incremental angle of the current block is identified by using a mapping table, wherein the index is mapped to the incremental angle of the current block in the mapping table.
9. The method according to any one of claims 1 to 3, wherein the decoding further comprises: Use signals to represent indexes; as well as The incremental angle of the current block is identified by using a mapping table, wherein in the mapping table, the index is mapped to the incremental angle of the current block and the incremental angle of the at least one neighboring block.
10. The method according to any one of claims 1 to 3, wherein: The current block is a luminance block.
11. A video encoding method, characterized in that: The method comprises: Encoding a code stream containing an encoded picture; the encoding includes: Obtaining a nominal angle of a current block of the picture, wherein the nominal angle of the current block is a basic directional mode for intra prediction of the current block; Obtaining a nominal angle of at least one neighboring block of the current block, where the nominal angle of the at least one neighboring block is a basic directional mode used for intra prediction of the at least one neighboring block; determining, based on a comparison between the nominal angle of the current block and the nominal angle of the at least one neighboring block, whether to signal all allowed incremental angles of the nominal angle of the current block or only a subset of the allowed incremental angles; Based on the determination, signaling in the codestream all of the allowed incremental angles or signaling the subset of the allowed incremental angles; and Output the code stream.
12. A method for storing a video stream, characterized in that: Execute the video encoding method according to claim 11 to generate a video code stream, and store the video code stream.
13. A method for transmitting a video stream, characterized in that: Execute the video encoding method according to claim 11 to generate a video code stream, and transmit the video code stream.
14. A video encoding and decoding system, characterized in that: The system comprises: at least one memory configured to store computer program code; and At least one processor is configured to access the computer program code and operate according to instructions of the computer program code to perform the method according to any one of claims 1 to 13.
15. A video decoding device, characterized in that: The device comprises: A decoding module is used to decode the received encoded picture, and the decoding module includes: A first acquisition module is configured to acquire a nominal angle of a current block of the encoded picture, where the nominal angle of the current block is a basic directional mode used for intra-frame prediction of the current block; a second acquisition module, acquiring a nominal angle of at least one neighboring block of the current block, where the nominal angle of the at least one neighboring block is a basic directional mode used for intra-frame prediction of the at least one neighboring block; a determination module that determines whether to signal all allowed incremental angles of the nominal angle of the current block or only a subset of the allowed incremental angles based on a comparison between the nominal angle of the current block and the nominal angle of the at least one neighboring block; a signaling module that, based on the determining, signals all of the allowed incremental angles or the subset of the allowed incremental angles; and A prediction module predicts the current block based on the signaling of all or a subset of the allowed delta angles.
16. A non-transitory computer-readable medium storing a computer program / instruction and a video code stream, wherein when the computer program / instruction is executed by a processor, the computer program / instruction implements the steps of the video encoding method of claim 11 to generate the video code stream.
Citation Information
Patent Citations
Directional intra-prediction coding
US20190124339A1