Method, apparatus and medium for multi-reference line intra prediction in video decoding
By adopting a multi-reference line selection scheme in video encoding and replacing non-adjacent reference lines with the top adjacent reference lines, the problems of high storage requirements and low signaling efficiency are solved, and the storage design of video encoding and decoding is optimized, and the performance is improved.
Patent Information
- Application Number
- CN202510590085.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-01-06
- Filing Date
- 2022-01-18
- Publication Date
- 2025-07-04
AI Technical Summary
The existing video encoding technology has problems with high storage requirements and low signaling efficiency in intra prediction, especially in multi-reference line selection schemes, where memory size and buffer requirements are too large.
Using a multi-reference row selection (MRLS) scheme, the encoded video code stream is received through the device, the parameters are extracted to indicate non-adjacent reference rows, and the top adjacent reference rows are used as the value of the non-adjacent reference rows at the current block boundary to reduce storage requirements.
It effectively reduces the amount of memory usage, optimizes the storage design and signaling efficiency, and improves the performance of video encoding and decoding.
Smart Images

Figure CN120264017A_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application with the application date of January 18, 2022, the Chinese patent application number of 202280003790.1, and the invention title of "Method, Apparatus, and Medium for Multi-Reference Line Intra Prediction in Video Decoding". Technical Field
[0002] Embodiments of the present application relate to video encoding and / or decoding technologies, and specifically relate to an improved low-storage design and signaling for a multi-reference line selection scheme. Background Art
[0003] The background description provided herein is for the purpose of generally presenting the content of the embodiments of the present application. The extent to which the work of the currently named inventors described in this background art section and various aspects of this specification has been carried out does not indicate that it was eligible as prior art at the time of filing of the present application, and has never been expressly or implicitly recognized as prior art for the content of the embodiments of the present application.
[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each picture having a spatial size of, for example, luminance samples of 1920×1080 and associated full or subsampled chrominance samples. The series of pictures can have a fixed or variable picture rate (alternatively, called frame rate) of, for example, 60 pictures per second or 60 frames per second. For streaming or data processing, uncompressed video has specific bitrate requirements. For example, a video with a pixel resolution of 1920×1080, a frame rate of 60 frames per second, a chrominance subsampling of 4:2:0, and 8 bits per pixel per color channel requires a bandwidth of nearly 1.5 Gbit / s. Such a one-hour video requires more than 600 GB of storage space.
[0005] One purpose of video encoding and decoding can be to reduce redundancy in an uncompressed input video signal through compression. Compression can help reduce the above-mentioned bandwidth and / or storage space requirements, in some cases by two orders of magnitude or more than two orders of magnitude. Lossless compression and lossy compression, as well as combinations thereof, can be employed. Lossless compression refers to techniques where an exact copy of the original signal can be reconstructed from the compressed original signal through the decoding process. Lossy compression refers to the following encoding / decoding process: the original video information is not fully retained during encoding and cannot be fully recovered during decoding. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal can be used for the intended application, although some information is lost. In the case of video, lossy compression is widely adopted in many applications. The amount of distortion that can be tolerated depends on the application. For example, users of some consumer video streaming applications can tolerate higher distortion compared to users of movie or television broadcast applications. The compression rate achievable through a specific encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortion generally allows the encoding algorithm to produce higher losses and higher compression rates.
[0006] Video encoders and decoders can utilize techniques and steps from several broad categories, which include, for example, motion compensation, Fourier transform, quantization, and entropy coding.
[0007] Video codec technology can include techniques known as intra-frame coding. In intra-frame coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into sample blocks. When all sample blocks are encoded in intra-frame mode, the picture can be referred to as an intra-frame picture. Intra-frame pictures and their derivatives (e.g., independent decoder refresh pictures) can be used to reset the decoder state and thus can be used as the first picture in an encoded video bitstream and a video session, or as a still image. Then, the samples of the blocks after intra-frame prediction can be transformed, transformed into the frequency domain, and the transform coefficients thus generated can be quantized before entropy coding. Intra-frame prediction represents a technique for minimizing sample values in the pre-transform domain. In some cases, the smaller the DC value and the smaller the AC coefficients after transformation, the fewer bits are required to represent the block after entropy coding for a given quantization step.
[0008] For example, traditional intra-coding known from, e.g., MPEG-2 generation coding techniques does not use intra-prediction. However, some more recent video compression techniques include techniques that attempt to encode / decode blocks based on, e.g., surrounding sample data and / or metadata that are obtained during spatially adjacent encoding and / or decoding and that are ordered in decoding order before the data block being intra-encoded or decoded. Such techniques are hereinafter referred to as "intra-prediction" techniques. It should be noted that, at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed and not reference data from reference pictures.
[0009] Intra-prediction can have many different forms. When more than one such technique is available in a given video coding technique, the technique in use may be referred to as an intra-prediction mode. One or more intra-prediction modes may be provided in a particular codec. In some cases, a mode may have sub-modes and / or may be associated with various parameters, and the mode / sub-mode information and the intra-coding parameters of a video block may be encoded separately or jointly included in a mode codeword. Which codeword is used for a given combination of mode, sub-mode, and / or parameters may affect the coding efficiency gain through intra-prediction and thus may affect the entropy coding technique used to convert the codeword into a bitstream.
[0010] H.264 introduced an intra-prediction mode that was refined in H.265 and further refined in more recent coding techniques such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Generally, for intra-prediction, available values belonging to already available neighboring samples can be used to form a predictor block. For example, the available values of a particular set of neighboring samples can be copied along a certain direction and / or row into the predictor block. The reference to the direction of use can be encoded in the bitstream or can itself be predicted.
[0011] Reference Figure 1A , a subset of 9 prediction directions specified among the 33 possible intra-prediction directions of H.265 (corresponding to the 33 angular modes of the 35 intra-modes specified in H.265) is depicted in the lower right. The point (101) where the arrows converge represents the sample being predicted. The arrows indicate the directions along which neighboring samples are used to predict the sample at 101. For example, arrow (102) indicates that the sample (101) is predicted based on one or more neighboring samples in the upper right, at a 45-degree angle to the horizontal direction. Similarly, arrow (103) indicates that the sample (101) is predicted based on one or more neighboring samples in the lower left of the sample (101), at a 22.5-degree angle to the horizontal direction.
[0012] Still referring to Figure 1A, a square block (104) of 4×4 samples is depicted in the upper left (indicated by the bold dashed line). The square block (104) includes 16 samples, each sample being labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (starting from the top) and the first sample in the X dimension (starting from the left). Similarly, sample S44 is the fourth sample in both the Y and X dimensions, within the block (104). Since the block size is 4×4 samples, S44 is located in the lower right corner. Exemplary reference samples following a similar numbering scheme are also shown. The reference samples are labeled with "R", their Y position (e.g., row index) relative to the block (104), and X position (column index). In H.264 and H.265, predictive samples adjacent to the block being reconstructed are used.
[0013] Intra picture prediction of block 104 can start by copying reference sample values from adjacent samples according to the predicted direction signaled. For example, assume that the encoded video bitstream includes signaling that indicates the predicted direction of arrow (102) for this block 104, i.e., samples are predicted based on one or more predictive samples at the upper right, at a 45-degree angle to the horizontal direction. In this case, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Then sample S44 is predicted based on reference sample R08.
[0014] In some cases, the values of multiple reference samples can be combined, for example, by interpolation, to calculate a reference sample; especially when the direction is not divisible by 45 degrees.
[0015] As video coding techniques continue to evolve, the number of possible directions increases. In H.264 (in 2003), for example, nine different directions can be used for intra prediction. In H.265 (in 2013), it increased to 33 directions, and at the time of the embodiments of this application, JEM / VVC / BMS can support up to 65 directions. Experimental studies have been conducted to help identify the most suitable intra prediction directions, and some techniques in entropy coding can be used to encode those most suitable directions with a small number of bits, accepting a certain bit cost for the directions. Additionally, sometimes the direction itself can be predicted based on the adjacent directions used in the intra prediction of the decoded adjacent blocks.
[0016] Figure 1B A schematic diagram (180) is shown, which depicts 65 intra prediction directions according to JEM, to illustrate the increase in the number of prediction directions in various coding techniques developed over time.
[0017] In an encoded video bitstream, the manner in which bits representing an intra prediction direction are mapped to that prediction direction may vary depending on the video coding technology; for example, the range can vary from a simple direct mapping of the prediction direction to an intra prediction mode, to a mapping of the prediction direction to a codeword, to a complex adaptive scheme involving the most probable mode, and similar techniques. However, in all cases, there may be certain intra prediction directions that are statistically less likely to occur in video content compared to certain other directions. Since the goal of video compression is to reduce redundancy, in a well-designed video coding technology, those less likely directions may be represented by more bits compared to the more likely directions.
[0018] Inter-picture prediction or inter prediction may be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or a portion thereof (reference picture) is spatially offset along a direction indicated by a motion vector (hereinafter referred to as MV) and can be used to predict a newly reconstructed picture or picture portion (e.g., block). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or may have three dimensions, with the third dimension being an indication of the reference picture being used (similar to a temporal dimension).
[0019] In some video compression techniques, a current MV applicable to a certain region of sample data may be predicted based on other MVs, for example, based on other MVs that are related to another region of sample data that is spatially adjacent to the region being reconstructed and that are in front of the current MV in decoding order. Doing so can greatly reduce the total amount of data required to encode the MVs by eliminating redundancy in the related MVs, thereby increasing the compression efficiency. MV prediction can work effectively, for example, because when encoding an input video signal obtained from a camera (referred to as natural video), there is the statistical likelihood that a region larger than the region to which a single MV applies moves in a similar direction in the video sequence. Thus, in some cases, similar motion vectors derived from the MVs of adjacent regions can be used to predict the larger region. This makes the actual MV for a given region similar to or the same as the MV predicted based on the surrounding MVs. Subsequently, after entropy coding, the MV can be represented using fewer bits compared to when the MV is directly encoded (rather than being predicted based on adjacent MVs). In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, for example, due to rounding errors that occur when calculating the predicted value based on multiple surrounding MVs, the MV prediction itself can be lossy.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (Recommendation ITU-T H.265, "High Efficiency Video Coding", December 2016). Among the various MV prediction mechanisms specified in H.265, the technique hereinafter referred to as "spatial merge" is described below.
[0021] Specifically, referring to Figure 2 , the current block (201) includes samples that have been found by the encoder during the motion search process, and the samples can be predicted based on a previous block of the same size that has generated a spatial offset. The MV can be derived from metadata associated with one or more reference pictures instead of directly encoding the MV. For example, the MV is derived from the nearest reference picture using the MV associated with any one of five surrounding samples labeled A0, A1 and B0, B1, B2 (corresponding to 202 to 206 respectively). In H.265, MV prediction can use the prediction value from the same reference picture that the adjacent blocks are using. SUMMARY OF THE INVENTION
[0022] Embodiments of the present application describe various embodiments of methods, devices, and computer-readable storage media for video coding and / or decoding.
[0023] According to one aspect, an embodiment of an embodiment of the present application provides a method for multi-reference row intra prediction in video decoding. The method includes:
[0024] Receiving an encoded video bitstream of a current block by a device, the device including a memory storing instructions and a processor communicating with the memory;
[0025] Extracting, by the device, parameters from the encoded video bitstream, the parameters indicating a non-adjacent reference row for intra prediction in the current block;
[0026] Dividing, by the device, the current block to obtain a plurality of sub-blocks; and
[0027] In response to a sub-block among the plurality of sub-blocks being located at the boundary of the current block and the multi-reference row selection being applied to the current block, using the top adjacent reference row as the value of all top non-adjacent reference rows of the coding block.
[0028] According to another aspect, an embodiment of an embodiment of the present application provides a device for multi-reference row intra prediction in video decoding. The device includes:
[0029] A memory storing instructions; and
[0030] A processor in communication with the memory, wherein when the processor executes the instructions, the processor is configured to cause the device to perform the method according to an embodiment of the present application.
[0031] In another aspect, an embodiment of an embodiment of the present application provides a non-transitory computer-readable medium storing instructions that, when executed by a processor, are configured to cause the processor to perform the method according to an embodiment of the present application.
[0032] In another aspect, an embodiment of an embodiment of the present application provides a method for storing a video bitstream, characterized in that the video bitstream is decoded based on the method according to an embodiment of the present application.
[0033] The above and other aspects and their implementation manners are described in more detail in the drawings, the description, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the drawings.
[0035] Figure 1A A schematic diagram showing an exemplary subset of intra prediction direction modes.
[0036] Figure 1B A diagram showing an exemplary intra prediction direction.
[0037] Figure 2 A schematic diagram showing a current block for motion vector prediction and its surrounding spatial merge candidates in one example.
[0038] Figure 3 A schematic diagram showing a simplified block diagram of a communication system according to an exemplary embodiment.
[0039] Figure 4 A schematic diagram showing a simplified block diagram of a communication system according to an exemplary embodiment.
[0040] Figure 5 A schematic diagram showing a simplified block diagram of a video decoder according to an exemplary embodiment.
[0041] Figure 6 A schematic diagram showing a simplified block diagram of a video encoder according to an exemplary embodiment.
[0042] Figure 7 A block diagram of a video encoder according to another exemplary embodiment is shown.
[0043] Figure 8Shows a block diagram of a video decoder according to another exemplary embodiment.
[0044] Figure 9 Shows a scheme for coding block partitioning according to an exemplary embodiment of the present application.
[0045] Figure 10 Shows another scheme for coding block partitioning according to an exemplary embodiment of the present application.
[0046] Figure 11 Shows another scheme for coding block partitioning according to an exemplary embodiment of the present application.
[0047] Figure 12 Shows another scheme for coding block partitioning according to an exemplary embodiment of the present application.
[0048] Figure 13 Shows a scheme for partitioning a coding block into multiple transform blocks and the coding order of the transform blocks according to an exemplary embodiment of the present application.
[0049] Figure 14 Shows another scheme for partitioning a coding block into multiple transform blocks and the coding order of the transform blocks according to an exemplary embodiment of the present application.
[0050] Figure 15 Shows another scheme for partitioning a coding block into multiple transform blocks according to an exemplary embodiment of the present application.
[0051] Figure 16 Shows an intra prediction scheme based on respective reference rows according to an exemplary embodiment of the present application.
[0052] Figure 17 Shows a flowchart of a method according to an exemplary embodiment of the present application.
[0053] Figure 18 Shows an intra prediction scheme based on multiple reference rows according to an exemplary embodiment of the present application.
[0054] Figure 19 Shows an intra prediction scheme based on multiple reference rows according to an exemplary embodiment of the present application.
[0055] Figure 20 Shows a schematic diagram of a computer system according to an exemplary embodiment of the present application. Detailed Description
[0056] Now, the present invention will be described in detail below with reference to the accompanying drawings, which form a part of the present invention and illustrate specific examples of embodiments by way of illustration. However, it should be noted that the present invention can be embodied in many different forms, and thus, the subject matter covered or claimed is intended to be construed as not limited to any of the embodiments set forth below. It should also be noted that the present invention can be embodied as a method, apparatus, component, or system. Thus, embodiments of the present invention can take, for example, the form of hardware, software, firmware, or any combination thereof.
[0057] Throughout the specification and claims, terms may have nuanced meanings that are revealed or implied in the context in addition to having the explicitly stated meaning. As used herein, phrases such as "in one embodiment" or "in some embodiments" do not necessarily refer to the same embodiment, and as used herein, phrases such as "in another embodiment" or "in other embodiments" do not necessarily refer to different embodiments. Similarly, as used herein, phrases such as "in one implementation" or "in some implementations" do not necessarily refer to the same implementation, and as used herein, phrases such as "in another implementation" or "in other implementations" do not necessarily refer to different implementations. For example, it is intended that the claimed subject matter includes combinations of all or part of the exemplary embodiments / implementations.
[0058] Generally, terms can be understood, at least in part, from their usage in context. For example, terms such as "and", "or", or "and / or" as used herein can include a variety of meanings that can depend, at least in part, on the context in which such terms are used. Generally, if "or" is used to associate a list (e.g., A, B, or C), then "or" is intended to mean A, B, and C (used herein in an inclusive sense) as well as A, B, or C (used herein in an exclusive sense). Additionally, terms such as "one or more" or "at least one" as used herein can, at least in part, depend on the context and can be used to describe any feature, structure, or property in a singular sense or can be used to describe a combination of features, structures, or properties in a plural sense. Similarly, terms such as "a", "an", or "the" can also be understood to convey singular usage or convey plural usage, at least in part, depending on the context. Additionally, the terms "based on" or "determined by" can be understood to not necessarily be intended to convey a set of exclusive factors but can allow for the existence of additional factors that are not necessarily explicitly described, at least in part, depending on the context.
[0059] Figure 3A simplified block diagram of a communication system (300) according to an embodiment of the present application is shown. The communication system (300) includes a plurality of terminal devices that can communicate with each other through, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected by a network (350). In Figure 3 the example, the first pair of terminal devices (310) and (320) can perform unidirectional data transmission. For example, the terminal device (310) can encode video data (e.g., video data of a video picture stream collected by the terminal device (310)) for transmission over the network (350) to another terminal device (320). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to recover the video pictures, and display the video pictures based on the recovered video data. Unidirectional data transmission can be implemented in applications such as media services.
[0060] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, which can be implemented, for example, during a video conferencing application. For bidirectional data transmission, in one example, each of the terminal devices (330) and (340) can encode video data (e.g., video data of a video picture stream collected by the terminal device) for transmission over the network (350) to the other of the terminal devices (330) and (340). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), decode the encoded video data to recover the video pictures, and display the video pictures on an accessible display device based on the recovered video data.
[0061] In Figure 3In the example, the terminal devices (310), (320), (330), and (340) can be implemented as servers, personal computers, and smart phones. However, the applicability of the underlying principles of the embodiments of this application is not limited thereto. The embodiments of the embodiments of this application can be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing devices, etc. The network (350) represents any number of networks that transmit the encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) can exchange data in circuit-switched channels, packet-switched channels, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless explicitly stated herein, the architecture and topology of the network (350) may be irrelevant to the operation of the embodiments of this application.
[0062] As an example of an application for the disclosed subject matter, Figure 4 illustrates the placement of a video encoder and a video decoder in a video streaming environment. The disclosed subject matter is equally applicable to other video applications, including, for example, video conferencing, digital TV, broadcasting, gaming, virtual reality, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0063] A video streaming system may include a video capture subsystem (413), and the capture subsystem (413) may include a video source (401) such as a digital camera. The video source (401) is used to create an uncompressed video picture or image stream (402). In one example, the video picture stream (402) includes samples recorded by the digital camera of the video source 401. Compared with the encoded video data (404) (or encoded video bitstream), the video picture stream (402), depicted as a thick line to emphasize the high data volume, can be processed by an electronic device (420), and the electronic device (420) includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter described in more detail below. Compared with the uncompressed video picture stream (402), the encoded video data (404) (or encoded video bitstream (404)), depicted as a thin line to emphasize the lower data volume, can be stored on a streaming server (405) for future use, or directly stored in a downstream video device (not shown). One or more streaming client subsystems, such as Figure 4The client subsystems (406) and (408) therein can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and generates an uncompressed output video picture stream (411) that can be presented on a display (412) (e.g., a display screen) or other rendering device (not depicted). The video decoder 410 can be configured to perform some or all of the various functions described in embodiments of the present application. In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally referred to as Next Generation Video Coding (VVC). The disclosed subject matter can be used in the context of VVC and other video coding standards.
[0064] It should be noted that the electronic devices (420) and (430) can include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and the electronic device (430) can further include a video encoder (not shown).
[0065] Hereinafter, Figure 5 A block diagram of a video decoder (510) according to any embodiment of the embodiments of the present application is shown. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., receiving circuitry). The video decoder (510) can be used to replace Figure 4 the video decoder (410) in the example of
[0066] A receiver (531) may receive one or more encoded video sequences to be decoded by a video decoder (510). In the same or another embodiment, one encoded video sequence may be decoded at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. Each video sequence may be associated with a plurality of video frames or pictures. The encoded video sequences may be received from a channel (501), which may be a hardware / software link leading to a storage device storing the encoded video data or a streaming source transmitting the encoded video data. The receiver (531) may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective processing circuits (not depicted). The receiver (531) may separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory (515) may be provided between the receiver (531) and an entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be located external to and separate from the video decoder (510) (not depicted). In still other applications, a buffer memory (not depicted) may be provided external to the video decoder (510) for purposes such as preventing network jitter, and another additional buffer memory (515) may be provided inside the video decoder (510) for purposes such as handling playout timing. When the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (515) may not be needed, or the buffer memory may be made smaller. For use on a service packet network such as the Internet, a buffer memory (515) of sufficient size may be needed, and the size of the buffer memory (515) may be relatively large. Such a buffer memory may be implemented with an adaptive size and may be implemented at least partially in an operating system or a similar element (not depicted) external to the video decoder (510).
[0067] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequences. The categories of these symbols include information for managing the operation of the video decoder (510) and potential information for controlling a rendering device such as a display (512) (e.g., a display screen), which may or may not be an integral part of the electronic device (530) but may be coupled to the electronic device (530), such as Figure 5As shown. The control information for presenting the device may be in the form of supplementary enhancement information (SEI message) or a video usability information (VUI) parameter set segment (not depicted). The parser (520) may perform parsing / entropy decoding on the encoded video sequence received by the parser (520). The entropy coding of the encoded video sequence may be performed according to a video coding technology or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (520) may also extract information from the encoded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc.
[0068] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).
[0069] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (521) may involve multiple different processing or functional units. The units involved and the way the units are involved may be controlled by the parser (520) through subgroup control information parsed from the encoded video sequence. For simplicity, such subgroup control information flows between the parser (520) and multiple processing or functional units below are not depicted.
[0070] In addition to the function blocks already mentioned, the video decoder (510) may be conceptually divided into multiple functional units as described below. In actual implementations operating under commercial constraints, many of these functional units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, multiple functions divided conceptually are adopted in the following embodiments of this application.
[0071] The first unit may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) may receive quantized transform coefficients as symbols (521) and control information from the parser (520), including information indicating which type of inverse transform to use, block size, quantization factor / parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (551) may output a block including sample values, and the sample values may be input into the aggregator (555).
[0072] In some cases, the output samples of the scaler / inverse transform (551) may belong to an intra-coded block; that is, a block that does not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra picture prediction unit (552). In some cases, the intra picture prediction unit (552) may use the surrounding block information that has been reconstructed and stored in the current picture buffer (558) to generate a block having the same size and shape as the block being reconstructed. For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, an aggregator (555) may add, on a per-sample basis, the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).
[0073] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-coded and potentially motion-compensated block. In such a case, a motion compensation prediction unit (553) may access a reference picture memory (557) to extract samples for inter picture prediction. After motion compensating the extracted samples according to the sign (521) belonging to the block, these samples may be added by an aggregator (555) to the output of the scaler / inverse transform unit (551) (the output of unit 551 may be referred to as residual samples or a residual signal), thereby generating output sample information. The extraction of prediction samples by the motion compensation prediction unit (553) from an address within the reference picture memory (557) may be controlled by a motion vector, which may be provided in the form of a sign (521) for use by the motion compensation prediction unit (553), and the sign (521) may have, for example, an X component, a Y component (offset), and a reference picture component (temporal). Motion compensation may also include interpolation of sample values extracted from the reference picture memory (557) when using sub-sample accurate motion vectors, and motion compensation may also be associated with a motion vector prediction mechanism, etc.
[0074] The output samples of the aggregator (555) may be subject to various loop filtering techniques in a loop filter unit (556). Video compression techniques may include in-loop filter techniques that are controlled by parameters included in an encoded video sequence (also referred to as an encoded video bitstream) and that may be available to the loop filter unit (556) as a sign (521) from a parser (520). However, video compression techniques may also respond to meta information obtained during the decoding of a previous (in decoding order) portion of an encoded picture or an encoded video sequence, and to previously reconstructed and loop-filtered sample values. Multiple types of loop filters may be included in various orders as part of the loop filter unit 556, as will be described in further detail below.
[0075] The output of the loop filter unit (556) can be a sample stream, which can be output to the rendering device (512) and stored in the reference picture memory (557) for future inter-picture prediction.
[0076] Once fully reconstructed, some of the encoded pictures can be used as reference pictures for future inter-picture prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (by, for example, the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before starting to reconstruct subsequent encoded pictures.
[0077] The video decoder (510) can perform decoding operations according to a predetermined video compression technique adopted in a standard such as ITU-T Recommendation H.265. In the sense that the encoded video sequence conforms to the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard, the encoded video sequence can conform to the syntax specified by the video compression technique or standard used. Specifically, the profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available under that profile. To conform to the standard, the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the hypothetical reference decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0078] In some exemplary embodiments, the receiver (531) can receive additional (redundant) data when receiving the encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by the video decoder (510) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data can take forms such as temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0079] Figure 6 A block diagram of a video encoder (603) according to an exemplary embodiment of the present application is shown. The video encoder (603) can be included in an electronic device (620). The electronic device (620) can also include a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used to replace Figure 4 the video encoder (403) in the example of
[0080] A video encoder (603) may receive video samples from a video source (601) (which is not part of the electronic device (620) in the example Figure 6 presented). The video source (601) may capture video images to be encoded by the video encoder (603). In another example, the video source (601) may be implemented as part of the electronic device (620).
[0081] The video source (601) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (603). The digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit...), any color space (e.g., BT.601 YCrCb, RGB, XYZ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) may be a storage device capable of storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. Video data may be provided as a plurality of individual pictures or images that, when viewed in sequence, are given motion. The pictures themselves may be constructed as spatial pixel arrays, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. A person of ordinary skill in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.
[0082] According to some exemplary embodiments, the video encoder (603) may encode and compress pictures of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed constitutes a function of the controller (650). In some embodiments, the controller (650) may be functionally coupled to other functional units as described below and control the other functional units. For simplicity, couplings are not depicted in the figure. Parameters set by the controller (650) may include rate control-related parameters (picture skip, quantizer, λ value of rate-distortion optimization techniques...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured to have other suitable functions that relate to optimizing the video encoder (603) for a particular system design.
[0083] In some exemplary embodiments, the video encoder (603) may be configured to operate in an encoding loop. As an overly simplified description, in one example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to how a (remote) decoder may create sample data, even though the embedded decoder 633 processes a video stream encoded by the source encoder 630 without entropy encoding (since in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols in the entropy encoding and the encoded video bitstream may be lossless). The reconstructed sample stream (sample data) is input to the reference picture memory (634). Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs, for example, when the synchronization cannot be maintained due to channel errors) is used to improve the encoding quality.
[0084] The operation of the "local" decoder (633) may be the same as that of a "remote" decoder such as the video decoder (510) described in detail above in conjunction with Figure 5 However, briefly referring additionally to Figure 5 , since the symbols are available and the entropy encoder (645) and the parser (520) are capable of encoding / decoding the symbols losslessly into the encoded video sequence, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633) in the encoder.
[0085] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding that may only exist in the decoder may also necessarily exist in the corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter may sometimes focus on the decoder operation, which is combined with the decoding part of the encoder. Thus, the description of the encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology described comprehensively. In the following, a more detailed description of the encoder is provided only in certain areas or aspects.
[0086] During operation, in some exemplary implementations, the source encoder (630) may perform motion-compensated predictive coding that predictive-codes an input picture by referencing one or more previously-encoded pictures designated as "reference pictures" in a video sequence. In this way, the encoding engine (632) encodes the differences (or residuals) in the color channels between pixel blocks of the input picture and pixel blocks of the reference picture, which may be selected as a prediction reference for the input picture.
[0087] The local video decoder (633) may decode the encoded video data of pictures that may be designated as reference pictures, based on symbols created by the source encoder (630). The operation of the encoding engine (632) may advantageously be a lossy process. When the encoded video data may be decoded in a video decoder ( Figure 6 not shown), the reconstructed video sequence may generally be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that may be performed by the video decoder on the reference pictures and may cause the reconstructed reference pictures to be stored in the reference picture cache (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference pictures that has the same content (absent transmission errors) as the reconstructed reference pictures that will be obtained by a distal (remote) video decoder.
[0088] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that may be used as a suitable prediction reference for the new picture. The predictor (635) may operate on a per-pixel-block basis of sample blocks to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (634).
[0089] The controller (650) may manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding the video data.
[0090] The outputs of all the above functional units may be entropy-encoded in the entropy encoder (645). The entropy encoder (645) losslessly compresses the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0091] The transmitter (640) may buffer the encoded video sequence created by the entropy encoder (645) to prepare for transmission over the communication channel (660), which may be a hardware / software link to a storage device that can store the encoded video data. The transmitter (640) may combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0092] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures may typically be assigned to any of the following picture types:
[0093] Intra pictures (I pictures), which may be pictures that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Instantaneous Decoder Refresh (“IDR”) pictures. Those of ordinary skill in the art are aware of these variants of I pictures and their corresponding applications and characteristics.
[0094] Predictive pictures (P pictures), which may be pictures that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0095] Bi - predictive pictures (B pictures), which may be pictures that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.
[0096] Source pictures can typically be spatially subdivided into multiple sample-encoding blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined by the encoding assignments of the corresponding pictures applied to the blocks. For example, blocks of an I picture can be non-predictively encoded, or blocks of an I picture can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be predictively encoded with reference to a previously encoded reference picture through spatial prediction or through temporal prediction. Blocks of a B picture can be predictively encoded with reference to one or two previously encoded reference pictures through spatial prediction or through temporal prediction. For other purposes, source pictures or pictures in intermediate processing can be subdivided into other types of blocks. The subdivision of encoding blocks and other types of blocks may or may not follow the same way, as described in further detail below.
[0097] The video encoder (603) can perform encoding operations according to a predetermined video encoding technique or standard such as ITU-T Recommendation H.265. In operation, the video encoder (603) can perform various compression operations, including predictive encoding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technique or standard being used.
[0098] In some exemplary embodiments, the transmitter (640) can transmit additional data when transmitting the encoded video. The source encoder (630) can include such data as part of the encoded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0099] The captured video can be a plurality of source pictures (video pictures) in a time series. Intra picture prediction (usually simplified to intra prediction) utilizes the spatial correlation within a given picture, while inter picture prediction utilizes the temporal or other correlations between pictures. For example, a particular picture being encoded / decoded can be divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0100] In some exemplary embodiments, bidirectional prediction techniques may be used for inter - picture prediction. According to such bidirectional prediction techniques, two reference pictures are used, e.g., a first reference picture and a second reference picture that are before the current picture in the video in decoding order (but may be past or future respectively in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be jointly predicted by a combination of the first reference block and the second reference block.
[0101] In addition, merge mode techniques may be used for inter - picture prediction to improve encoding efficiency.
[0102] According to some exemplary embodiments of the embodiments of the present application, predictions such as inter - picture prediction and intra - picture prediction are performed on a per - block basis. For example, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression. CTUs in a picture may have the same size, e.g., 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU may include three parallel coding tree blocks (CTBs): one luminance CTB and two chrominance CTBs. Each CTU may be recursively split into one or more coding units (CUs) in a quadtree manner. For example, a 64×64 - pixel CTU may be split into a 64×64 - pixel CU, or 4 32×32 - pixel CUs. Each of one or more 32×32 blocks may be further split into 4 16×16 - pixel CUs. In some exemplary embodiments, each CU may be analyzed during encoding to determine the prediction type for the CU among multiple prediction types, e.g., an inter - frame prediction type or an intra - frame prediction type. Depending on temporal and / or spatial predictability, a CU may be split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In one embodiment, prediction operations in encoding (encoding / decoding) are performed on a per - prediction - block basis. The splitting of a CU into PUs (or PBs of different color channels) may be performed in various spatial patterns. For example, a luminance or chrominance PB may include a matrix of values for samples (e.g., luminance values), and the samples may be, for example, 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0103] Figure 7 A diagram of a video encoder (703) according to another exemplary embodiment of the embodiments of the present application is shown. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a video picture sequence, and encode the processing block into an encoded picture that is part of an encoded video sequence. In an example, the video encoder (703) may be used to replace Figure 4 the video encoder (403) in the example of
[0104] For example, the video encoder (703) receives a matrix of sample values for processing a block, such as a prediction block of 8×8 samples. Then, the video encoder (703) uses, for example, rate distortion optimization (RD O) to determine whether to use an intra mode, an inter mode, or a bi-prediction mode to optimally encode the processing block. When it is determined to encode the processing block in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into an encoded picture; and when it is determined to encode the processing block in the inter mode or the bi-prediction mode, the video encoder (703) may use inter prediction or bi-prediction techniques, respectively, to encode the processing block into an encoded picture. In some exemplary embodiments, a merge mode may be used as a sub-mode of inter-picture prediction, where a motion vector is derived from one or more motion vector predictors without resorting to encoded motion vector components external to the predictor. In some other exemplary embodiments, there may be motion vector components applicable to the subject block. Thus, the video encoder (703) may include components not explicitly shown in Figure 7 , such as a mode decision module for determining the prediction mode of the processing block.
[0105] In Figure 7 's example, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together as shown in the exemplary arrangement in Figure 7 .
[0106] The inter encoder (730) is configured to receive samples of a current block (e.g., the processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a later picture in display order), generate inter prediction information (e.g., a description of redundant information, a motion vector, merge mode information according to inter coding techniques), and calculate an inter prediction result (e.g., a predicted block) based on the inter prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on the encoded video information using a decoding unit 633 embedded in the Figure 6 exemplary encoder 620 in Figure 7 , as shown by the residual decoder 728 in
[0107] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with encoded blocks in the same picture, and generate quantized coefficients after transformation, and in some cases also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). The intra encoder (722) can calculate an intra prediction result (e.g., a predicted block) based on the intra prediction information and a reference block in the same picture.
[0108] The general controller (721) can be configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines the prediction mode of a block and provides a control signal to the switch (726) based on the prediction mode. For example, when the prediction mode is the intra mode, the general controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream; and when the prediction mode of the block is the inter mode, the general controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.
[0109] The residual calculator (723) can be configured to calculate the difference (residual data) between the received block and the prediction result of a block selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) can be configured to encode the residual data to generate transform coefficients. For example, the residual encoder (724) can be configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various exemplary embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture can be buffered in a memory circuit (not shown) and used as a reference picture.
[0110] The entropy encoder (725) can be configured to format the bitstream to include the encoded blocks and perform entropy encoding. The entropy encoder (725) is configured to include various information in the bitstream. For example, the entropy encoder (725) can be configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. When encoding a block in the merge submode of the inter mode or the bi - prediction mode, there may be no residual information.
[0111] Figure 8 FIG. shows an exemplary video decoder (810) according to another embodiment of the embodiments of the present application. The video decoder (810) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In one example, the video decoder (810) can be used instead of Figure 4 the video decoder (410) in the example of
[0112] In Figure 8 the example of Figure 8 the video decoder (810) includes an entropy decoder (871), an inter - frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra - frame decoder (872) coupled together as shown in the exemplary arrangement of
[0113] The entropy decoder (871) can be configured to reconstruct certain symbols from the encoded picture, where these symbols represent the syntax elements that make up the encoded picture. Such symbols can include, for example, the mode for encoding a block (e.g., intra mode, inter mode, bi - prediction mode, merge submode, or another submode), prediction information (e.g., intra prediction information or inter prediction information) that can identify certain samples or metadata for use by the intra - frame decoder (872) or the inter - frame decoder (880) for prediction, residual information in the form of, for example, quantized transform coefficients, etc. In one example, when the prediction mode is inter or bi - prediction mode, the inter - frame prediction information is provided to the inter - frame decoder (880); and when the prediction type is intra - frame prediction type, the intra - frame prediction information is provided to the intra - frame decoder (872). The residual information can be inverse - quantized and provided to the residual decoder (873).
[0114] The inter - frame decoder (880) can be configured to receive the inter - frame prediction information and generate an inter - frame prediction result based on the inter - frame prediction information.
[0115] The intra - frame decoder (872) can be configured to receive the intra - frame prediction information and generate a prediction result based on the intra - frame prediction information.
[0116] The residual decoder (873) can be configured to perform inverse quantization to extract the dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) can also utilize certain control information (used to include quantization parameter (QP)), and this control information can be provided by the entropy decoder (871) (the data path is not depicted because this is only low-data-volume control information).
[0117] The reconstruction module (874) can be configured to combine, in the spatial domain, the residual output by the residual decoder (873) with the prediction result (which can be output by the inter-frame prediction module or the intra-frame prediction module, as the case may be) to form a reconstructed block, and the reconstructed block forms part of the reconstructed picture, and the reconstructed picture is part of the reconstructed video. It should be noted that other suitable operations such as deblocking operations can also be performed to improve the visual quality.
[0118] It should be noted that any suitable technology can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In some exemplary embodiments, one or more integrated circuits can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In another embodiment, one or more processors executing software instructions can be used to implement the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810).
[0119] Turning to the coding block partitioning, in some exemplary implementations, a predetermined pattern can be applied. As Figure 9 shown, an exemplary 4-way partitioning tree can be used that starts from a first predetermined level (e.g., the 64×64 block level) and drills down to a second predetermined level (e.g., the 4×4 level). For example, the basic block can be subject to four partitioning options indicated by 902, 904, 906, and 908, where the partitions marked as R can be allowed for recursive partitioning because, as Figure 9 indicated, the same partitioning tree can be repeated on a small scale until the smallest level (e.g., the 4×4 level). In some implementations, additional restrictions can be applied to the Figure 9 partitioning scheme. In the Figure 9 implementation, rectangular partitioning (e.g., 1:2 / 2:1 rectangular partitioning) can be allowed, but rectangular partitioning may not be allowed to be recursive, while square partitioning can be recursive. If needed, following Figure 9Recursively partitioning generates a set of final coded blocks. This scheme can be applied to one or more color channels.
[0120] Figure 10 Another exemplary predefined partitioning pattern is shown that allows recursive partitioning to form a partitioning tree. As Figure 10 shown, an exemplary 10-way partitioning structure or pattern can be predefined. The root block can start from a predefined level (e.g., start from the 128×128 level or the 64×64 level). Figure 10 The exemplary partitioning structure of includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. In Figure 10 the second row having a partitioning type of 3-way partitioning indicated by 1002, 1004, 1006, and 1008 can be referred to as a "T-shaped" partition. The "T-shaped" partitions 1002, 1004, 1006, and 1008 can be referred to as the left T-shaped, top T-shaped, right T-shaped, and bottom T-shaped. In some implementations, further subdivision of Figure 10 the rectangular partitions is not allowed. The coding tree depth can be further defined to indicate the split depth starting from the root node or root block. For example, for the root node or root block (e.g., for a 128×128 block), the coding tree depth can be set to 0, and after further splitting the root block once, the coding tree depth increases by 1. In some implementations, only all the square partitions in 1010 are allowed to follow Figure 10 the pattern to be recursively partitioned into the next level of the partitioning tree. In other words, recursive partitioning may not be allowed for the square partitions having the patterns 1002, 1004, 1006, and 1006. If needed, following Figure 10 recursively partitioning generates a set of final coded blocks. This scheme can be applied to one or more color channels. Figure 10 Recursively partitioning generates a set of final coded blocks. This scheme can be applied to one or more color channels.
[0121] After splitting or partitioning the base block following any of the above partitioning processes or other processes, a set of final partitions or coded blocks can be obtained again. Each of these partitions can be at one of various partitioning levels. Each partition can be referred to as a coded block (CB). For the various exemplary partitioning implementations above, each generated CB can have any allowed size and partitioning level. The partition is called a coded block because the partition can form a unit for which some basic encoding / decoding decisions can be made, and the encoding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest level in the final partition represents the depth of the coded block partitioning tree. The coded block can be a luminance coded block or a chrominance coded block.
[0122] In some other exemplary implementations, a quadtree structure can be used to recursively split a basic luminance block and chrominance blocks into coding units. This splitting structure can be referred to as a Coding Tree Unit (CTU), and the CTU is split into Coding Units (CUs) by using the quadtree structure so that the partitioning adapts to various local characteristics of the basic CTU. In such an implementation, an implicit quadtree split can be performed at the picture boundary so that the blocks will maintain the quadtree split until the size fits the picture boundary. The term CU is used to commonly refer to units of luminance coding blocks and chrominance coding blocks (CBs).
[0123] In some implementations, the CB can be further partitioned. For example, the CB can be further partitioned into a plurality of Prediction Blocks (PBs) for intra prediction or inter prediction during the encoding and decoding processes. In other words, the CB can be further split into different sub-partitions in which independent prediction decisions / configurations can be made. In parallel, the CB can be further partitioned into a plurality of Transform Blocks (TBs) to describe the levels used to perform the transform or inverse transform of the video data. The partitioning schemes for splitting the CB into PBs and TBs can be the same or can be different. For example, each partitioning scheme can be performed using its own process based on various characteristics of the video data, for example. In some exemplary implementations, the PB partitioning scheme and the TB partitioning scheme can be independent. In some other exemplary implementations, the PB partitioning scheme, the TB partitioning scheme, and the boundaries can be interrelated. In some implementations, for example, the TB can be partitioned after the PB partitioning. In particular, after determining each PB following the partitioning of the coding block, each PB can be further partitioned into one or more TBs. For example, in some implementations, a PB can be split into one TB, two TBs, four TBs, or other numbers of TBs.
[0124] In some implementations, to divide a basic block into coding blocks and further into prediction blocks and / or transform blocks, the luminance channel and the chrominance channel may be processed differently. For example, in some implementations, for the luminance channel, it may be allowed to divide a coding block into prediction blocks and / or transform blocks, while for the chrominance channel, it may not be allowed to divide a coding block into prediction blocks and / or transform blocks in such a way. Thus, in such an implementation, only the transform and / or prediction of luminance blocks can be performed at the coding block level. For another example, the minimum transform block size of the luminance channel and the chrominance channel may be different. For example, compared with the chrominance channel, it may be allowed to divide the coding blocks of the luminance channel into smaller transform blocks and / or prediction blocks. For yet another example, between the luminance channel and the chrominance channel, the maximum depth of dividing a coding block into transform blocks and / or prediction blocks may be different. For example, compared with the chrominance channel, it may be allowed to divide the coding blocks of the luminance channel into deeper transform blocks and / or prediction blocks. For a specific example, a luminance coding block may be divided into transform blocks of multiple sizes, which may be represented by a recursive division down to up to 2 levels, and transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4 and transform block sizes from 4×4 to 64×64 may be allowed. However, for chrominance blocks, only the maximum possible transform blocks specified for luminance blocks may be allowed.
[0125] In some exemplary implementations, to divide a coding block into PBs, the depth, shape, and / or other characteristics of the PB division may depend on whether the PB is intra-coded or inter-coded.
[0126] The division of a coding block (or prediction block) into transform blocks can be implemented in various exemplary schemes, including but not limited to recursive or non-recursive quadtree splitting and predetermined pattern splitting, and additionally considering transform blocks located at the boundaries of the coding block or prediction block. Generally, the resulting transform blocks may be at different split levels, may not have the same size, and their shape may not need to be square (e.g., the transform block may be a rectangle with some allowed sizes and aspect ratios).
[0127] In some exemplary implementations, an encoded partitioning tree scheme or structure may be used. The encoded partitioning tree schemes for the luminance channel and the chrominance channel may not need to be the same. In other words, the luminance channel and the chrominance channel may have different encoded tree structures. Further, whether the luminance channel and the chrominance channel use the same or different encoded partitioning tree structures, and the actual encoded partitioning tree structure to be used, may depend on whether the slice to be encoded is a P slice, a B slice, or an I slice. For example, for an I slice, the chrominance channel and the luminance channel may have different encoded partitioning tree structures or encoded partitioning tree structure patterns, while for a P slice or a B slice, the luminance channel and the chrominance channel may share the same encoded partitioning tree scheme. When different encoded partitioning tree structures or patterns are applied, the luminance channel may be partitioned into CUs by one encoded partitioning tree structure, and the chrominance channel may be partitioned into chrominance CUs by another encoded partitioning tree structure.
[0128] Specific exemplary implementations of encoding block and transform block partitioning are described below. In such an exemplary implementation, the basic encoding block may be split into encoding blocks using the recursive quadtree splitting described above. At each level, whether further quadtree splitting of a particular partition should continue may be determined by local video data characteristics. The resulting CUs may be at various quadtree split levels of various sizes. A decision may be made at the CU level (or at the CU level for all three color channels) as to whether to use inter-picture (temporal) prediction or intra-picture (spatial) prediction to encode a picture region. Depending on the PB split type, each CU may be further split into one PB, two PBs, four PBs, or some other number of PBs. Within one PB, the same prediction process may be applied, and relevant information may be sent to the decoder based on the PB. After obtaining the residual blocks by applying the prediction process based on the PB split type, the CU may be partitioned into TUs according to another quadtree structure similar to the encoding tree of the CU. In this particular implementation, the CU or TU may be limited to a square, but does not have to be limited to a square. Further, in this particular example, for inter prediction, the PB may be square or rectangular, while for intra prediction, the PB can only be square. For example, an encoding block may be further split into four square TUs. Each TU may be further recursively split (using quadtree splitting) into smaller TUs, which are called residual quadtrees (RQTs).
[0129] Another specific example of partitioning a basic encoding block into CUs and other PBs and / or TUs is described below. For example, instead of using as Figure 10Rather than the multiple partition unit types shown, a quadtree with a nested multi-type tree can be used, which uses a binary and ternary split segmentation structure. The separation of the CB, PB, and TB concepts (i.e., dividing the CB into PB and / or TB, and dividing the PB into TB) can be abandoned, unless when the CB needs to have a size that is too large for the maximum transform length (in which case, further splitting of such a CB may be required). This exemplary splitting scheme can be designed to support greater flexibility in the CB partition shape, such that prediction and transformation can be performed at the CB level without the need for further partitioning. In such a coding tree structure, the CB can be square or rectangular. Specifically, the coding tree block (CTB) can first be partitioned by a quadtree structure. Then, the quadtree leaf nodes can be further partitioned by a multi-type tree structure. In Figure 11 An example of a multi-type tree structure is shown. Specifically, Figure 11 The exemplary multi-type tree architecture includes four splitting types, which are called vertical binary split (split_BT_VER) (1102), horizontal binary split (split_BT_HOR) (1104), vertical ternary split (split_TT_VER) (1106), and horizontal ternary split (split_TT_HOR) (1108). Then, the CB corresponds to the leaf of the multi-type tree. In this exemplary implementation, unless the CB is too large for the maximum transform length, this segmentation is used for prediction and transformation processing without any further partitioning. This means that in most cases, the CB, PB, and TB have the same block size in a quadtree with a nested multi-type tree coding block structure. An exception occurs when the supported maximum transform length is less than the width or height of the CB color component.
[0130] In Figure 12 An example of a quadtree with a nested multi-type tree coding block structure for the block partitioning of a CTB is shown. More specifically, Figure 12 It is shown that the CTB 1200 is a quadtree, split into four square partitions 1202, 1204, 1206, and 1208. For each quadtree split partition, a decision is made to further split using Figure 11 the multi-type tree structure. In Figure 12 the example, partition 1204 is not further split. Both partitions 1202 and 1208 use another quadtree split. For partition 1202, the upper left, upper right, lower left, and lower right partitions of the second-level quadtree split use the third-level quadtree split, Figure 11 1104 of, no split, and Figure 11 1108 of. Partition 1208 uses another quadtree split, and the upper left, upper right, lower left, and lower right partitions of the second-level quadtree split use Figure 11The third-level splitting 1106, no splitting, no splitting, and Figure 11 of 1104. The two sub-divisions of the upper-left third-level division of 1208 are further split according to 1104 and 1108. The division 1206 adopts a second-level splitting pattern following Figure 11 of 1102, and is split into two divisions, and these two divisions are further split according to Figure 11 of 1108 and 1102 at the third level. According to Figure 11 of 1104, the fourth-level splitting is further applied to one of the divisions.
[0131] For the above specific example, the maximum luminance transform size can be 64×64, and the maximum chrominance transform size supported can be different from the luminance, for example, at 32×32. When the width or height of a luminance coding block or a chrominance coding block is greater than the maximum transform width or height, the luminance coding block or the chrominance coding block can be automatically split along the horizontal direction and / or the vertical direction to meet the transform size limit along that direction.
[0132] In a specific example, in order to divide a basic coding block into the above-mentioned CBs, the coding tree scheme can support the ability of luminance and chrominance to have different block tree structures. For example, for P slices and B slices, the luminance CTB and the chrominance CTB in a CTU can share the same coding tree structure. For example, for I slices, the luminance and chrominance can have different coding block tree structures. When different block tree modes are applied, the luminance CTB can be divided into luminance CBs through one coding tree structure, and the chrominance CTB can be divided into chrominance CBs through another coding tree structure. This means that a CU in an I slice can be composed of coding blocks of the luminance component or coding blocks of the two chrominance components, and a CU in a P slice or a B slice is always composed of coding blocks of all three color components, unless the video is monochromatic.
[0133] Exemplary implementations of dividing coding blocks or prediction blocks into transform blocks and the coding order of transform blocks are further described in detail below. In some exemplary implementations, the transform division can support transform blocks of various shapes such as 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, where the size range of the transform blocks is, for example, from 4×4 to 64×64. In some implementations, if the coding block is less than or equal to 64×64, the transform block division can only be applied to the luminance component, so that for chrominance blocks, the transform block size is the same as the coding block size. Otherwise, if the coding block width or height is greater than 64, the luminance coding block and the chrominance coding block can be implicitly split into multiple min(W, 64)×min(H, 64) transform blocks and min(W, 32)×min(H, 32) transform blocks respectively.
[0134] In some exemplary implementations, for intra-coded blocks and inter-coded blocks, the coded block can be further divided into a plurality of transform blocks, where the division depth can reach a predetermined number of levels (e.g., 2 levels). The transform block division depth and size can be related. An exemplary mapping from the transform size at the current depth to the transform size at the next depth is shown in Table 1 below.
[0135] Table 1: Transform Division Size Settings
[0136] Transformation size at the current depth Transformation size at the next depth TX_4×4 TX_4×4 TX_8×8 TX_4×4 TX_16×16 TX_8×8 TX_32×32 TX_16×16 TX_64×64 TX_32×32 TX_4×8 TX_4×4 TX_8×4 TX_4×4 TX_8×16 TX_8×8 TX_16×8 TX_8×8 TX_16×32 TX_16×16 TX_32×16 TX_16×16 TX_32×64 TX_32×32 TX_64×32 TX_32×32 TX_4×16 TX_4×8 TX_16×4 TX_8×4 TX_8×32 TX_8×16 TX_16×64 TX_16×32 TX_64×16 TX_32×16
[0137] According to the exemplary mapping in Table 1, for a 1:1 square block, the next-level transform split can create four 1:1 square sub-transform blocks. The transform split can stop, for example, at 4×4. Thus, the transform size 4×4 at the current depth corresponds to the same size 4×4 at the next depth. In the example of Table 1, for a 1:2 / 2:1 non-square block, the next-level transform split will create two 1:1 square sub-transform blocks, and for a 1:4 / 4:1 non-square block, the next-level transform split will create two 1:2 / 2:1 sub-transform blocks.
[0138] In some exemplary implementations, for the luminance component of intra-coded blocks, additional restrictions can be applied. For example, for each transform division level, all sub-transform blocks can be restricted to be of equal size. For example, for a 32×16 coded block, the level 1 transform split will create two 16×16 sub-transform blocks, and the level 2 transform split will create eight 8×8 sub-transform blocks. In other words, the second-level split must be applied to all first-level sub-blocks to maintain the equal size of the transform units. In Figure 13 An example of the transform block division of an intra-coded square block following Table 1, and the coding order indicated by the arrows are shown. Specifically, 1302 shows the square coded block. The first-level split into 4 transform blocks of equal size according to Table 1 is shown in 1304, where the coding order is indicated by the arrows. The second-level split of all first-level blocks of equal size into 16 transform blocks of equal size according to Table 1 is shown in 1306, where the coding order is indicated by the arrows.
[0139] In some exemplary implementations, for the luminance component of inter-coded blocks, the above restrictions for intra-coding may not be applied. For example, after the first-level transform split, any one of the sub-transform blocks can be further independently split one more level. Thus, the resulting transform blocks can be of the same size or can be of different sizes. In Figure 14 An example of splitting an inter-coded block into transform blocks, and the coding order of the transform blocks are shown. In Figure 14In the example of Figure 14 , according to Table 1, the inter-coded block 1402 is split into transform blocks at two levels. At the first level, the inter-coded block is split into four equally sized transform blocks. Then, only one of the four transform blocks (not all of them) is further split into four sub-transform blocks, resulting in a total of seven transform blocks with two different sizes, as shown in 1404. The exemplary coding order of these seven transform blocks is as Figure 14 shown by the arrows in 1404 of
[0140] In some exemplary implementations, for the chrominance component, some additional restrictions can be applied to the transform blocks. For example, for the chrominance component, the transform block size can be as large as the coding block size, but not less than a predetermined size, such as 8×8.
[0141] In some other exemplary implementations, for coding blocks with a width (W) or height (H) greater than 64, the luminance coding block and the chrominance coding block can be implicitly split into multiples of min(W, 64)×min(H, 64) transform units and min(W, 32)×min(H, 32) transform units, respectively.
[0142] Figure 15 A further exemplary alternative for partitioning a coding block or a prediction block into transform blocks is shown. As Figure 15 shown, as an alternative to using recursive transform partitioning, a predetermined set of partitioning types can be applied to the coding block according to the transform type of the coding block. In the Figure 15 specific example shown, one of six exemplary partitioning types can be applied to split the coding block into various numbers of transform blocks. This scheme can be applicable to coding blocks or prediction blocks.
[0143] More specifically, as Figure 15 shown, Figure 15 the partitioning scheme provides up to six partitioning types for any given transform type. In this scheme, a transform type can be assigned to each coding block or prediction block based on, for example, rate-distortion cost. In one example, the partitioning type assigned to a coding block or a prediction block can be determined based on the transform partitioning type of the coding block or the prediction block. A specific partitioning type can correspond to a transform block partitioning size and pattern (or partitioning type), as shown by the four partitioning types in Figure 15 shown. The correspondence between various transform types and various partitioning types can be predefined. An exemplary correspondence is shown below, where uppercase letter labels indicate the transform types that can be assigned to coding blocks or prediction blocks based on rate-distortion cost:
[0144] · PARTITION_NONE (Partition_None): Assign a transform size equal to the block size.
[0145] ·PARTITION_SPLIT (Partition_Split): Allocate the following transform size, which is half of the width of the block size and half of the height of the block size.
[0146] ·PARTITION_HORZ (Partition_HORZ): Allocate the following transform size, which has the same width as the block size and is half of the height of the block size.
[0147] ·PARTITION_VERT (Partition_VERT): Allocate the following transform size, which is half of the width of the block size and has the same height as the block size.
[0148] ·PARTITION_HORZ4 (Partition_HORZ4): Allocate the following transform size, which has the same width as the block size and is one quarter of the height of the block size.
[0149] ·PARTITION_VERT4 (Partition_VERT4): Allocate the following transform size, which is one quarter of the width of the block size and has the same height as the block size.
[0150] In the above example, for the transformed blocks of the partition, as Figure 15 shown, all partition types include a unified transform size. This is only an example, not a limitation. In some other exemplary implementations, in a specific partition type (or mode), a mixed transform block size can be used for the transformed blocks of the partition.
[0151] Returning to intra prediction, in some exemplary implementations, the prediction of samples in an encoded block or a prediction block may be based on one reference line among a set of reference lines. In other words, instead of always using the nearest neighbor line (e.g., the line immediately above or immediately to the left of the prediction block as shown in FIG. 1 above), multiple reference lines may be provided as selection options for intra prediction. This intra prediction implementation may be referred to as multi-reference line selection (MRLS). In these implementations, the encoder decides and signals which one of the multiple reference lines is used to generate the intra prediction. On the decoder side, after parsing the reference line index, the reconstructed reference samples may be identified by looking up the specified reference line according to the intra prediction mode (e.g., the directional intra prediction mode, the non-directional intra prediction mode, and other intra prediction modes), thereby generating the intra prediction of the current intra prediction block. In some implementations, the reference line index may be signaled at the encoded block level, and only one of the multiple reference lines can be selected for the intra prediction of one encoded block. In some examples, more than one reference line may be selected together for intra prediction. For example, more than one reference line may be combined (with or without weights), averaged, interpolated, or otherwise processed to generate the prediction. In some exemplary implementations, MRLS may only be applicable to the luminance component and not to the chrominance component.
[0152] In Figure 16 , an example of 4-reference-line MRLS is depicted. As Figure 16 the example of shows, the intra-coded block 1602 may be predicted based on one horizontal reference line among the four horizontal reference lines 1604, 1606, 1608, and 1610 and one vertical reference line among the four vertical reference lines 1612, 1614, 1616, and 1618. Among these reference lines, 1610 and 1618 are directly adjacent reference lines. The reference lines may be indexed according to the distance from the encoded block. For example, the reference lines 1610 and 1618 may be referred to as zero reference lines, while the other reference lines may be referred to as non-zero reference lines. Specifically, the reference lines 1608 and 1616 may be referred to as the first reference lines; the reference lines 1606 and 1614 may be referred to as the second reference lines; and the reference lines 1604 and 1612 may be referred to as the third reference lines.
[0153] There may be some problems / difficulties associated with implementing a multi-reference line selection (MRLS) scheme. In some implementations without multi-reference line selection, only the most recent (or neighboring) top (or above) and / or left reference lines are relevant, and only samples from the most recent (or neighboring) top (or above) and / or left reference lines need to be buffered / stored for intra prediction of the current block. When MRLS is applied to an intra-coded block, 4 top (or above) reference lines and 4 left reference lines can be used for intra prediction. Therefore, the buffer for storing adjacent reference samples needs to be increased by 3 times. In addition, for a super block, the length of the buffer can be as long as the length of the reference lines located at the boundary of the super block, and in some cases, it can be equal to the picture width in a hardware decoder. Therefore, a 3-fold increase in buffer size is a heavy burden on the hardware decoder.
[0154] Embodiments of the present application describe various embodiments for improving the low-storage design and / or signaling of a multi-reference line selection scheme for intra prediction in video encoding and / or decoding, thereby solving at least one of the problems / difficulties discussed above.
[0155] In various embodiments, refer to Figure 17 , method 1700 for multi-reference line intra prediction in video decoding. Method 1700 may include some or all of the following steps: step 1710, receiving an encoded video bitstream of a current block by a device including a memory storing instructions and a processor communicating with the memory; step 1720, extracting, by the device, from the encoded video bitstream a parameter indicating a non-adjacent reference line for intra prediction in the current block; step 1730, partitioning, by the device, the current block to obtain a plurality of sub-blocks; and / or step 1740, in response to a sub-block of the plurality of sub-blocks being located at the boundary of the current block, using, by the device, a top adjacent reference line as the value of all top non-adjacent reference lines of the sub-block.
[0156] In some implementations, the current block may be referred to as a block, and the plurality of sub-blocks obtained by partitioning the current block may be referred to as a plurality of coded blocks. In some other implementations, step 1720 may include: extracting, by the device, from the encoded video bitstream a parameter indicating a reference line for intra prediction in the current block.
[0157] In various embodiments of the embodiments of the present application, the size of a block (such as, but not limited to, a coded block, a prediction block, or a transform block) may refer to the width or height of the block. The width or height of the block may be an integer in pixels.
[0158] In various embodiments of the embodiments of the present application, the size of a block (such as, but not limited to, a coding block, a prediction block, or a transform block) may refer to the area size of the block. The area size of the block may be an integer in pixels calculated by multiplying the width of the block by the height of the block.
[0159] In some different embodiments of the embodiments of the present application, the size of a block (such as, but not limited to, a coding block, a prediction block, or a transform block) may refer to the minimum value of the width or height of the block, the minimum value of the height or width of the block, or the aspect ratio of the block. The aspect ratio of the block may be calculated by dividing the width of the block by the height, or may be calculated by dividing the height of the block by the width.
[0160] In an embodiment of the present application, the reference row index indicates a reference row among a plurality of reference rows. In various embodiments, if the reference row index of a block is 0, it may indicate an adjacent reference row of the block, and the adjacent reference row is also the nearest reference row of the block. For example, in the block (1602) in Figure 16 , the top reference row (1610) is the top adjacent reference row of the block (1602) and also the top nearest reference row of the block; the left reference row (1618) is the left adjacent reference row of the block (1602) and also the left nearest reference row of the block. If the reference row index of a block is greater than 0, it indicates a non - adjacent reference row of the block, and the non - adjacent reference row is also a non - nearest reference row of the block. For example, in the block (1602) in Figure 16 , if the reference row index is 1, it may indicate the top reference row (1608) and / or the left reference row (1616); if the reference row index is 2, it may indicate the top reference row (1606) and / or the left reference row (1614); and / or if the reference row index is 3, it may indicate the top reference row (1604) and / or the left reference row (1612).
[0161] Referring to step 1710, the device may be Figure 5 the electronic device (530) in Figure 8 or the video decoder (810) in Figure 6 . In some implementations, the device may be the decoder (633) in the encoder (620) in Figure 5 . In other implementations, the device may be a part of the electronic device (530) in Figure 8 , a part of the video decoder (810) in Figure 6 , or a part of the decoder (633) in the encoder (620) in Figure 8 . The encoded video bitstream may be the encoded video sequence in Figure 6 or Figure 7 the intermediate encoded data in
[0162] In some implementations, a block may be referred to as a superblock. A superblock may refer to the largest coding block, such as but not limited to a coding tree block (CTB) and / or a largest coding unit (LCU). In some other implementations, a superblock may refer to a predetermined block size, such as but not limited to 32×32, 64×64, 128×128, and / or 256×256.
[0163] Referring to step 1720, the device may extract parameters from the encoded video bitstream, and the parameters may be used for MR LS and indicate the reference lines for intra prediction in the block. In some implementations, the reference line indicated by the parameter may be one of N reference lines, where N is an integer greater than 1. In some other implementations, the reference line indicated by the parameter may be one of N non - adjacent reference lines.
[0164] In some implementations, method 1700 may further include: using, by the device, the left - adjacent reference line of the block. In some other implementations, the left - adjacent reference line may be one of N left - hand reference lines of the block.
[0165] Referring to step 1730, the device may partition the block to obtain a plurality of coding blocks. In some implementations, the device may partition the block to obtain a coding block partition tree. The coding block partition tree may include a plurality of coding blocks.
[0166] Referring to step 1740, for a coding block from among the plurality of coding blocks, when the coding block is at the boundary of the block, the device may use the top - adjacent reference line as the value of all the top non - adjacent reference lines of the coding block. In some implementations, the boundary (or boundaries) of the block (e.g., superblock) may refer to only the top boundary of the block, or only the left boundary of the block, or the left boundary and the top boundary of the block.
[0167] In various embodiments, to reduce the memory size for storing samples from multiple reference lines, when the current intra - coded block is at the boundary of the superblock, the samples in N left - hand reference lines (columns) may be used for intra prediction, and the samples in the nearest (or adjacent) upper (or top) reference line may be used for intra prediction.
[0168] N is a positive integer greater than 1, such as 2, 3, or 4. For example, when N = 3, for a coding block located at the boundary of the superblock, three left - hand reference lines and the adjacent top reference line may be used for intra prediction. In some implementations, the three left - hand reference lines may be the three nearest left - hand reference lines of the block, such as Figure 16 the reference lines (1618, 1616, and 1614) in
[0169] In some implementations, the reference line index may be signaled and encoded into the bitstream, regardless of whether the current intra - coded block is at the boundary of the superblock.
[0170] Referring again to step 1740, the step of using the top adjacent reference line as the value of all top non - adjacent reference lines of the coding block may include: copying the samples from the top adjacent reference line to all other top non - adjacent reference lines, such that the memory size for storing the samples from all other top non - adjacent reference lines is significantly reduced.
[0171] In some implementations, if the in - frame coding block within the current frame is located at the boundary of a block (e.g., a super - block), and the reference line index indicates a non - zero (or non - adjacent) reference line for in - frame prediction of the current coding block, then the samples in the top (or upper) non - zero reference line are derived by copying from the nearest top (or upper) reference line.
[0172] As an example, in Figure 18 , the coding block (also referred to as an encoded block or a code block) (1802) is located at the top boundary (1830) and the left boundary (1840) of a block (e.g., a super - block). As Figure 18 shown, the top boundary (1830) and the left boundary (1840) of the super - block may be indicated by thick lines. The top - left samples (A - 4, A - 3, A - 2, and / or A - 1) and the top samples (A+0, …, A+N) in the nearest upper reference line (1810) of the coding block (1802) are copied to the upper non - zero reference lines (1808, 1806, and / or 1804). In some implementations, multiple left reference lines (1818, 1816, 1814, and / or 1812) may be allowed to store their respective samples.
[0173] In the above - described various implementations, taking the top adjacent reference line as an example, a similar scheme may apply to the left adjacent reference line, where, in order to improve the low - storage design, the samples in the left non - zero reference line may be derived from the nearest (or adjacent) left reference line, for example, by copying the left adjacent reference line.
[0174] In various embodiments, optionally, method 1700 may further include: dividing the coding block by the device to obtain a plurality of transform blocks. In some implementations, the device may divide the coding block to obtain a transform - block partition tree including a plurality of transform blocks. In some implementations, optionally, method 1700 may include: in response to a first transform block among the plurality of transform blocks being located at the top boundary of the coding block, using the top adjacent reference line as the value of all top non - adjacent reference lines of the first transform block by the device; and / or in response to a second transform block among the plurality of transform blocks not being located at the top boundary of the coding block, using, by the device, the reference line indicated by the parameters of the second transform block for the second transform block. In some implementations, the reference line indicated by the parameters may be one of the N top reference lines of the second transform block.
[0175] In some implementations, a current coding block may be split into more than one transform block (TB) and / or more than one transform unit (TU). When the current coding block is at the boundary of a superblock and the reference line index indicates a non-zero reference line for the current coding block, only the samples in the nearest upper reference line are available for intra prediction of the TU located at the boundary of the superblock; when the current coding block is at the boundary of a superblock and the reference line index indicates a non-zero reference line for the current coding block, the samples in the non-zero upper reference lines are still available for the TUs not located at the boundary of the superblock.
[0176] As an example, referring to Figure 19 , the thick solid line (1930) indicates the top boundary of the superblock. Multiple transform units (TU1, TU2, TU3, and / or TU4) (1904, 1906, 1908, and 1910) are located in a coding block. TU1 (1904) and TU2 (1906) are located at the boundary of the superblock (indicated by the horizontal thick solid line (1930)), while TU3 (1908) and TU4 (1910) are not located at the top boundary of the superblock. Therefore, only the samples in the nearest upper reference line are available for TU1 and TU2, and the samples in the non-zero reference lines are still available for TU3 and TU4. Similarly, in some other implementations, when the transform block within a coding block is at the left boundary of the superblock, only the samples in the nearest left reference line are available for the transform block.
[0177] The embodiments of the present application can be used alone or in any order combination. In addition, each method (or embodiment), encoder, and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. The embodiments of the present application can be separately applied to more than one color component, or can be applied together to more than one color component.
[0178] The above technology can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 20 shows a computer system (2600) suitable for implementing certain embodiments of the disclosed subject matter.
[0179] The computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be subject to mechanisms such as assembly, compilation, linking, or the like to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode execution, etc.
[0180] Instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0181] Figure 20 The components of the illustrated computer system (2600) are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present application. The configuration of the components should also not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system (2600).
[0182] The computer system (2600) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to one or more human users through inputs such as, for example: tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, clapping), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, captured images from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0183] The human-machine interface input devices may include one or more of the following (only one of each is shown): keyboard (2601), mouse (2602), touchpad (2603), touch screen (2610), data glove (not shown), joystick (2605), microphone (2606), scanner (2607), camera (2608).
[0184] The computer system (2600) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of the touch screen (2610), data glove (not shown) or joystick (2605), but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers (2609), headphones (not depicted)), visual output devices (e.g., screens (2610) including CRT screens, LCD screens, plasma screens, OLED screens, each screen having or not having touch screen input functionality, each screen having or not having tactile feedback functionality, some of which are capable of outputting two-dimensional visual output or output beyond three dimensions through means such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted)).
[0185] The computer system (2600) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (2620) with media such as CD / DVD (2621), thumb drives (2622), removable hard disk drives or solid state drives (2623), traditional magnetic media such as magnetic tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), etc.
[0186] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0187] The computer system (2600) may also include an interface (2654) to one or more communication networks (2655). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local area network, a wide area network, a metropolitan area network, vehicle and industrial networks, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial television broadcasting, vehicle and industrial networks including CAN bus, etc. Some networks typically require an external network interface adapter attached to certain common data ports or peripheral buses (2649) (e.g., the USB port of the computer system (2600)); as described below, other network interfaces are typically integrated into the kernel of the computer system (2600) by attaching to the system bus (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system). The computer system (2600) may communicate with other entities using any of these networks. Such communication may be one-way reception only (e.g., television broadcasting), one-way transmission only (e.g., CAN bus connected to certain CANBus devices), or two-way, for example, using a local area network or a wide area network digital network to connect to other computer systems. As described above, certain protocols and protocol stacks may be used on each of those networks and network interfaces.
[0188] The above-described human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the kernel (2640) of the computer system (2600).
[0189] The kernel (2640) may include one or more central processing units (CPUs) (2641), a graphics processing unit (GPU) (2642), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (2643), a hardware accelerator (2644) for certain tasks, a graphics adapter (2650), etc. These devices, as well as a read-only memory (ROM) (2645), a random access memory (2646), and an internal mass storage (2647) such as an internal hard disk drive, SSD, etc., which are not accessible to users, may be connected via a system bus (2648). In some computer systems, the system bus (2648) can be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices can be directly attached to the system bus (2648) of the kernel or attached to the system bus (2648) of the kernel via a peripheral bus (2649). In one example, a screen (2610) can be connected to the graphics adapter (2650). The architecture of the peripheral bus includes PCI, USB, etc.
[0190] The CPU (2641), GPU (2642), FPGA (2643), and accelerator (2644) can execute certain instructions, which can be combined to form the above-mentioned computer code. The computer code can be stored in the ROM (2645) or the RAM (2646). Transitional data can also be stored in the RAM (2646), while permanent data can be stored, for example, in the internal mass storage (2647). Fast storage and retrieval to any storage device can be achieved by using a cache, which can be closely associated with one or more CPUs (2641), GPUs (2642), mass storage (2647), ROM (2645), RAM (2646), etc.
[0191] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be media and computer code that are specifically designed and constructed for the purposes of the embodiments of this application, or the medium and the computer code can be of the types that are well-known and available to those skilled in the field of computer software.
[0192] As a non-limiting example, a computer system having an architecture (2600), particularly a core (2640), can provide functionality due to software executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) included in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage introduced above, as well as certain non-transitory memories of the core (2640), such as on-core mass memory (2647) or ROM (2645). The software implementing the various embodiments of the embodiments of the present application can be stored in such devices and executed by the core (2640). Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause the core (2640), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute the specific processes or specific parts of the specific processes described herein, including defining data structures stored in RAM (2646) and modifying such data structures according to the processes defined by the software. In addition or as an alternative, a computer system can provide functionality due to logic hardwired or otherwise embodied in circuitry (e.g., accelerator (2644)), which can replace the software or operate together with the software to execute the specific processes or specific parts of the specific processes described herein. In appropriate cases, portions referring to software can include logic, and vice versa. In appropriate cases, portions referring to computer-readable media can include circuitry (e.g., integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or including both. Embodiments of the present application include any suitable combination of hardware and software.
[0193] Although a particular invention is described with reference to illustrative embodiments, the description is not meant to be limiting. Various modifications to the illustrative embodiments of the invention and additional embodiments will be apparent to those of ordinary skill in the art from this description. Those skilled in the art will readily recognize that these and various other modifications can be made to the exemplary embodiments illustrated and described herein without departing from the spirit and scope of the invention. Accordingly, it is contemplated that the appended claims will cover any such modifications and alternative embodiments. Certain ratios in the figures may be exaggerated while others may be minimized. Accordingly, the disclosure and the figures should be regarded as illustrative, rather than restrictive.
Claims
1. A method for multi-reference line intra prediction in video decoding, characterized in that, The method includes: Receiving, by a device, an encoded video bitstream of a current block, the device including a memory storing instructions and a processor communicating with the memory; Extracting, by the device, parameters from the encoded video bitstream, the parameters indicating a non-adjacent reference row for intra prediction in the current block; Partitioning, by the device, the current block to obtain a plurality of sub-blocks; and In response to a sub-block of the plurality of sub-blocks being located at a boundary of the current block and the multi-reference row selection being applied to the current block, using a top adjacent reference row as values of all top non-adjacent reference rows of an encoded block.
2. The method according to claim 1, wherein The method further includes: Using, by the device, a left adjacent reference row of the sub-block.
3. The method according to claim 1, wherein The current block includes at least one of the following: a super block, a largest coding block, a coding tree block CTB, a largest coding unit LCU, a predetermined block having a predetermined size.
4. The method according to claim 1, wherein The boundary of the current block includes one of the following: a top boundary of the current block, a left boundary of the current block, or a left boundary and a top boundary of the current block.
5. The method according to claim 1, characterized in that The using of the top adjacent reference row as values of all top non-adjacent reference rows of the encoded block includes: Copying samples from the top adjacent reference row to all other top non-adjacent reference rows.
6. The method according to claim 5, characterized in that The method further includes: Partitioning, by the device, the sub-block to obtain a plurality of transform blocks.
7. The method according to any one of claims 1 to 6, characterized in that The method further includes: In response to a first transform block of the plurality of transform blocks being located at a top boundary of the sub-block, using, by the device, the top adjacent reference row as values of all top non-adjacent reference rows of the first transform block; and In response to a second transform block of the plurality of transform blocks not being located at a top boundary of the sub-block, using, by the device, a reference row indicated by parameters of the second transform block for the second transform block.
8. An apparatus for multi-reference line intra prediction in video decoding, characterized in that, The apparatus includes: A memory storing instructions; and A processor communicating with the memory, wherein when the processor executes the instructions, the processor is configured to cause the apparatus to perform the method according to any one of claims 1 to 7.
9. A non-transitory computer-readable storage medium storing instructions, characterized in that, When the instructions are executed by the processor, the instructions are configured to cause the processor to perform the method according to any one of claims 1 to 7.
10. A method for storing a video stream, characterized in that The video bitstream is decoded based on the method according to any one of claims 1 to 7.