Video code stream coding method and device, video code stream decoding method and device and storage medium
By dividing the intra mode list into multiple intra mode sets in video encoding technology, and encoding and decoding using set index and pattern index, the problem of intra prediction mode decoding in the prior art is solved, and more efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202510161102.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2021-12-29
- Filing Date
- 2022-01-18
- Publication Date
- 2025-05-13
AI Technical Summary
Existing video encoding technologies have problems with inefficiency in intra prediction mode decoding, especially when dealing with multi-directional prediction, more bits are required during encoding and decoding, resulting in increased bandwidth and storage space requirements.
By dividing the intra mode list into multiple intra mode sets and using set indexes and pattern indexes to encode and decode video code streams, the encoding method of intra prediction mode is optimized and unnecessary bit use is reduced.
Improves the efficiency of video encoding and decoding, reduces bandwidth and storage space requirements, and enhances the performance of encoding and decoding processes.
Smart Images

Figure CN119996659A_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application with an application date of January 18, 2022, Chinese patent application number 202280006553.0, and invention name “Method, device and storage medium for intra-frame prediction mode decoding”. Technical Field
[0002] The present application relates to video encoding and / or decoding technology, in particular, to improved design and writing of intra-frame prediction mode encoding, and more specifically to a method, device and storage medium for intra-frame prediction mode encoding. Background Art
[0003] The background technology description provided herein is for the purpose of generally describing the contents of the embodiments of the present application. Certain works of the inventors (i.e., works described in this background technology section) and the contents in the specification concerning certain prior art that has not yet become the prior art before the filing date are not considered to be the prior art with respect to the embodiments of the present application, whether explicitly or implicitly.
[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each picture having a spatial size of, for example, 1920×1080 luminance samples and associated fully sampled or subsampled chrominance samples. The series of pictures may have a fixed or variable picture rate (alternatively, referred to as a frame rate) of, for example, 60 pictures per second or 60 frames per second. Uncompressed video has specific bit rate requirements for streaming or data processing. For example, a video with a pixel resolution of 1920×1080, a frame rate of 60 frames per second, and a chrominance subsampling of 4:2:0 with 8 bits per pixel per color channel requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video requires more than 600 GB of storage space.
[0005] One purpose of video encoding and decoding can be to reduce redundancy in an uncompressed input video signal by compression. Compression can help reduce the above-mentioned bandwidth and / or storage space requirements, in some cases by two orders of magnitude or more. Lossless compression and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal through a decoding process. Lossy compression refers to a coding / decoding process in which the original video information is not fully retained during encoding and the original video information is not fully restored during decoding. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal usable for the intended application despite the loss of some information. In the case of video, lossy compression is widely used in many applications. The amount of distortion that can be tolerated depends on the application. For example, users of some consumer video streaming applications can tolerate higher distortion than users of movie or television broadcast applications. The compression rate achievable by a particular encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortions generally allow the encoding algorithm to produce higher losses and higher compression rates.
[0006] Video encoders and decoders may utilize techniques from multiple categories and steps including, for example, motion compensation, Fourier transforms, quantization, and entropy coding.
[0007] Video codec techniques may include a technique known as intra-coding. In intra-coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. When all blocks of samples are encoded in intra-mode, the picture may be referred to as an intra-picture. Intra-pictures and their derivatives (e.g., independent decoder refresh pictures) may be used to reset the decoder state and may therefore be used as the first picture in an encoded video stream and video session, or as a still image. The samples of the block after intra-prediction may then be transformed in the frequency domain, and the transform coefficients so generated may be quantized prior to entropy coding. Intra-prediction represents a technique for minimizing sample values in the pre-transform domain. In some cases, the smaller the DC value after transformation and the smaller the AC coefficient, the fewer bits are needed to represent the entropy coded block at a given quantization step size.
[0008] Traditional intra-frame coding, such as known from coding techniques such as MPEG-2, does not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to encode / decode blocks based on surrounding sample data and / or metadata obtained, for example, during spatially adjacent encoding and / or decoding, preceding the data block encoded or decoded within the frame in decoding order. Such techniques are hereinafter referred to as "intra-frame prediction" techniques. It should be noted that, at least in some cases, intra-frame prediction uses only reference data from the current picture being reconstructed, and does not use reference data from other reference pictures.
[0009] Intra-frame prediction can take many different forms. When more than one such technique can be used in a given video coding technique, the technique in use may be referred to as an intra-frame prediction mode. One or more intra-frame prediction modes may be provided in a particular codec. In some cases, a mode may have sub-modes, and / or may be associated with various parameters, and the mode / sub-mode information and intra-frame coding parameters for a video block may be encoded separately or included together in a mode codeword. Which codeword is used for a given mode, sub-mode, and / or parameter combination may have an impact on the coding efficiency gain through intra-frame prediction, and the entropy coding technique used to convert the codeword into a bitstream may also have an impact on it.
[0010] H.264 introduced a certain intra-frame prediction mode, which was improved in H.265 and further improved in new coding techniques such as the Joint Exploration Model (JEM), the next generation video coding (Versatile Video Coding, VVC), and the Benchmark Set (BMS). Typically, for intra-frame prediction, the values of neighboring samples that have become available can be used to form a prediction block. For example, the available values of a particular set of neighboring samples can be copied to the prediction block along some directions and / or lines. A reference to the direction used can be encoded in the bitstream, or it can be predicted itself.
[0011] refer to Figure 1A , a subset of 9 prediction directions specified in the 33 possible intra-frame prediction directions of H.265 (corresponding to the 33 angular modes of the 35 intra-frame modes specified in H.265) is depicted at the bottom right. The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction along which the neighboring samples are used to predict the sample at 101. For example, arrow (102) indicates that sample (101) is predicted from one or more neighboring samples at a 45 degree angle to the horizontal direction at the upper right. Similarly, arrow (103) indicates that sample (101) is predicted from one or more neighboring samples at a 22.5 degree angle to the horizontal direction at the lower left of sample (101).
[0012] Still referring to Figure 1, a square block (104) of 4×4 samples is depicted in the upper left (indicated by bold dashed lines). The square block (104) contains 16 samples, each sample is labeled with "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (starting from the top) and the first sample in the X dimension (starting from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the Y dimension and the X dimension. Since the size of the block is 4×4 samples, S44 is in the lower right corner. Example reference samples following a similar numbering scheme are also shown. The reference samples are labeled with R and their Y position (e.g., row index) and X position (column index) relative to the block (104). In H.264 and H.265, prediction samples adjacent to the block being reconstructed are used.
[0013] The intra picture prediction of block 104 may start by copying reference sample values from neighboring samples according to a signaled prediction direction. For example, assuming that the encoded video code stream includes signaling indicating the prediction direction of arrow (102) for this block 104, i.e., predicting samples from one or more prediction samples at the top right and at a 45 degree angle to the horizontal direction. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted based on reference sample R08.
[0014] In some cases, the values of multiple reference samples may be combined, such as by interpolation, in order to calculate the reference sample, especially when the direction is not divisible by 45 degrees.
[0015] As video coding technology continues to develop, the number of possible directions increases. For example, in H.264 (2003), nine different directions can be used for intra-frame prediction. In H.265 (2013), this increased to 33 directions, and in the embodiments of the present application, JEM / VVC / BMS can support up to 65 directions. Experimental studies have been conducted to help identify the most appropriate intra-frame prediction directions, and some techniques in entropy coding can be used to encode those most appropriate directions with a small number of bits, accepting a certain bit cost for the direction. In addition, the direction itself can sometimes be predicted from the adjacent directions used in the intra-frame prediction of the adjacent blocks that have been decoded.
[0016] Figure 1B A schematic diagram (180) is shown depicting 65 intra prediction directions according to JEM to illustrate the increase in the number of prediction directions in various coding techniques developed over time.
[0017] The manner in which the bits representing the intra prediction direction in the coded video bitstream are mapped to the prediction direction may vary between video coding techniques; for example, it may range from a simple direct mapping of the prediction direction to the intra prediction mode, to a codeword, to a complex adaptive scheme involving the most probable mode, and similar techniques. However, in all cases, for intra prediction, there may be certain directions that are statistically less likely to occur in the video content than certain other directions. Since the goal of video compression is to reduce redundancy, in a well-designed video coding technique, those less likely directions will be represented by more bits than the more likely directions.
[0018] Inter-picture prediction or inter-prediction may be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or portion thereof (reference picture) may be used to predict a newly reconstructed picture or picture portion (e.g., block) after being spatially shifted in a direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, with the third dimension indicating the reference picture being used (similar to the time dimension).
[0019] In some video compression techniques, the current MV applicable to a region of sample data may be predicted based on other MVs, for example, other MVs related to other regions of sample data that are spatially adjacent to the region being reconstructed and that precede the current MV in decoding order. This can greatly reduce the total amount of data required to encode the MV by eliminating redundancy in the related MVs, thereby increasing compression efficiency. MV prediction can work effectively, for example, because when encoding an input video signal obtained from a camera (called natural video), there is a statistical probability that a larger area than the area to which a single MV applies moves in a similar direction in a video sequence, so in some cases, similar motion vectors derived from MVs of adjacent areas can be used to predict the larger area. This makes the actual MV for a given region similar or identical to the MV predicted from the surrounding MVs. Further, after entropy coding, the MV can be represented with fewer bits than when the MV is encoded directly (rather than predicted from adjacent MVs). In some cases, MV prediction can be an example of losslessly compressing a signal (i.e., an MV) derived from an original signal (i.e., a sample stream). In other cases, the MV prediction itself may be lossy, for example due to round-off errors when computing the prediction value from multiple surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, "High Efficiency Video Coding", December 2016). Among the multiple MV prediction mechanisms specified by H.265, this article describes a technique referred to below as "spatial merging".
[0021] Specifically, refer to Figure 2 , the current block (201) includes samples that have been discovered by the encoder during the motion search process, and the samples can be predicted based on the previous block of the same size that has generated a spatial offset. In addition, the MV can be derived from metadata associated with one or more reference pictures instead of encoding the MV directly. For example, the MV associated with any of the five surrounding samples A0, A1 and B0, B1, B2 (corresponding to 202 to 206, respectively) is used to derive the MV from the metadata of the nearest reference picture (in decoding order). In H.265, MV prediction can use prediction values from the same reference picture that is being used by neighboring blocks. Summary of the invention
[0022] The embodiments of the present application describe various embodiments of methods, devices, and computer-readable storage media for video encoding and / or decoding.
[0023] In some embodiments, an embodiment of the present application provides a video code stream encoding method, the method comprising:
[0024] A device receives a video code stream of a block in an intra prediction mode, the device comprising a memory storing instructions and a processor in communication with the memory;
[0025] The device divides the intra-mode list into a plurality of intra-mode sets for the block based on mode information of each intra-mode in the intra-mode list, the intra-mode list corresponding to an intra-prediction mode of at least one neighboring block of the block;
[0026] The device extracts, from the video code stream, a set index indicating an intra-mode set from the multiple intra-mode sets;
[0027] The device extracts, from the video code stream, a mode index indicating an intra-frame prediction mode from the one intra-frame mode set; and
[0028] The device encodes the video code stream into an encoded video code stream based on the set index and the mode index.
[0029] In some embodiments, an embodiment of the present application provides a method for decoding a video code stream, the method comprising:
[0030] Receiving an encoded video stream for a block by a device, the device comprising a memory storing instructions and a processor in communication with the memory;
[0031] The device divides the intra-mode list into a plurality of intra-mode sets for the block based on mode information of each intra-mode in the intra-mode list, the intra-mode list corresponding to an intra-prediction mode of at least one neighboring block of the block;
[0032] The device extracts, from the encoded video code stream, a set index indicating an intra-mode set from the plurality of intra-mode sets;
[0033] The device extracts a mode index indicating an intra-frame prediction mode from the one intra-frame mode set from the encoded video code stream; and
[0034] The apparatus determines the intra prediction mode of the block based on the set index and the mode index.
[0035] In some embodiments, an embodiment of the present application provides a device for encoding a video code stream, the device comprising:
[0036] a memory for storing instructions; and
[0037] A processor in communication with the memory, wherein when the processor executes the instruction, the processor is configured to cause the device to perform the video code stream encoding method described in the embodiment of the present application.
[0038] In some embodiments, an embodiment of the present application provides a device for decoding a video code stream, the device comprising:
[0039] a memory for storing instructions; and
[0040] A processor in communication with the memory, wherein when the processor executes the instruction, the processor is configured to cause the device to perform the method for decoding a video code stream described in the embodiment of the present application.
[0041] In some embodiments, an embodiment of the present application provides a non-temporary computer-readable storage medium storing instructions. When the instructions are executed by a processor, the instructions are configured to cause the processor to execute the method for decoding a video stream described in the embodiment of the present application, or to execute the method for encoding a video stream described in the embodiment of the present application.
[0042] In some embodiments, an embodiment of the present application provides a method for storing or sending a video stream, wherein the video stream is decoded based on the video stream decoding method described in the embodiment of the present application, or the video stream is generated based on the video stream encoding method described in the embodiment of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Further features, properties and various advantages of the subject matter of the embodiments of the present application will become more apparent through the following detailed description and accompanying drawings.
[0044] Figure 1A A schematic diagram showing an exemplary subset of intra prediction direction modes.
[0045] Figure 1B An illustration of exemplary intra prediction directions is shown.
[0046] Figure 2 A schematic diagram showing a current block and surrounding spatially merged candidate blocks for motion vector prediction in an example.
[0047] Figure 3 A schematic diagram showing a simplified block diagram of a communication system according to an example embodiment.
[0048] Figure 4 A schematic diagram showing a simplified block diagram of a video streaming system according to an example embodiment.
[0049] Figure 5 A schematic diagram showing a simplified block diagram of a video decoder according to an example embodiment.
[0050] Figure 6 A schematic diagram showing a simplified block diagram of a video encoder according to an example embodiment.
[0051] Figure 7 A block diagram of a video encoder according to another example embodiment is shown.
[0052] Figure 8 A block diagram of a video decoder according to another example embodiment is shown.
[0053] Fig. 9 The directional intra prediction mode according to an exemplary embodiment of the present application is shown.
[0054] Fig.10 A non-directional intra prediction mode according to an exemplary embodiment of an embodiment of the present application is shown.
[0055] Fig.11 The recursive intra prediction mode according to an exemplary embodiment of the present application is shown.
[0056] Fig.12 An intra prediction scheme based on different reference lines according to an exemplary embodiment of the present application is shown.
[0057] Fig.13An offset-based refinement for intra prediction according to an exemplary embodiment of the present application is shown.
[0058] Fig.14A Another schematic diagram of offset-based refinement for intra prediction according to an exemplary embodiment of the present application is shown.
[0059] Fig. 14B Another schematic diagram of offset-based refinement for intra prediction according to an exemplary embodiment of the present application is shown.
[0060] Fig.15 A flowchart of a method according to an exemplary embodiment of the present application is shown.
[0061] Fig.16 A schematic diagram of a computer system according to an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0062] The present invention will be described in detail below with reference to the accompanying drawings, which form a part of the present invention and illustrate specific examples of embodiments by way of illustration. However, please note that the present invention may be implemented in a variety of different forms, and therefore, the subject matter covered or claimed is intended to be construed as not limited to any embodiment set forth below. Please also note that the present invention may be embodied as a method, device, component or system. Therefore, embodiments of the present invention may, for example, take the form of hardware, software, firmware or any combination thereof.
[0063] Throughout the specification and claims, terms may have nuanced meanings that are suggested or implied by the context beyond the explicitly stated meanings. The phrases "in one embodiment" or "in some embodiments" used herein do not necessarily refer to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" used herein do not necessarily refer to different embodiments. Similarly, the phrases "in one embodiment" or "in some embodiments" used herein do not necessarily refer to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" used herein do not necessarily refer to different embodiments. For example, the claimed subject matter includes a combination of all or part of the exemplary embodiments / embodiments.
[0064] In general, terms can be understood at least in part from usage in context. For example, terms such as "and", "or" or "and / or" used herein can include various meanings, which can depend at least in part on the context in which these terms are used. Generally, if "or" is used to associate a list such as A, B or C, it is intended to represent A, B and C (here for inclusion) and A, B or C (here for exclusion). In addition, the terms "one or more" or "at least one" used herein, at least in part depending on the context, can be used to describe any feature, structure or characteristic in a singular sense, or can be used to describe a combination of features, structures or characteristics in a plural sense. Similarly, terms such as "one", "an" or "the" can also be understood to convey singular usage or to convey plural usage, which depends at least in part on the context. In addition, the term "based on" or "determined by..." can be understood to not necessarily be intended to convey a set of exclusive factors, but may allow the presence of other factors that are not necessarily explicitly described, which also depends at least in part on the context.
[0065] Figure 3 A simplified block diagram of a communication system (300) according to one embodiment of the present application is shown. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). Figure 3 In the example of , the first terminal device pair (310) and (320) can perform unidirectional data transmission. For example, the terminal device (310) can encode video data (for example, video data of a video picture stream collected by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data is transmitted in the form of one or more encoded video code streams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to restore the video picture, and display the video picture based on the restored video data. The unidirectional data transmission can be implemented in applications such as media services.
[0066] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, which can be implemented, for example, during a video conferencing application. For bidirectional data transmission, in one example, each of the terminal devices (330) and (340) can encode video data (e.g., video data of a video picture stream collected by the terminal device) for transmission to the other terminal device (330) and (340) via a network (350). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other terminal device (330) and (340), and can decode the encoded video data to restore the video picture, and can display the video picture on an accessible display device based on the restored video data.
[0067] exist Figure 3 In the example of, terminal device (310), terminal device (320), terminal device (330) and terminal device (340) can be implemented as a server, a personal computer and a smart phone, but the applicability of the basic principle of the embodiment of the present application may not be limited to this. The embodiment of the embodiment of the present application can be implemented on a desktop computer, a laptop computer, a tablet computer, a media player, a wearable computer, a dedicated video conferencing device, and / or the like. The network (350) represents any number or type of network that transmits encoded video data between terminal device (310), terminal device (320), terminal device (330) and terminal device (340), including, for example, a wired (wired) and / or wireless communication network. The communication network (350) can exchange data in a circuit switching channel, a packet switching channel, and / or other types of channels. Representative networks include telecommunication networks, local area networks, wide area networks and / or the Internet. For the purpose of this discussion, unless explicitly stated herein, the architecture and topology of the network (350) may be irrelevant to the operation of the embodiment of the present application.
[0068] As examples of applications for the disclosed subject matter, Figure 4 is a schematic diagram of a simplified block diagram of a video streaming system according to an example embodiment. Figure 4 The placement of the video encoder and the video decoder in a video streaming environment is shown. The disclosed subject matter is equally applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., and the like.
[0069] The video streaming system may include a video acquisition subsystem (413), which may include a video source (401), such as a digital camera, which creates an uncompressed video picture stream or image (402). In one example, the video picture stream (402) includes samples recorded by the digital camera of the video source 401. Compared to the encoded video data (404) (or the encoded video bitstream), the video picture stream (402) is depicted as a thick line to emphasize the high data volume of the video picture stream, and the video picture stream (302) can be processed by an electronic device (420), which includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), depicted as thin lines to emphasize the lower data volume compared to the uncompressed video picture stream (402), can be stored on the streaming server (405) for future use, or directly stored in a downstream video device (not shown). One or more streaming client subsystems, such as Figure 4 The client subsystem (406) and the client subsystem (408) in the streaming server (405) can access the streaming server (405) to retrieve the copy (407) and the copy (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and generates an uncompressed output video picture stream (411) that can be presented on a display (412) (e.g., a display screen) or other presentation device (not depicted). The video decoder 410 may be configured to perform some or all of the various functions described in the embodiments of the present application. In some streaming transmission systems, the encoded video data (404), the video data (407), and the video data (409) (e.g., a video bitstream) may be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T H.265. In an embodiment, the video coding standard under development is informally referred to as the next generation video coding (Versatile Video Coding, VVC). The disclosed subject matter may be used in the context of VVC, and may be used in other video coding standards.
[0070] It should be noted that the electronic device (420) and the electronic device (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).
[0071] In the following, Figure 5A block diagram of a video decoder (510) according to any embodiment of the present application is shown. The video decoder (510) may be arranged in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used to replace Figure 4 A video decoder (410) in an example of FIG.
[0072] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510). In the same or another embodiment, one encoded video sequence may be decoded at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. Each video sequence may be associated with multiple video frames or images. The encoded video sequence may be received from a channel (501), which may be a hardware / software link leading to a storage device storing the encoded video data or a streaming source sending the encoded video data. The receiver (531) may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective processing circuits (not depicted). The receiver (531) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) may be provided between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be located outside the video decoder (510) and separate from the video decoder (510) (not depicted). In still other applications, a buffer memory (not depicted) is provided outside the video decoder (510) to, for example, prevent network jitter, and another additional buffer memory (515) may be provided inside the video decoder (510) to, for example, handle broadcast timing. When the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer memory (515), or the buffer memory may be made smaller. In order to be used on a service packet network such as the Internet, a buffer memory (515) of sufficient size may be required, and the size of the buffer memory (515) may be relatively large. Such a buffer memory may be implemented to have an adaptive size and may be implemented at least partially in an operating system or a similar element (not depicted) outside the video decoder (510).
[0073] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510) and potential information for controlling a rendering device such as a display (512) (e.g., a display screen), which may or may not be an integral part of the electronic device (530), but may be coupled to the electronic device (530), such as Figure 5 As shown. The control information for the rendering device may be in the form of a Supplemental Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser (520) may parse / entropy decode the encoded video sequence received by the parser (520). The entropy encoding of the encoded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (520) may also extract information from the encoded video sequence, such as transform coefficients (eg, Fourier transform coefficients), quantizer parameter values, motion vectors, etc.
[0074] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515), thereby creating symbols (521).
[0075] Depending on the type of coded video picture or portion of coded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbol (521) may involve multiple different processing or functional units. Which units are involved and how they are involved can be controlled by the parser (520) through subgroup control information parsed from the coded video sequence. For simplicity, such subgroup control information flow between the parser (520) and the multiple processing or functional units described below is not depicted.
[0076] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units as described below. In actual implementations that operate under commercial constraints, many of these functional units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, the conceptual subdivision into multiple functional units is adopted below in the embodiments of the present application.
[0077] The first unit may include a sealer / inverse transform unit (551). The sealer / inverse transform unit (551) may receive quantized transform coefficients as symbols (521) and control information from the parser (520), including information indicating which type of inverse transform to use, block size, quantization factors / parameters, quantization scaling matrices, etc. The sealer / inverse transform unit (551) may output a block including sample values, which may be input into an aggregator (555).
[0078] In some cases, the output samples of the sealer / inverse transform (551) may belong to an intra-coded block; that is, a block that does not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may use surrounding block information that has been reconstructed and stored in a current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (555) may add the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the sealer / inverse transform unit (551) on a per-sample basis.
[0079] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-frame coded and potentially motion compensated block. In this case, the motion compensated prediction unit (553) may access the reference picture memory (557) to extract samples for inter-frame picture prediction. After the extracted samples are motion compensated according to the symbols (521) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (551) by the aggregator (555) (the output of unit 551 may be referred to as residual samples or residual signals), thereby generating output sample information. The extraction of prediction samples by the motion compensated prediction unit (553) from the address in the reference picture memory (557) may be controlled by a motion vector, and the motion vector may be provided to the motion compensated prediction unit (553) in the form of a symbol (521), which may have, for example, an X component, a Y component (offset) and a reference picture component (time). Motion compensation may also include interpolation of sample values extracted from the reference picture memory (557) when using sub-sample accurate motion vectors, and may also be associated with motion vector prediction mechanisms, etc.
[0080] The output samples of the aggregator (555) may be employed by various loop filtering techniques in a loop filter unit (556). The video compression techniques may include in-loop filter techniques that are controlled by parameters included in an encoded video sequence (also referred to as an encoded video bitstream) and that are available to the loop filter unit (556) as symbols (521) from the parser (520). However, in other embodiments, the video compression techniques may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, as well as to previously reconstructed and loop filtered sample values. A variety of types of loop filters may be included as part of the loop filter unit 556 in various orders, as will be described in further detail below.
[0081] The output of the loop filter unit (556) may be a sample stream that may be output to a rendering device (512) and stored in a reference picture memory (557) for future inter-picture prediction.
[0082] Once fully reconstructed, certain coded pictures may be used as reference pictures for future inter-picture prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) may become part of the reference picture memory (557), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.
[0083] The video decoder (510) may perform decoding operations according to a predetermined video compression technique employed in a standard such as ITU-T H.265 Recommendation. The coded video sequence may conform to the syntax specified by the video compression technique or standard used in the sense that the coded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technique or standard as the only tools available for use under the profile. In order to conform to the standard, the complexity of the coded video sequence may be within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further defined by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management signaled in the coded video sequence.
[0084] In some exemplary embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0085] Figure 6 6 is a block diagram of a video encoder (603) according to an exemplary embodiment disclosed in the present application. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., a transmission circuit). The video encoder (603) may be used to replace Figure 4 A video encoder (403) in an example.
[0086] The video encoder (603) can be used to obtain the video source (601) (not Figure 6 In an example of an electronic device (620) receiving video samples, a video source (601) can capture video images to be encoded by a video encoder (603). In another example, the video source (601) can be implemented as a part of the electronic device (620).
[0087] The video source (601) may provide a source video sequence in the form of a digital video sample stream to be encoded by a video encoder (603), the digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCB, RGB, XYZ, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) may be a storage device capable of storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures or images that are given motion when viewed sequentially. The picture itself may be constructed as a spatial pixel array, wherein each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be easily understood by a person skilled in the art. The following description focuses on the samples.
[0088] According to some exemplary embodiments, the video encoder (603) may encode and compress the pictures of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed constitutes a function of the controller (650). In some embodiments, the controller (650) may be functionally coupled to other functional units as described below and control the other functional units. For simplicity, the coupling is not depicted in the figure. The parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, etc.), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) can be used to have other suitable functions that are related to the video encoder (603) optimized for a certain system design.
[0089] In some exemplary embodiments, the video encoder (603) may be configured to operate in an encoding loop. As an oversimplified description, in one example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to the way the (remote) decoder creates sample data, even if the embedded decoder 633 processes the encoded video stream through the source encoder 630 without entropy coding (because in the video compression techniques considered in the disclosed subject matter, any compression between the symbols in the entropy coding and the encoded video code stream can be lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since the decoding of the symbol stream produces a bit-accurate result that is independent of the decoder location (local or remote), the contents in the reference picture memory (634) are also bit-accurately corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, e.g. due to channel errors) is used to improve encoding quality.
[0090] The operation of the "local" decoder (633) can be combined with the above Figure 5 The "remote" decoder described in detail for the video decoder (510) is identical. However, additional brief reference is made to Figure 5 , since the symbols are available and the entropy encoder (645) and the parser (520) are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633), in the encoder.
[0091] At this point, it can be observed that any decoder technology, except for the parsing / entropy decoding that may only exist in the decoder, must also exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter sometimes focuses on the decoder operation, which cooperates with the decoding part of the encoder. Therefore, the description of the encoder technology can be simplified because the encoder technology is mutually inverse to the decoder technology described comprehensively. Only some areas or aspects of the encoder are described in more detail below.
[0092] During operation, in some example implementations, the source encoder (630) may perform motion compensated predictive coding that predictively encodes an input picture by referencing one or more previously encoded pictures from a video sequence designated as "reference pictures." In this manner, the encoding engine (632) encodes the differences (or residuals) in color channels between pixel blocks of the input picture and pixel blocks of a reference picture that may be selected as a prediction reference for the input picture.
[0093] The local video decoder (633) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbol created by the source encoder (630). The operation of the encoding engine (632) may be a lossy process. When the encoded video data may be decoded at the video decoder ( Figure 6 When the video encoder (603) is decoded on a local (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that can be performed by the video decoder on the reference picture and can cause the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) can store a copy of the reconstructed reference picture locally that has common content (absent transmission errors) with the reconstructed reference picture that will be obtained by the far-end (remote) video decoder.
[0094] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may serve as appropriate prediction references for the new picture. The predictor (635) may operate pixel-by-pixel based on sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (635), it may be determined that the input picture may have prediction references taken from a plurality of reference pictures stored in the reference picture memory (634).
[0095] The controller (650) may manage encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.
[0096] The outputs of all the above functional units may be entropy encoded in an entropy encoder (645). The entropy encoder (645) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.
[0097] The transmitter (640) may buffer the encoded video sequence created by the entropy encoder (645) in preparation for transmission over a communication channel (660), which may be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (640) may combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).
[0098] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:
[0099] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow for different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those of ordinary skill in the art are aware of the variations of I pictures and their corresponding applications and features.
[0100] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict sample values for each block.
[0101] Bidirectional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using intra prediction or inter prediction, which uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0102] The source picture may typically be spatially subdivided into a plurality of coding blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined by the coding allocation applied to the corresponding picture of the block. For example, a block of an I picture may be non-predictively coded, or the block may be predictively coded (spatial prediction or intra-frame prediction) with reference to an already coded block of the same picture. A pixel block of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. A block of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction. For other purposes, the source picture or the intermediate processed picture may be subdivided into other types of blocks. The division of coding blocks and other types of blocks may or may not follow the same approach, as described in further detail below.
[0103] The video encoder (603) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (603) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.
[0104] In an exemplary embodiment, the transmitter (640) may transmit additional data when transmitting the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0105] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. For example, a particular picture being encoded / decoded may be divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.
[0106] In some exemplary embodiments, a bidirectional prediction technique may be used for inter-picture prediction. According to this bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that precede the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be jointly predicted by a combination of the first reference block and the second reference block.
[0107] In addition, merge mode technology can be used for inter-picture prediction to improve coding efficiency.
[0108] According to some exemplary embodiments of the embodiments of the present application, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, a picture in a video picture sequence is divided into coding tree units (CTUs) for compression, and the CTUs in the picture may have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU may include three parallel coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Each CTU may be recursively divided into one or more coding units (CUs) in a quadtree. For example, a 64×64 pixel CTU may be divided into a 64×64 pixel CU, or four 32×32 pixel CUs. Each of the one or more 32×32 blocks may be further divided into four 16×16 pixel CUs. In some exemplary embodiments, each CU may be analyzed during encoding to determine a prediction type for the CU among various prediction types, such as an inter-prediction type or an intra-prediction type. According to temporal and / or spatial predictability, the CU can be divided into one or more prediction units (PUs). Typically, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in the encoding (encoding / decoding) is performed in units of prediction blocks. The division of the CU into PUs (or PBs of different color channels) can be performed in various spatial modes. For example, a luminance or chrominance PB may include a matrix of values for samples (e.g., luminance values), and the samples are, for example, 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 samples, etc.
[0109] Figure 7A diagram of another exemplary embodiment of a video encoder (703) according to an embodiment of the present application is shown. The video encoder (703) is configured to receive a processed block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures and encode the processed block into an encoded picture that is part of an encoded video sequence. The exemplary video encoder (703) may be used instead of Figure 4 A video encoder (403) in an example.
[0110] For example, the video encoder (703) receives a matrix of sample values for a processing block, such as a prediction block of 8×8 samples, etc. The video encoder (703) then uses, for example, rate-distortion optimization (RDO) to determine whether to use intra mode, inter mode, or bidirectional prediction mode to best encode the processing block. When it is determined that the processing block is encoded in intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into an encoded picture; and when it is determined that the processing block is encoded in inter mode or bidirectional prediction mode, the video encoder (703) may use inter prediction or bidirectional prediction techniques to encode the processing block into an encoded picture, respectively. In some exemplary embodiments, merge mode may be used as an inter-picture prediction submode, in which motion vectors are derived from one or more motion vector predictors without the aid of encoded motion vector components external to the predictor. In some other exemplary embodiments, there may be motion vector components applicable to the subject block. Therefore, the video encoder (703) may include a submode that is not in Figure 7 Components explicitly shown in, such as a mode decision module for determining a prediction mode for a processing block.
[0111] exist Figure 7 In the example of FIG. 7 , the video encoder ( 703 ) includes Figure 7 The exemplary arrangement shown is an inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721) and an entropy encoder (725) coupled together.
[0112] The inter-frame encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block to one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture in display order), generate inter-frame prediction information (e.g., redundant information description according to an inter-frame coding technique, motion vectors, merge mode information), and calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique. In some examples, the reference picture is based on the encoded video information, using a decoded reference picture decoded by the decoding unit 633, which is embedded in the decoding unit 633. Figure 6 In the exemplary encoder 620 (shown as Figure 7 residual decoder 728, as described in further detail below).
[0113] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with an encoded block in the same picture, generate quantization coefficients after transformation, and in some cases also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). The intra encoder (722) can calculate an intra prediction result (e.g., a predicted block) based on the intra prediction information and a reference block in the same picture.
[0114] The general controller (721) can be configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines a prediction mode of a block and provides a control signal to a switch (726) based on the prediction mode. For example, when the prediction mode is an intra-frame mode, the general controller (721) controls the switch (726) to select an intra-frame mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select intra-frame prediction information and include the intra-frame prediction information in the bitstream; and when the prediction mode of the block is an inter-frame mode, the general controller (721) controls the switch (726) to select an inter-frame prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select inter-frame prediction information and include the inter-frame prediction information in the bitstream.
[0115] The residual calculator (723) may be configured to calculate the difference (residual data) between the received block and the block prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) may be configured to encode the residual data to generate a transform coefficient. For example, the residual encoder (724) may be configured to transform the residual data from the spatial domain to the frequency domain to generate a transform coefficient. The transform coefficient is then quantized to obtain a quantized transform coefficient. In various exemplary embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is used to perform an inverse transform and generate decoded residual data. The decoded residual data may be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) may generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) may generate a decoded block based on the decoded residual data and the intra prediction information. The decoded blocks are appropriately processed to generate decoded pictures, and the decoded pictures may be buffered in a memory circuit (not shown) and used as reference pictures.
[0116] The entropy encoder (725) may be configured to format the code stream to include the encoded blocks and perform entropy encoding. The entropy encoder (725) may be configured to include various information in the code stream. For example, the entropy encoder (725) may be configured to include general control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the code stream. When the block is encoded in the inter-frame mode or the merge sub-mode of the bidirectional prediction mode, the residual information may not be present.
[0117] Figure 8 FIG. 8 is a diagram of an exemplary video decoder (810) according to another embodiment of an embodiment of the present application. The video decoder (810) is used to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed picture. In one example, the video decoder (810) can be used instead of Figure 4 A video decoder (410) in an example of FIG.
[0118] exist Figure 8 In the example of FIG. 8 , the video decoder ( 810 ) includes Figure 8 The exemplary arrangement shown is an entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874) and an intra-frame decoder (872) coupled together.
[0119] The entropy decoder (871) can be used to reconstruct certain symbols from the encoded picture, which represent syntax elements that constitute the encoded picture. Such symbols may include, for example, a mode for encoding a block (e.g., intra mode, inter mode, bidirectional prediction mode, merge submode, or another submode), prediction information (e.g., intra prediction information or inter prediction information) that can identify certain samples or metadata used by the intra decoder (872) or the inter decoder (880) for prediction, respectively, residual information in the form of, for example, quantized transform coefficients, etc. In one example, when the prediction mode is inter or bidirectional prediction mode, the inter prediction information is provided to the inter decoder (880); and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information can be inverse quantized and provided to the residual decoder (873).
[0120] The inter-frame decoder (880) may be configured to receive the inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.
[0121] The intra decoder (872) may be configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0122] The residual decoder (873) may be configured to perform inverse quantization to extract dequantized transform coefficients, and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also use certain control information (including quantizer parameters (QP)) that may be provided by the entropy decoder (871) (data path not depicted as this is only low data volume control information).
[0123] The reconstruction module (874) may be configured to combine the residual output by the residual decoder (873) with the prediction result (which may be output by the inter-frame prediction module or the intra-frame prediction module, as appropriate) in the spatial domain to form a reconstructed block, which forms part of a reconstructed picture, which is part of the reconstructed video. It should be noted that other suitable operations such as deblocking operations may also be performed to improve visual quality.
[0124] It should be noted that the video encoder (403), video encoder (603) and video encoder (703) and video decoder (410), video decoder (510) and video decoder (810) may be implemented using any suitable technology. In some exemplary embodiments, the video encoder (403), video encoder (603) and video encoder (703) and video decoder (410), video decoder (510) and video decoder (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoder (403), video encoder (603) and video encoder (603) and video decoder (410), video decoder (510) and video decoder (810) may be implemented using one or more processors that execute software instructions.
[0125] Returning to the intra prediction process, in this process, samples in a block (e.g., a luma or chroma prediction block, or a coding block if it has not been further partitioned into prediction blocks) are predicted by samples of an adjacent line, the next adjacent line, or one or more other lines or a combination thereof to generate a prediction block. The residual between the prediction block and the actual block being encoded can then be processed by a transform and subsequent quantization. Various intra prediction modes can be made available, and parameters related to intra mode selection and other parameters can be written to the bitstream. For example, various intra prediction modes can involve one or more line positions for predicting samples, the direction of selecting prediction samples from one or more prediction lines, and other special intra prediction modes.
[0126] For example, a set of intra prediction modes (interchangeably referred to as "intra modes") may include a predetermined number of directional intra prediction modes. As described above with respect to the example implementation of FIG. 1, these intra prediction modes may correspond to a predetermined number of directions along which samples outside the block are selected as predictions for samples predicted in a particular block. In another specific example implementation, 8 main directional modes corresponding to angles from 45 degrees to 207 degrees relative to the horizontal axis may be supported and predefined.
[0127] In some other embodiments of intra prediction, in order to further exploit more types of spatial redundancy in directional textures, the directional intra mode can be further extended to an angle set with finer granularity. Fig. 9 As shown, the above 8-angle implementation can be configured to provide eight nominal angles, called V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, and for each nominal angle, a predetermined number (e.g., 7) of finer angles can be added. With such an extension, a larger total number (e.g., 56 in this example) of directional angles corresponding to the same number of predetermined directional intra-frame modes can be used for intra-frame prediction. The predicted angle can be represented by the nominal intra-frame angle plus the angle increment. For the specific example of having 7 finer angular directions for each nominal angle above, the angle increment can be a step size of -3 to 3 times 3 degrees.
[0128] The above-mentioned directional intra prediction may also be referred to as unidirectional intra prediction, which is different from the bidirectional intra prediction (also referred to as bidirectional intra prediction) described in the later part of the embodiments of the present application.
[0129] In some embodiments, instead of or in addition to the above-mentioned directional intra-frame mode, a predetermined number of non-directional intra-frame prediction modes may also be predefined and available. For example, 5 non-directional intra-frame modes called smooth intra-frame prediction modes may be specified. These non-directional intra-frame mode prediction modes may be specifically referred to as DC, PAETH, SMOOTH, SMOOTH_V and SMOOTH_H intra-frame modes. Fig.10 The prediction of samples for a particular block under these exemplary non-directional modes is shown in FIG. Fig.10A 4×4 block 1002 is shown that is predicted by samples from an upper neighboring line and / or a left neighboring line. A particular sample 1010 in the block 1002 may correspond to a sample 1004 directly above the upper neighboring line of the block 1002, a sample 1006 above and to the left of the intersection of the upper and left neighboring lines of the sample 1010, and a sample 1008 directly to the left of the left neighboring line of the block 1002 of the sample 1010. For the exemplary DC intra prediction mode, an average of the left neighboring sample 1008 and the upper neighboring sample 1004 may be used as a predictor for the sample 1010. For the exemplary PAETH intra prediction mode, the upper, left, and upper left reference samples 1004, 1008, and 1006 may be extracted, and then the value closest to (above + left - upper left) of the three reference samples may be set as the predicted value for the sample 1010. For the exemplary SMOOTH_V intra prediction mode, sample 1010 can be predicted by quadratic interpolation in the vertical direction of the upper left neighboring sample 1006 and the left neighboring sample 1008. For the exemplary SMOOTH_H intra prediction mode, sample 1010 can be predicted by quadratic interpolation in the horizontal direction of the upper left neighboring sample 1006 and the upper neighboring sample 1004. For the exemplary smooth intra prediction mode, sample 1010 can be predicted by the average of the quadratic interpolation in the vertical and horizontal directions. The above implementation of the non-directional intra mode is shown only as a non-limiting example. Other neighboring lines and other non-directional sample selections for predicting specific samples in the prediction block are also considered, as well as ways to combine prediction samples.
[0130] The encoder selects a specific intra-frame prediction mode from the directional or non-directional mode above to be written into the code stream at various coding levels (picture, slice, block, unit, etc.). In some exemplary embodiments, exemplary 8 nominal directional modes together with 5 non-angle smooth modes (13 options in total) can be written first. Then, if the mode written is one of the intra-frame modes of 8 nominal angles, the index is further written to indicate the selected angle increment to the corresponding nominal angle of writing. In some other example embodiments, all intra-frame prediction modes can be added with indexes together (for example, 56 directional modes plus 5 non-directional modes to produce 61 intra-frame prediction modes) for writing.
[0131] In some example embodiments, the example 56 or other number of directional intra prediction modes may be implemented using a unified directional predictor that projects each sample of the block to a reference subsample location and interpolates the reference samples via a 2-tap bilinear filter.
[0132] In some embodiments, in order to capture the attenuated spatial correlation of references on edges, additional filter modes, referred to as filter intra modes, may be designed. For these modes, in addition to samples outside the block, the predicted samples within the block may be used as intra prediction reference samples for some subblocks within the block. For example, these modes may be predefined and may be used for intra prediction of at least luminance blocks (or only luminance blocks). A predetermined number (e.g., five) of filter intra modes may be predesigned, each mode being represented by a set of n-tap filters (e.g., 7-tap filters) that reflect the correlation between samples in, for example, a 4×2 subblock and its n adjacent neighboring blocks. In other words, the weighting factors of the n-tap filters may be position-dependent. Taking an 8×8 block, a 4×2 subblock, and a 7-tap filter as an example, Fig.11 As shown, the 8×8 block 1102 can be divided into eight 4×2 sub-blocks. Fig.11 In the example, B0, B1, B2, B3, B4, B5, B6 and B7 are used to represent each sub-block. Fig.11 For sub-block B0, all neighboring blocks may have been reconstructed. However, for other sub-blocks, since some neighboring blocks are in the current block and may not have been reconstructed, the predicted values of the neighboring blocks are used as reference. For example, Fig.11 All neighboring blocks of the shown sub-block B7 are not reconstructed, so prediction samples of neighboring blocks, such as part of B4, B5 and / or B6, are used instead.
[0133] In some embodiments of intra prediction, one color component may be predicted using one or more other color components. The color component may be any one of the components in YCrCb, RGB, XYZ color space, etc. For example, it may be possible to predict a chrominance component (e.g., a chrominance block) from a luma component (e.g., a luma reference sample), referred to as Chroma from Luma (CfL). In some example embodiments, cross color prediction may only allow chrominance to be predicted from luma. For example, the chrominance samples in a chrominance block may be modeled as a linear function of the coincident reconstructed luma samples. CfL prediction may be implemented as follows:
[0134] CfL(α)=α×L AC +DC (1)
[0135] Among them, L ACrepresents the AC contribution of the luma component, α represents the parameters of the linear model, and DC represents the DC contribution of the chroma component. For example, an AC component is obtained for each sample of a block, and a DC component is obtained for the entire block. Specifically, the reconstructed luma samples can be subsampled to the chroma resolution, and then the average luma value (the DC of luma) can be subtracted from each luma value to form the AC contribution of luma. The AC contribution of luma is then used in the linear mode of equation (1) to predict the AC value of the chroma component. In order to approximate or predict the chroma AC component based on the luma AC contribution, instead of requiring the decoder to calculate the scaling parameters, an exemplary CfL implementation can determine the parameter α based on the original chroma samples and write it into the bitstream. This reduces the complexity of the decoder and produces a more accurate prediction. As for the DC contribution of the chroma component, in some example embodiments, it can be calculated using the intra-frame DC mode within the chroma component.
[0136] Returning to intra prediction, in some example embodiments, the prediction of samples in a coding block or prediction block can be based on one of a set of reference lines. In other words, multiple reference lines can be provided as options for selection for intra prediction, rather than always using the nearest neighbor line (e.g., the neighbor line immediately above the top or the neighbor line immediately to the left of the prediction block as shown in FIG. 1 above). This intra prediction implementation may be referred to as multiple reference line selection (MRLS). In these embodiments, the encoder determines which of the multiple reference lines is used to generate the intra predictor and writes it. On the decoder side, after parsing the reference line index, the reconstructed reference sample can be identified by looking up the specified reference line according to the intra prediction mode (e.g., directional, non-directional, and other intra prediction modes), thereby generating an intra prediction of the current intra prediction block. In some embodiments, the reference line index can be written into the coding block level, and only one of the multiple reference lines can be selected and used for intra prediction of a coding block. In some examples, multiple reference lines can be selected together for intra prediction. For example, multiple reference lines may be combined, averaged, interpolated, or in any other manner weighted or unweighted to generate a prediction. In some example embodiments, MRLS may be applied only to the luma component and may not be applied to the chroma components.
[0137] exist Fig.12 An example of a 4-reference line MRLS is depicted in FIG. Fig.12As shown in the example, the intra-frame coding block 1202 can be predicted based on one of the four horizontal reference lines 1204, 1206, 1208 and 1210 and the four vertical reference lines 1212, 1214, 1216 and 1218. Among these reference lines, 1210 and 1218 are adjacent reference lines. The reference lines can be indexed according to their distance from the coding block. For example, reference lines 1210 and 1218 can be referred to as zero reference lines, and other reference lines can be referred to as non-zero reference lines. Specifically, reference lines 1208 and 1216 can be referenced as the first reference line; reference lines 1206 and 1214 can be referenced as the second reference line; and reference lines 1204 and 1212 can be referenced as the third reference line.
[0138] In some embodiments, for a particular coding block, coding unit, prediction block, or prediction unit that is intra-coded, its intra mode needs to be written into the bitstream by one or more syntax elements. As described above, the number of possible intra prediction modes can be huge, and 62 intra prediction modes may be available: 56 directional intra prediction modes, 5 non-directional modes, and one chroma mode from luminance (e.g., only for chroma components). In order to write these intra prediction modes, a first syntax may be written to indicate which nominal angle or non-directional mode is equal to the nominal mode of the current block. Then, if the mode of the current block is a directional mode, a second syntax may be written to indicate which delta angle is equal to the delta angle of the current block. In some cases during video encoding and / or decoding, there may be a strong correlation between the intra prediction mode of the current block and its neighboring blocks.
[0139] In various embodiments, this correlation can be exploited to design more efficient syntax for intra-mode encoding. In some implementations, the available intra-prediction modes of the current block can be split into multiple intra-prediction mode sets based on the intra-prediction modes of its neighboring blocks. To obtain the intra-prediction mode of the current block, first, a set index can be written to indicate the set index of the intra-prediction mode of the current block; second, a mode index can be written to indicate the index of the intra-prediction mode within the mode set.
[0140] In various embodiments of the embodiments of the present application, "XYZ is written" may refer to XYZ being encoded into an encoded code stream during the encoding process; and / or, after the encoded code stream is sent from one device to another device, "XYZ is written" may refer to decoding / extracting XYZ from the encoded code stream during the decoding process.
[0141] For example, in some embodiments described above, the number of available intra prediction modes may include 62 different modes, including, for example, 56 directional intra prediction modes (e.g., 8 nominal directions with 7 fine angles in each nominal direction), 5 non-directional modes, and a mode from luma to chroma (for chroma components only). Once an intra mode is selected during the encoding process of a particular coding block, coding unit, prediction block, or prediction, the signaling corresponding to the selected intra mode needs to be included in the bitstream. The write syntax must be able to distinguish all 62 modes in some way. For example, the 62 modes can be written using a single syntax for 62 indexes, each index corresponding to a mode. In some other example embodiments, a syntax may be written to indicate which nominal angle or non-directional mode is used as the nominal mode in the current block, and then, if the nominal mode of the current block is a directional mode, another syntax may be additionally written to indicate which incremental angle is selected for the current block.
[0142] Since various syntaxes related to intra-frame coding usually occupy a large part of the code stream, and intra-frame mode selection must be frequently written, for example, at various coding levels, reducing the number of bits used for intra-frame mode writing becomes crucial in improving video coding efficiency. In practice, the use of various intra-frame prediction modes can follow certain statistical patterns, and this usage pattern can be used to design the index of intra-frame mode and writing syntax, so that writing efficiency can be improved. In addition, on average, there may be some correlation between intra-frame mode selections between blocks. This correlation can be obtained offline on a statistical basis and considered in the syntax design for writing the selection of intra-frame mode. The goal is to reduce the number of bits written in the coded code stream on average. For example, some general statistics can show that there may be a strong correlation between the best intra-frame prediction mode of the current block and its neighboring blocks. This correlation can be exploited when designing syntax for intra-frame mode coding.
[0143] In some embodiments, in order to improve video encoding / decoding performance, offset-based refinement for intra prediction (ORIP) may be used after generating intra prediction samples. When ORIP is applied, the prediction samples are refined by adding offset values.
[0144] like Fig.13As shown, intra prediction is performed based on reference samples (1330). The reference samples may include samples from one or more left reference lines (1312) and / or one or more top reference lines (1310). Offset-based intra prediction refinement (ORIP) (1350) may use neighboring reference samples to generate offset values. In some embodiments, the neighboring reference samples used for ORIP may be the same set as the reference samples used for intra prediction. In some other embodiments, the neighboring reference samples used for ORIP may be a different set than the reference samples used for intra prediction.
[0145] In reference Fig.14A and Fig. 14B In some embodiments, ORIP can be performed at the 4×4 sub-block level. For each 4×4 sub-block (1471, 1472, 1473, and / or 1474), an offset is generated from its neighboring samples. For example, for the first sub-block (1471), an offset is generated from its top neighboring samples (P1, P2, P3, and P4 in 1420), left neighboring samples (P5, P6, P7, and P8 in 1410), and / or upper left neighboring samples (P0) (1401). In some embodiments, the top neighboring samples may include top neighboring samples (P1, P2, P3, and P4 in 1420) and upper left neighboring samples (P0) (1401). In some other embodiments, the left neighboring samples may include left neighboring samples (P5, P6, P7, and P8 in 1410) and upper left neighboring samples (P0) (1401).
[0146] The first sub-block (1471) includes 4×4 pixels, and each pixel in the 4×4 pixels corresponds to predN, which is the Nth neighboring prediction sample before refinement, such as pred0, pred1, pred2, ... pred16.
[0147] In various embodiments, the offset value of each pixel of a given sub-block may be calculated based on neighboring samples according to a formula, which may be a predefined formula or a formula indicated by a parameter encoded in an encoded bitstream.
[0148] In reference Fig. 14B In some implementations of , the offset value (offset(k)) of the kth position of a given sub-block may be generated as follows:
[0149]
[0150] pred_refined k =clip3(pred k +offset(k)) (3)
[0151] W kn is a predefined weight used for offset calculation. n are the values of neighboring samples (e.g., P0, P1, P2, ..., P8). k is the predicted value of the pixel after applying intra prediction or other prediction (e.g., inter prediction). pred_refined k is the thinning value of the pixel after applying ORIP. clipP3() is the clip 3 math function. n is an integer from 0 to 8, and k is an integer from 0 to 15.
[0152] In some embodiments, W kn It can be predefined and can be obtained according to Table 1.
[0153] Table 1 Predefined weights for offset calculation
[0154] k <![CDATA[W k0 ]]> <![CDATA[W k1 ]]> <![CDATA[W k2 ]]> <![CDATA[W k3 ]]> <![CDATA[W k4 ]]> <![CDATA[W k5 ]]> <![CDATA[W k6 ]]> <![CDATA[W k7 ]]> <![CDATA[W k8 ]]> 0 4 16 4 0 0 16 4 0 0 1 2 4 16 4 0 8 2 0 0 2 1 0 4 16 4 4 1 0 0 3 0 0 2 4 16 2 0 0 0 4 2 8 2 0 0 4 16 4 0 5 0 2 8 2 0 2 8 2 0 6 0 0 2 8 2 1 4 1 0 7 0 0 0 2 8 1 2 0 0 8 0 4 0 0 0 0 4 16 4 9 0 0 4 0 0 0 2 8 2 10 0 0 1 4 1 0 1 4 1 11 0 0 0 2 4 0 0 4 0 12 0 0 1 0 0 0 2 4 16 13 0 0 0 1 0 0 1 2 8 14 0 0 1 2 1 0 0 1 4 15 0 0 0 1 2 0 0 1 2
[0155] In some other embodiments, sub-block based ORIP may be applied only to a predefined set of intra prediction modes and / or luma and chroma may be different, depending on the intra prediction mode. Table 2 shows an embodiment of sub-block based ORIP according to various intra prediction modes and luma or chroma channels. Taking the luma channel as an example: when the prediction mode is DC or SMOOTH, ORIP is always on and no additional signaling is required; when the prediction mode is HOR / VER and the delta angle (angle_delta) is equal to 0, block-level signaling is required to enable / disable ORIP; and / or when the intra prediction mode is other modes, ORIP is always off and no additional signaling is required.
[0156] Table 2 Modes in the proposed method depending on on / off
[0157]
[0158] Back to the second 4×4 sub-block (1473), due to its relative position to the first 4×4 sub-block (1471), the top neighboring samples of the second sub-block may be some pixels of the first sub-block: P1 of the second sub-block may be pred12 of the first block, P2 of the second sub-block may be pred13 of the first sub-block, p3 of the second sub-block may be pred14 of the first sub-block, and p4 of the second sub-block may be pred15 of the first sub-block. The top left neighboring sample (P0) of the second sub-block may be the left neighboring sample (P8) of the first sub-block.
[0159] In various embodiments, the available intra prediction modes or mode options for the current block being encoded may be divided into a plurality of intra prediction mode sets. Each set may be assigned a set index. The set index of a set is an integer representing the index of the set in a plurality of intra prediction mode sets. In some embodiments, the set index may be an integer equal to or greater than 0; or in some other embodiments, the set index may be an integer equal to or greater than 1. Each set may contain a plurality of intra mode prediction modes. The mode index of an intra mode prediction mode is an integer representing the index of the intra mode prediction mode in a plurality of intra mode prediction modes. In some embodiments, the mode index may be an integer equal to or greater than 0; or in some other embodiments, the mode index may be an integer equal to or greater than 1. Based on the correlation between intra prediction modes between blocks, the manner in which the available intra prediction modes are split and sorted and the intra prediction modes are sorted in each mode set may be determined at least in part according to the intra prediction modes used by its neighboring blocks. The intra prediction mode used by the neighboring blocks may be referred to as a "reference intra prediction mode" or a "reference mode". The intra prediction mode of a particular unit may be determined and selected. The selection of the intra prediction mode may be written. First, a set index may be written to indicate the set index of the intra-prediction mode set containing the selected intra-prediction mode. Second, a mode index (alternatively referred to as a mode position index within a set) may be written to indicate the index of the selected intra-prediction mode within the mode set.
[0160] The general implementation of intra-frame prediction mode division and sorting above and the specific examples below use statistical effects and neighboring correlations to dynamically index these modes so that the syntax design for writing their selection into the encoded video bitstream can be optimized to improve coding efficiency. For example, these implementations help reduce the amount of syntax written and help generate entropy coded contexts more efficiently.
[0161] The various embodiments and / or implementations described in the embodiments of the present application can be used alone or in combination in any order. In addition, a part, all or any part or all combination of these embodiments and / or implementations can be embodied as a part of an encoder and / or decoder, and can be implemented with hardware and / or software. For example, they can be hard-coded in a dedicated processing circuit (e.g., one or more integrated circuits). In another example, they can be implemented by one or more processors executing a program stored in a non-transitory computer-readable medium.
[0162] When applying ORIP, there may be some issues / problems related to intra-frame mode coding. For example, in ORIP design, if the nominal mode is HOR / VER, an additional increment angle can be added to indicate the use of ORIP; in intra-frame mode coding design, the writing of the nominal mode and the increment angle can be modified, resulting in at least one issue / problem, that is, it is impossible to directly merge ORIP and intra-frame mode coding together.
[0163] The embodiments of the present application describe various embodiments of intra-frame prediction mode coding in video encoding and / or decoding, solve at least one of the problems / difficulties discussed above, and achieve an effective combination of ORIP and improved intra-frame mode coding.
[0164] In various embodiments, Fig.15 A method 1500 for intra-frame prediction mode encoding in video decoding is shown, and the method 1500 may include some or all of the following steps: step 1510, receiving an encoded video code stream for a block through a device including a memory storing instructions and a processor communicating with the memory; step 1520, the device divides the intra-frame mode list into multiple intra-frame mode sets for the block based on mode information of each intra-frame mode in the intra-frame mode list, and the intra-frame mode list corresponds to the intra-frame prediction mode of at least one neighboring block of the block; step 1530, the device extracts a set index indicating an intra-frame mode set from the multiple intra-frame mode sets from the encoded video code stream; step 1540, the device extracts a mode index indicating an intra-frame prediction mode from the intra-frame mode set from the encoded video code stream; and step 1550, the device determines the intra-frame prediction mode of the block based on the set index and the mode index.
[0165] In some embodiments, the mode information of the intra mode may include at least one of the following: a directional mode, a nominal angle of a directional mode, an offset angle of a directional mode, a non-directional mode, a smooth mode (e.g., smooth, smooth_v, smooth_h), a DC mode, a PAETH mode, and / or a mode for generating prediction samples according to a given prediction direction. In some other embodiments, in a loose classification, the directional mode may broadly include: any mode that is not a smooth mode (smooth, smooth_v, smooth_h), a DC or a PAETH mode; and any mode for generating prediction samples according to a given prediction direction. In some other embodiments, the non-directional mode may include a smooth mode (e.g., smooth, smooth_v, smooth_h), a DC mode, a PAETH mode, and a luma-for-chroma mode. In some other embodiments, in a loose classification, the non-directional mode may broadly include any mode that is not a directional mode.
[0166] In various embodiments of the embodiments of the present application, the size of a block (such as but not limited to a coding block, a prediction block, or a transform block) may refer to the width or height of the block. The width or height of the block may be an integer in pixels. In various embodiments of the embodiments of the present application, the size of a block may refer to the area size of the block. The area size of a block may be an integer calculated by multiplying the width of the block in pixels by the height of the block. In some different embodiments of the embodiments of the present application, the size of a block may refer to the maximum value of the width or height of the block, the minimum value of the width or height of the block, or the aspect ratio of the block. The aspect ratio of the block may be calculated by dividing the width of the block by the height, or may be calculated by dividing the height of the block by the width.
[0167] In some embodiments, the available intra-frame prediction modes of the current block can be divided / split into multiple intra-frame prediction mode sets according to the intra-frame prediction modes of its neighboring blocks. In order to obtain the intra-frame prediction mode of the current block, first, the set index can be written to indicate the set index of the intra-frame prediction mode of the current block; secondly, the mode index can be written to indicate the index of the intra-frame prediction mode within the mode set. In some embodiments, all non-directional modes can be included in the first mode set. The set index of the set is an integer representing the index of the set in multiple sets; and / or the mode index of the mode is an integer representing the index of the mode in the mode set. In some embodiments, the set index and / or the mode index can be an integer equal to or greater than 0. In some other embodiments, the set index and / or the mode index can be an integer equal to or greater than 1.
[0168] In various embodiments of the embodiments of the present application, the "first" mode set refers not only to "one" mode set, but also to the "first" mode set with the minimum set index, the "second" mode set refers not only to "another" mode set, but also to the "second" mode set with the second minimum set index, and so on. For example, multiple intra-frame prediction mode sets can be indicated by M, and the set index can be in the range of, for example, 1 to M or 0 to M-1. When the set index ranges from 1 to M, the "first" mode set is the "first" mode set with a set index of 1, the "second" mode set is the "second" mode set with a set index of 2, and so on. When the set index ranges from 0 to M-1, the "first" mode set is the "first" mode set with a set index of 0, the "second" mode set is the "second" mode set with a set index of 1, and so on.
[0169] In various embodiments of the embodiments of the present application, "XYZ is written" may refer to XYZ being encoded into an encoded code stream during the encoding process; and / or, after the encoded code stream is sent from one device to another device, "XYZ is written" may refer to decoding / extracting XYZ from the encoded code stream during the decoding process.
[0170] Referring to step 1510, the device may be Figure 5 An electronic device (530) or Figure 8 In some embodiments, the device may be a video decoder (810) in Figure 6 In other embodiments, the device may be Figure 5 a portion of an electronic device (530) in Figure 8 A portion of a video decoder (810) in Figure 6 The encoded video stream may be a part of the decoder (633) in the encoder (620) in FIG. Figure 8 The coded video sequence in Figure 6 or Figure 7 The block can be referred to as a coding block or a coded block.
[0171] Referring to step 1520, the device may divide the intra-mode list into multiple intra-mode sets for the block based on the mode information of each intra-mode in the intra-mode list. In some embodiments, step 1520 may include determining multiple intra-mode sets by the device based on one or more factors, such as but not limited to the mode type of each intra-mode, the orientation angle of the directional intra-mode, and / or the intra-mode of the neighboring blocks of the block. The mode type of the intra-mode may include whether the intra-mode is a directional mode or a non-directional mode, and / or whether the intra-mode is an ORIP mode or a non-ORIP mode. In some other embodiments, the intra-mode list may correspond to at least one intra-prediction mode of the neighboring blocks of the block; and / or the intra-mode list may be used to derive the intra-prediction mode of the block.
[0172] In various embodiments, the intra-frame mode list may include a predetermined number of intra-frame prediction modes, for example, 61.
[0173] In some embodiments, the intra-mode list includes a first sub-list at the top of the intra-mode list; and the first sub-list includes all non-directional intra-prediction modes. For example, when the intra-mode list is divided into a plurality of intra-mode sets, since the non-directional intra-prediction mode is at the top of the intra-mode list, the non-directional intra-prediction mode can be divided into the first intra-mode set.
[0174] In some other embodiments, in response to a neighboring block using a directional intra prediction mode, the intra mode list includes a second sublist adjacent to the first sublist in the intra mode list; and the second sublist includes a plurality of derived intra prediction modes based on the directional intra prediction mode of the neighboring block. For example, when the intra mode list is divided into a plurality of intra mode sets, because the plurality of derived intra prediction modes are adjacent to the non-directional intra prediction mode from the top of the intra mode list, the plurality of derived intra prediction modes based on the directional intra prediction mode of the neighboring block may be divided into the first intra mode set or the second intra mode set. In some embodiments, the neighboring blocks of the current block may include the top (above) block of the current block, the left block of the current block, or both the top (above) block and the left block of the current block.
[0175] In some other embodiments, by adding an offset [0, -1, +1, -2, +2, -3, +3, -4, +4] to the directional intra prediction mode of the neighboring block, the multiple derived intra prediction modes include nine directional intra prediction modes. For example, when the directional intra prediction mode of the neighboring block has a specific directional angle (x degrees) and the step size of the directional angle is 3 degrees, the multiple derived intra prediction modes may include 9 directional intra prediction modes, and the directional angles are x degrees, x±3 degrees, x±6 degrees, x±9 degrees, and x±12 degrees.
[0176] In some other embodiments, the intra-frame mode list includes a third sub-list adjacent to the second sub-list in the intra-frame mode list; and the third sub-list includes at least one default intra-frame prediction mode. In some other embodiments, the default intra-frame prediction mode includes at least one nominal directional intra-frame prediction mode with a zero delta angle. For example, the default intra-frame prediction mode may include all other nominal directional intra-frame prediction modes with zero delta angles that are not already included in the above-described derived intra-frame prediction mode based on the directional intra-frame prediction mode of the neighboring block.
[0177] In some other embodiments, in response to a neighboring block using a non-directional intra-frame prediction mode: the intra-frame mode list includes a second sub-list located adjacent to the first sub-list in the intra-frame mode list; the second sub-list includes a plurality of nominal directional intra-frame prediction modes with zero delta angles. For example, when both neighboring blocks are non-directional modes, all directional intra-frame prediction modes with zero delta angles will be added to the second sub-list. In some other embodiments, in response to all neighboring blocks using a non-directional intra-frame prediction mode: the intra-frame mode list includes a second sub-list located adjacent to the first sub-list in the intra-frame mode list; the second sub-list includes a plurality of nominal directional intra-frame prediction modes with zero delta angles. All neighboring blocks may include a top neighboring block (i.e., a neighboring block located above the current block) and a left neighboring block (i.e., a neighboring block located to the left of the current block); and all neighboring blocks may not include a right neighboring block (i.e., a neighboring block located to the right of the current block) or a bottom neighboring block (i.e., a neighboring block located below the current block).
[0178] For example, first, all non-directional intra prediction modes can be added to the intra mode list. Secondly, the offset can be added to the directional intra prediction mode of the neighboring block to derive the intra prediction mode, and the derived intra prediction mode can be added to the intra mode list. Finally, after adding all the derived intra prediction modes, if the intra prediction mode list is still not full, the default mode can be used to fill the remaining positions in the intra mode list.
[0179] In another example, the available intra prediction modes for the current block are 61, including 5 non-directional modes and 56 directional modes. When only one of the neighboring blocks of the current block is encoded with a directional intra prediction mode, the 5 non-directional modes are first added to the mode list; secondly, 9 directional intra prediction modes are derived by adding offsets [0, -1, +1, -2, +2, -3, +3, -4, +4] to the directional modes of the neighboring blocks. After that, only 14 modes are added to the mode list, and 47 (61-14=47) positions in the intra mode list are not filled. Then, the default mode is added to the mode list. In some embodiments, if the intra mode list is not full, a nominal angle with an incremental angle equal to zero is first used as a default mode to fill the intra mode list.
[0180] In various embodiments, multiple intra-mode sets include a first intra-mode set; and the first intra-mode set includes all non-directional intra-prediction modes. For example, non-directional modes are always included in the first mode set. Non-directional modes may include smooth modes (e.g., smooth, smooth_v, smooth_h), DC modes, PAETH modes, and luminance-chrominance modes. In some other embodiments, in a loose classification, non-directional modes may broadly include any mode that is not a directional mode.
[0181] In various embodiments, the plurality of intra mode sets include a first intra mode set; and the first intra mode set consists of all non-directional intra prediction modes. For example, the first mode set includes only non-directional intra prediction modes.
[0182] In various embodiments, the plurality of intra mode sets include a first intra mode set; and the first intra mode set includes all non-directional intra prediction modes and one or more offset based refinement for intra prediction (ORIP) modes. In some embodiments, the one or more ORIP modes include a vertical ORIP mode and a horizontal ORIP mode. For example, two additional modes VER_ORIP and HOR_ORIP are specified to indicate the use of ORIP on VER and HOR modes with an increment angle equal to 0. These two modes can be marked as non-directional modes and divided into the first mode set.
[0183] In various embodiments, the one or more ORIP modes include all ORIP modes. For example, all intra prediction modes to which ORIP is applied are considered non-directional modes.
[0184] In various embodiments, the plurality of intra-mode sets include N intra-mode sets and M intra-mode sets, wherein: the N intra-mode sets precede the M intra-mode sets in ascending order of set index, N is a positive integer, and M is a positive integer; and the number of intra-prediction modes in each of the M intra-mode sets is equal to a power of 2. In some implementations, N is, for example but not limited to, 1 or 2. For example, N is 1 and M is 4; the plurality of intra-mode sets include, in ascending order of intra-mode set index, a first intra-mode set, a second intra-mode set, a third intra-mode set, a fourth intra-mode set, and a fifth intra-mode set; the number of intra-prediction modes in the first intra-mode set is one of the following: 5 or 7; the number of intra-prediction modes in the second intra-mode set is 8; the number of intra-prediction modes in the third intra-mode set is 16; the number of intra-prediction modes in the fourth intra-mode set is 16; and the number of intra-prediction modes in the fifth intra-mode set is 16.
[0185] For another example, except for the first N intra mode sets, the number of modes in each mode set is equal to a power of 2; N is a positive integer, such as 1 or 2. When N is 1, it may mean that only the number of modes in the first mode set is not equal to a power of 2.
[0186] In one example, the available intra prediction modes are divided into 5 mode sets; and the number of modes in each mode set is (5, 8, 16, 16, 16). The number of the first mode set is 5, which is equal to the number of non-directional modes.
[0187] In another example, the available intra prediction modes are divided into 5 mode sets. The number of modes in each mode set is (7, 8, 16, 16, 16). The number of the first mode set is 7, which is equal to the number of non-directional modes plus VER_ORIP and HOR_ORIP.
[0188] The embodiments in the embodiments of the present application can be used alone or in combination in any order. In addition, each method (or embodiment), encoder and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-temporary computer-readable medium. The embodiments in the embodiments of the present application can be applied to luminance blocks or chrominance blocks. And in the chrominance block, the embodiments can be applied to more than one color component individually, or can be applied to more than one color component together.
[0189] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.16A computer system (2600) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0190] As a non-limiting example, a computer system having an architecture (2600), particularly a kernel (2640), can provide functionality due to one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, as well as certain non-temporary kernel (2640) memories, such as kernel internal mass storage (2647) or ROM (2645). Software for implementing various embodiments of the present application can be stored in such devices and executed by the kernel (2640). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can enable the kernel (2640), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific processes or specific parts of specific processes described herein, including defining data structures (2646) stored in RAM and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality due to hardwiring or otherwise embodied in logic in a circuit (e.g., an accelerator (2644)) that may replace software or operate in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, a reference to a portion of software may include logic and vice versa. Where appropriate, a reference to a portion of a computer-readable medium may include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both. Embodiments of the present application include any suitable combination of hardware and software.
[0191] Computer software may be encoded using any suitable machine code or computer language that may be subjected to assembly, compilation, linking or similar mechanisms to create code comprising instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpreted code, microcode, etc.
[0192] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, IoT devices, etc.
[0193] Fig.16The components shown for the computer system (2600) are exemplary in nature and are not intended to impose any limitations on the scope of use or functionality of computer software implementing embodiments of the present application. The configuration of the components should not be interpreted as having any dependency or requirement related to any one or combination of components shown in the exemplary embodiment of the computer system (2600).
[0194] The computer system (2600) may include certain human-machine interface input devices. Such human-machine interface input devices may be responsive to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). Human-machine interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.
[0195] The input human-machine interface device may include one or more of the following (only one of each is shown): keyboard (2601), mouse (2602), touchpad (2603), touch screen (2610), data gloves (not shown), joystick (2605), microphone (2606), scanner (2607), camera (2608).
[0196] The computer system (2600) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate one or more human user senses, for example, through tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen (2610), a data glove (not shown), or a joystick (2605), but may also be a tactile feedback device that is not an input device), audio output devices (e.g., speakers (2609), headphones (not shown)), visual output devices (e.g., screens (2610) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which are capable of outputting two-dimensional visual output or more than three-dimensional output through devices such as stereo image output, virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).
[0197] The computer system (2600) may also include human-accessible storage devices and their associated media: for example, optical media including CD / DVD ROM / RW (2620) with CD / DVD etc. media (2621), thumb drives (2622), removable hard drives or solid-state drives (2623), traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD based devices such as security dongles (not shown), etc.
[0198] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.
[0199] The computer system (2600) may also include an interface (2654) to one or more communication networks (2655). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial television including CAN Bus, etc. Some networks typically require an external network interface adapter connected to some common data port or peripheral bus (2649) (e.g., a USB port of the computer system (2600)); as described below, other network interfaces are typically integrated into the kernel of the computer system (2600) by connecting to the system bus (e.g., connecting to an Ethernet interface in a PC computer system or connecting to a cellular network interface in a smartphone computer system). The computer system (2600) can use any of these networks to communicate with other entities. Such communications may be one-way receive only (e.g., broadcast television), one-way send only (e.g., a CANbus connected to certain CANbus devices), or bidirectional, for example, connecting to other computer systems using a LAN or WAN digital network. As described above, certain protocols and protocol stacks may be used on each of those networks and network interfaces.
[0200] The above-mentioned human interface device, human-accessible storage device, and network interface may be attached to the kernel (2640) of the computer system (2600).
[0201] The core (2640) may include one or more central processing units (CPUs) (2641), graphics processing units (GPUs) (2642), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (2643), hardware accelerators for certain tasks (2644), graphics adapters (2650), etc. These devices, as well as read-only memory (ROM) (2645), random access memory (2646), internal mass storage (2647) such as internal non-user accessible hard drives, SSDs, etc., may be connected via a system bus (2648). In some computer systems, the system bus (2648) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus (2648) or to the core's system bus (1848) via a peripheral bus (2649). In one example, a screen (2610) may be connected to a graphics adapter (2650). The architecture of the peripheral bus includes PCI, USB, etc.
[0202] The CPU (2641), GPU (2642), FPGA (2643) and accelerator (2644) can execute certain instructions, which can be combined to form the above-mentioned computer code. The computer code can be stored in ROM (2645) or RAM (2646). Transition data can also be stored in RAM (2646), while permanent data can be stored in, for example, internal mass storage (2647). Fast storage and retrieval of any storage device can be performed by using a high-speed cache, which can be closely associated with the following: one or more CPUs (2641), GPUs (2642), mass storage (2647), ROM (2645), RAM (2646), etc.
[0203] The computer readable medium may have thereon computer code for performing various computer-implemented operations. The medium and computer code may be specially designed and constructed for the purposes of the embodiments of the present application, or the medium and computer code may be of a type known and available to those skilled in the art of computer software.
[0204] Although a particular invention has been described with reference to illustrative embodiments, this description is not meant to be limiting. Based on this description, various modifications of the illustrative embodiments and additional embodiments of the present invention will be apparent to those of ordinary skill in the art. Those skilled in the art will readily recognize that these and various other modifications may be made to the exemplary embodiments illustrated and described herein without departing from the spirit and scope of the present invention. Therefore, it is contemplated that the appended claims will cover any such modifications and alternative embodiments. Certain proportions in the accompanying drawings may be exaggerated, while other proportions may be minimized. Therefore, the present application embodiments and the accompanying drawings should be considered illustrative, rather than limiting.
Claims
1. A video code stream encoding method, characterized in that: The method comprises: A device receives a video code stream of a block in an intra prediction mode, the device comprising a memory storing instructions and a processor in communication with the memory; The device divides the intra-mode list into a plurality of intra-mode sets for the block based on mode information of each intra-mode in the intra-mode list, the intra-mode list corresponding to an intra-prediction mode of at least one neighboring block of the block; The device extracts, from the video code stream, a set index indicating an intra-mode set from the multiple intra-mode sets; The device extracts, from the video code stream, a mode index indicating an intra-frame prediction mode from the one intra-frame mode set; and The device encodes the video code stream into an encoded video code stream based on the set index and the mode index.
2. The method according to claim 1, characterized in that: The intra-mode list includes a predetermined number of intra-prediction modes.
3. The method according to claim 2, characterized in that: The predetermined number is 61.
4. The method according to claim 1, characterized in that: The first intra mode set is located at the top of the intra mode list, the first intra mode set including all non-directional intra prediction modes and one or more offset-based intra prediction refinement (ORIP) modes.
5. The method according to any one of claims 1 to 4, characterized in that: In response to a neighboring block of the block using a directional intra prediction mode: The intra-mode list includes a second intra-mode set immediately adjacent to the first intra-mode set in the intra-mode list; The second intra-mode set includes a plurality of derived intra-prediction modes based on the directional intra-prediction mode of the neighboring block.
6. The method according to claim 5, characterized in that: The plurality of derived intra prediction modes include 9 directional intra prediction modes by adding offsets [0, -1, +1, -2, +2, -3, +3, -4, +4] to the directional intra prediction modes of the neighboring blocks.
7. The method according to claim 5, characterized in that: The intra-mode list includes a third intra-mode set immediately adjacent to the second intra-mode set in the intra-mode list; and The third intra-mode set includes at least one default intra-prediction mode.
8. The method according to claim 7, characterized in that: The at least one default intra prediction mode includes at least one nominal directional intra prediction mode having a delta angle of zero.
9. The method according to any one of claims 1 to 4, characterized in that: In response to a neighboring block of the block using a non-directional intra prediction mode: The intra-mode list includes a second intra-mode set immediately adjacent to the first intra-mode set in the intra-mode list; and The second set of intra modes includes a plurality of nominal directional intra prediction modes having zero delta angles.
10. The method according to claim 1, characterized in that: The multiple intra mode sets include N intra mode sets and M intra mode sets, wherein: The N intra mode sets are located before the M intra mode sets in ascending order of set indexes. N is a positive integer, M is a positive integer; and The number of intra prediction modes in each of the M intra mode sets is equal to a power of 2.
11. The method according to claim 10, characterized in that: N is 1, M is 4; The multiple intra-frame mode sets include, in ascending order of the intra-frame mode set index, a first intra-frame mode set, a second intra-frame mode set, a third intra-frame mode set, a fourth intra-frame mode set and a fifth intra-frame mode set; The number of intra prediction modes in the first intra mode set is one of: 5 or 7; The number of intra prediction modes in the second intra mode set is 8; The number of intra prediction modes in the third intra mode set is 16; The number of intra prediction modes in the fourth intra mode set is 16; and The number of intra prediction modes in the fifth intra mode set is 16.
12. A method for decoding a video code stream, characterized in that: The method comprises: Receiving an encoded video stream for a block by a device, the device comprising a memory storing instructions and a processor in communication with the memory; The device divides the intra-mode list into a plurality of intra-mode sets for the block based on mode information of each intra-mode in the intra-mode list, the intra-mode list corresponding to an intra-prediction mode of at least one neighboring block of the block; The device extracts, from the encoded video code stream, a set index indicating an intra-mode set from the plurality of intra-mode sets; The device extracts a mode index indicating an intra-frame prediction mode from the one intra-frame mode set from the encoded video code stream; and The apparatus determines the intra prediction mode of the block based on the set index and the mode index.
13. A device for encoding a video code stream, characterized in that: The device comprises: a memory for storing instructions; and A processor in communication with the memory, wherein when the processor executes the instruction, the processor is configured to cause the device to perform the video code stream encoding method according to any one of claims 1 to 11.
14. A device for decoding a video code stream, characterized in that: The device comprises: a memory for storing instructions; and A processor in communication with the memory, wherein when the processor executes the instruction, the processor is configured to cause the device to perform the method for decoding a video code stream as claimed in claim 12.
15. A non-transitory computer-readable storage medium storing instructions, characterized in that: When the instruction is executed by a processor, the instruction is configured to cause the processor to execute the method for decoding a video code stream according to claim 12, or to execute the method for encoding a video code stream according to any one of claims 1 to 11.
16. A method for storing or sending a video code stream, characterized in that: The video code stream is decoded based on the method for decoding the video code stream according to claim 12, or the video code stream is generated based on the method for encoding the video code stream according to any one of claims 1 to 11.