Intra-frame prediction method and device
By acquiring and replacing unavailable reference samples, the availability of reference samples during video decoding is improved, and the problem of low video compression and decompression efficiency is solved, and more efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202510317680.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-11-21
- Filing Date
- 2019-11-21
- Publication Date
- 2025-07-04
AI Technical Summary
In the case of limited bandwidth communication networks and limited storage resources, existing video decoding technologies are difficult to achieve efficient video compression and decompression while maintaining image quality.
By acquiring the intra prediction mode of the current block, checking the availability of the reference sample, and replacing the unavailable sample with available reference samples, performing intra prediction reconstruction, especially for the Y component, Cb component or Cr component, improving the available information acquisition of the reference sample.
It improves the information accuracy in the intra prediction process, enhances the efficiency of video encoding and decoding, and is suitable for various video decoding devices and systems.
Smart Images

Figure CN120263968A_ABST
Abstract
Description
[0001] This application is a divisional application. The application number of the original application is 201980069241.2, the original application date is November 21, 2019, and the entire content of the original application is incorporated herein by reference.
[0002] Cross-reference to related applications
[0003] This application claims priority to U.S. Provisional Patent Application No. 62 / 770,736, filed on November 21, 2018, with the title "Intra Prediction Method and Device", the disclosure of which is incorporated herein by reference. Technical Field
[0004] Embodiments of the present invention generally relate to the field of video coding, and more particularly to the field of intra prediction methods and devices. Background Art
[0005] Even for short videos, a large amount of video data needs to be described, which can cause difficulties when the data is to be sent over a communication network with limited bandwidth capacity or otherwise. Therefore, video data is usually compressed first and then sent in modern telecommunication networks. Since memory resources may be limited, the size of the video may also be a problem when storing the video in a storage device. Video compression devices typically use software and / or hardware on the source side to decode the video data before sending or storing it, thereby reducing the amount of data required to represent the digital video image. Then, the compressed data is received at the destination side by a video decompression device for decoding the video data. In the context of limited network resources and the growing demand for higher video quality, improved compression and decompression techniques are needed that can increase the compression ratio with little impact on image quality. Summary of the Invention
[0006] Embodiments of the present invention provide an intra prediction apparatus and method for encoding and decoding images. Embodiments of the present invention should not be construed as limiting the examples set forth herein.
[0007] The method provided by the first aspect of the present invention includes: obtaining an intra prediction mode of a current block; obtaining the availability of reference samples of a component of the current block; replacing unavailable reference samples with available reference samples; obtaining a prediction of the current block according to the intra prediction mode and the replaced reference samples; and reconstructing the current block according to the prediction. In one embodiment, the component includes a Y component, a Cb component, or a Cr component. In another embodiment, the component includes a luminance component or a chrominance component.
[0008] Due to the availability of reference samples in each component in the first aspect of the present invention, available information can be provided more accurately.
[0009] According to one implementation of the first aspect, all Cb samples in the reconstruction block are marked as available; alternatively, all Cr samples in the reconstruction block are marked as available.
[0010] Since this implementation of the first aspect stores available information for each sample in each component, available information can be provided more accurately during the intra prediction process.
[0011] According to a second aspect of the present invention, the decoder includes processing circuitry for performing the steps of the above method.
[0012] According to a third aspect of the present invention, the encoder includes processing circuitry for performing the steps of the above method.
[0013] According to a fourth aspect of the present invention, the computer program product includes program code that, when executed by a processor, performs the above method.
[0014] According to a fifth aspect of the present invention, the decoder for intra prediction includes one or more processing units and a non-transitory computer-readable storage medium that is coupled to the one or more processing units and stores program instructions, and the one or more processing units execute the program instructions to perform the above method.
[0015] According to a sixth aspect of the present invention, the encoder for intra prediction includes one or more processing units and a non-transitory computer-readable storage medium that is coupled to the one or more processing units and stores program instructions, and the one or more processing units execute the program instructions to perform the above method.
[0016] Embodiments of the present invention also provide a decoding device and an encoding device for performing the above method.
[0017] For clarity, any of the above embodiments can be combined with any of the other above embodiments to create new embodiments within the scope of the present invention.
[0018] These and other features will be more clearly understood from the following detailed description in conjunction with the accompanying drawings and the claims. Description of the Drawings
[0019] To more fully understand the present invention, reference is made to the following brief description taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like parts.
[0020] Figure 1A It is a block diagram of an exemplary decoding system in which embodiments of the present invention can be implemented.
[0021] Figure 1B It is a block diagram of another example decoding system that can implement the embodiments of the present invention.
[0022] Figure 2 It is a block diagram of an example video encoder that can implement the embodiments of the present invention.
[0023] Figure 3 It is a block diagram of an example of a video decoder that can implement the embodiments of the present invention.
[0024] Figure 4 It is a schematic diagram of a network device provided by an exemplary embodiment of the present invention.
[0025] Figure 5 It is provided by an exemplary embodiment of the present invention and can be used as Figure 1A a simplified block diagram of a device that can be one or both of a source device and a target device.
[0026] Figure 6 It is a description of the intra prediction algorithm of H.265 / HEVC that can be used in the description of the present invention.
[0027] Figure 7 It is a schematic diagram of a reference sample provided by an exemplary embodiment of the present invention.
[0028] Figure 8 It is a schematic diagram of available and unavailable reference samples provided by an exemplary embodiment of the present invention.
[0029] Figure 9 It is a simplified flowchart for obtaining a reconstructed signal provided by an exemplary embodiment of the present invention.
[0030] Figure 10 It is a simplified flowchart for obtaining a prediction signal provided by an exemplary embodiment of the present invention.
[0031] Figure 11 It is a schematic diagram in which samples in the current block are marked as available provided by an exemplary embodiment of the present invention.
[0032] Figure 12 It is a schematic diagram in which samples in the current block are marked as available at the right and bottom boundaries of the unit N×N provided by an exemplary embodiment of the present invention.
[0033] Figure 13 It is a schematic diagram in which samples at the right and bottom boundaries of the current block are marked as available provided by an exemplary embodiment of the present invention.
[0034] Figure 14 It is a schematic diagram in which samples at the right and bottom boundaries of the current block in the unit N are marked as available provided by an exemplary embodiment of the present invention.
[0035] Figure 15 Schematic diagram of samples at the right and bottom boundaries of the current block in unit N×N provided for an exemplary embodiment of the present invention being marked as available.
[0036] Figure 16 Schematic diagram of a method for marking samples of a chrominance component provided for an exemplary embodiment of the present invention.
[0037] Figure 17 Block diagram of an exemplary structure of a content supply system 3100 for implementing a content distribution service.
[0038] Figure 18 Block diagram of an exemplary structure of a terminal device. Detailed implementation
[0039] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using a variety of techniques, whether the techniques are currently known or existing techniques. The present invention should not be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but can be modified within the full scope of the appended claims and their equivalents.
[0040] Figure 1A Schematic block diagram of an exemplary decoding system 10 that can use bidirectional prediction techniques. As Figure 1A shown, the decoding system 10 includes a source device 12 that provides encoded video data, and a target device 14 that decodes the encoded video data. In particular, the source device 12 can provide the video data to the target device 14 via a computer-readable medium 16. The source device 12 and the target device 14 can include any of a variety of devices, including desktop computers, laptop computers (i.e., notebook computers), tablet computers, set-top boxes, handheld phones (e.g., smart phones, smart tablets), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source device 12 and the target device 14 can be used for wireless communication.
[0041] The target device 14 may receive encoded video data to be decoded via a computer-readable medium 16. The computer-readable medium 16 may include any type of medium or device capable of moving the encoded video data from the source device 12 to the target device 14. In one example, the computer-readable medium 16 may include a communication medium to enable the source device 12 to send the encoded video data directly to the target device 14 in real time. The encoded video data may be modulated according to a communication standard (e.g., a wireless communication protocol) and sent to the target device 14. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network (such as the Internet). The communication medium may include routers, switches, base stations, or any other devices that may help facilitate communication from the source device 12 to the target device 14.
[0042] In some examples, the encoded data can be output from the output interface 22 to a storage device. Similarly, the encoded data can be accessed from the storage device through the input interface. The storage device can include any of a variety of distributed or locally accessible data storage media, such as a hard disk drive, a Blu-ray disc, a digital video disk (DVD), a compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In another example, the storage device can correspond to a file server or can correspond to other intermediate storage devices that can store the encoded video generated by the source device 12. The target device 14 can access the stored video data from the storage device by streaming or downloading. The file server can be any type of server capable of storing the encoded video data and sending the encoded video data to the target device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. The target device 14 can access the encoded video data through any standard data connection, including an Internet connection. This can include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., digital subscriber line (DSL), cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device can be a streaming transmission, a download transmission, or a combination thereof.
[0043] The techniques of the present invention are not necessarily limited to wireless applications or settings. These techniques can be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission (such as HTTP-based dynamic adaptive streaming over HTTP (DASH)), digital video encoded into a data storage medium, decoding digital video stored in a data storage medium, or other applications. In some examples, the decoding system 10 can be used to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0044] In Figure 1AIn the example, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The target device 14 includes an input interface 28, a video decoder 30, and a display device 32. According to the present invention, the video encoder 200 of the source device 12 and / or the video decoder 300 of the target device 14 may apply bidirectional prediction technology. In other examples, the source device and the target device may include other components or devices. For example, the source device 12 may receive video data from an external video source (such as an external camera). Similarly, the target device 14 may be connected to an external display device instead of including an integrated display device.
[0045] Figure 1A The decoding system 10 shown is merely an example. The bidirectional prediction technology may be performed by any digital video encoding and / or decoding device. Although the technology of the present invention is generally performed by a video decoding device, these technologies may also be performed by a video encoder / decoder, commonly referred to as a "CODEC". In addition, the technology of the present invention may also be performed by a video preprocessor. The video encoder and / or decoder may be a graphics processing unit (GPU) or a similar device.
[0046] The source device 12 and the target device 14 are merely examples of such decoding devices, where the source device 12 generates encoded video data and sends it to the target device 14. In some examples, the source device 12 and the target device 14 may operate in a substantially symmetric manner, such that each of the source device 12 and the target device 14 includes a video encoding component and a decoding component. Therefore, the decoding system 10 may support unidirectional or bidirectional video transmission between the video devices 12, 14, such as for video streaming, video playback, video broadcasting, or video telephony.
[0047] The video source 18 of the source device 12 may include a video capture device (such as a video camera), a video archive including previously captured video, and / or a video feed interface for receiving video from a video content provider. As another alternative, the video source 18 may generate computer graphics-based data as the source video, or generate a combination of real-time video, archived video, and computer-generated video.
[0048] In some cases, when the video source 18 is a video camera, the source device 12 and the target device 14 may form a webcam or a video phone. However, as described above, the technology described in the present invention generally applies to video decoding and may be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. Then, the encoded video information may be output by the output interface 22 into the computer-readable medium 16.
[0049] The computer-readable medium 16 can include a transient medium, such as a wireless broadcast or a wired network transmission, or a storage medium (i.e., a non-transient storage medium), such as a hard disk, a flash drive, an optical disc, a digital video disc, a Blu-ray disc, or other computer-readable media. In some examples, a network server (not shown) can receive the encoded video data from the source device 12 and provide the encoded video data to the target device 14, for example, via a network transmission. Similarly, a computing device of a media production facility (such as an optical disc stamping facility) can receive the encoded video data from the source device 12 and produce an optical disc including the encoded video data. Thus, in various examples, the computer-readable medium 16 can be understood to include one or more computer-readable media in various forms.
[0050] The input interface 28 of the target device 14 receives information from the computer-readable medium 16. The information of the computer-readable medium 16 can include syntax information defined by the video encoder 20, which is also used by the video decoder 30. The syntax information includes syntax elements that describe the characteristics and / or processing of blocks and other decoding units (e.g., groups of pictures (GOPs)). The display device 32 displays the decoded video data to the user and can include any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.
[0051] Video encoder 200 and video decoder 300 may operate according to a video coding standard, such as the High Efficiency Video Coding (HEVC) standard currently under development, and may conform to the HEVC Test Model (HM). Alternatively, video encoder 200 and video decoder 300 may operate according to other proprietary or industry standards, such as the International Telecommunications Union Telecommunication Standardization Sector (ITU-T) H.264 standard, or what is known as the Moving Picture Expert Group (MPEG)-4, Part 10, Advanced Video Coding (AVC), H.265 / HEVC, or extensions of such standards. However, the techniques of the present invention are not limited to any particular coding standard. Other examples of video coding standards include MPEG-2 and ITU-T H.263. Although not shown in Figure 1A In some aspects, video encoder 200 and video decoder 300 may be integrated with an audio encoder and decoder, respectively, and may include appropriate multiplexer-demultiplexer (MUX-DEMUX) units or other hardware and software to encode audio and video in a common data stream or separate data streams. If applicable, the MUX-DEMUX unit may conform to the ITU H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP).
[0052] The video encoder 200 and the video decoder 300 can be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When these techniques are implemented partially in software, the device can store the software instructions in a suitable non-transitory computer-readable medium and execute these instructions in hardware by one or more processors to perform the techniques of the present invention. Each of the video encoder 200 and the video decoder 300 can be included in one or more encoders or decoders, and any one of the encoders or decoders can be integrated as part of a combined encoder / decoder (codec) in the corresponding device. Devices such as the video encoder 200 and / or the video decoder 300 can include integrated circuits, microprocessors, and / or wireless communication devices, such as cellular phones.
[0053] Figure 1B An example video coding system 40 including Figure 2 the encoder 200 and / or Figure 3 the decoder 300 is provided for an exemplary embodiment. The system 40 can implement the techniques of the present invention, such as fusion estimation in inter-frame prediction. In the illustrated implementation, the video coding system 40 can include an imaging device 41, a video encoder 20, a video decoder 300 (and / or a video decoder implemented by the logic circuit 47 of the processing unit 46), an antenna 42, one or more processors 43, one or more memories 44, and / or a display device 45.
[0054] As shown, the imaging device 41, the antenna 42, the processing unit 46, the logic circuit 47, the video encoder 20, the video decoder 30, the processor 43, the memory 44, and / or the display device 45 can communicate with each other. As discussed, although both the video encoder 20 and the video decoder 30 are shown, in various practical scenarios, the video coding system 40 can include only the video encoder 20 or only the video decoder 30.
[0055] As shown in the figure, in some examples, the video decoding system 40 may include an antenna 42. For example, the antenna 42 may be used to transmit or receive an encoded bitstream of video data. Additionally, in some examples, the video decoding system 40 may include a display device 45. The display device 45 may be used to present video data. As shown in the figure, in some examples, the logic circuit 47 may be implemented by the processing unit 46. The processing unit 46 may include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. The video decoding system 40 may further include an optional processor 43, and the optional processor 43 may similarly include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. In some examples, the logic circuit 54 may be implemented by hardware, video decoding-specific hardware, and the like, and the processor 43 may implement general software, an operating system, and the like. Additionally, the memory 44 may be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting example, the memory 44 may be implemented by a cache memory. In some examples, the logic circuit 47 may access the memory 44 (e.g., for implementing an image buffer). In other examples, the logic circuit 47 and / or the processing unit 46 may include a memory (e.g., a cache, etc.) for implementing an image buffer and the like.
[0056] In some examples, the video encoder 200 implemented by the logic circuit may include an image buffer (e.g., implemented by the processing unit 46 or the memory 44) and a graphics processing unit (e.g., implemented by the processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the video encoder 200 implemented by the logic circuit 47 to implement the various modules discussed in conjunction with Figure 2 and / or any other encoder system or subsystem described herein. The logic circuit may be used to perform the various operations described herein.
[0057] The video decoder 300 may be implemented in a similar manner as implemented by the logic circuit 47 to implement in conjunction with Figure 3The various modules discussed for decoder 300 and / or any other decoder system or subsystem described herein. In some examples, video decoder 300, which can be implemented by logic circuitry, can include an image buffer (e.g., implemented by processing unit 46 or memory 44) and a graphics processing unit (e.g., implemented by processing unit 46). The graphics processing unit can be communicatively coupled to the image buffer. The graphics processing unit can include video decoder 300 implemented by logic circuitry 47 to implement in conjunction with Figure 3 The various modules discussed and / or any other decoder system or subsystem described herein.
[0058] In some examples, antenna 42 of video decoding system 40 can be used to receive an encoded bitstream of video data. As discussed, the encoded bitstream can include data related to video frame encoding discussed herein, indicators, index values, mode selection data, etc., such as data related to decoding partitions (e.g., transform coefficients or quantized transform coefficients, optional indicators as discussed, and / or data defining decoding partitions). Video decoding system 40 can also include video decoder 300 coupled to antenna 42, and video decoder 300 is used to decode the encoded bitstream. Display device 45 is used to present video frames.
[0059] Figure 2 FIG. is a block diagram of an example of video encoder 200 that can implement the technology of the present application. Video encoder 200 can perform intra and inter prediction on video blocks within a video slice. Intra prediction reduces or eliminates spatial redundancy of video in a given video frame or image through spatial prediction. Inter prediction reduces or eliminates temporal redundancy of video in adjacent frames or images of a video sequence through temporal prediction. Intra mode (I mode) can refer to any one of several spatial-based prediction modes. Inter modes (e.g., unidirectional prediction (P mode) or bidirectional prediction (B mode)) can refer to any one of several time-based prediction modes.
[0060] Figure 2 FIG. is a schematic / conceptual block diagram of an exemplary video encoder 200 for implementing the technology of the present invention. In Figure 2In the example, the video encoder 200 includes a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a buffer 216, a loop filter unit 220, a decoded picture buffer (DPB) 230, a prediction processing unit 260, and an entropy encoding unit 270. The prediction processing unit 260 may include an inter-frame estimation 242, an inter-frame prediction unit 244, an intra-frame estimation unit 252, an intra-frame prediction unit 254, and a mode selection unit 262. The inter-frame prediction unit 244 may further include a motion compensation unit (not shown). According to the hybrid video codec, Figure 2 the illustrated video encoder 200 may also be referred to as a hybrid video encoder or a video encoder.
[0061] For example, the residual calculation unit 204, the transform processing unit 206, the quantization unit 208, the prediction processing unit 260, and the entropy encoding unit 270 form the forward signal path of the encoder 200, while for example, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoder buffer (decoded picture buffer, DPB) 230, and the prediction processing unit 260 form the reverse signal path of the encoder, where the reverse signal path of the encoder corresponds to the signal path of a decoder (see Figure 3 the decoder 300).
[0062] For example, the encoder 200 is used to receive an image 201 or a block 203 of the image 201 through an input terminal 202. The image 201 is an image that constitutes an image sequence of a video or a video sequence, for example. The image block 203 may also be referred to as a current image block or a block to be decoded, and the image 201 is referred to as a current image or an image to be decoded (especially in video decoding, the current image is distinguished from other images, such as images that have been previously encoded and / or decoded in the same video sequence (i.e., the video sequence that also includes the current image)).
[0063] Segmentation
[0064] An embodiment of the encoder 200 may include a segmentation unit ( Figure 2 not shown in ), which is used to segment the image 201 into a plurality of blocks (such as the block 203), usually into a plurality of non-overlapping blocks. The segmentation unit may be used to use the same block size for all images in the video sequence and the corresponding grid defining the block size, or change the block size between images or subsets or groups of images, and segment each image into corresponding blocks.
[0065] In HEVC and other video coding specifications, to generate an encoded representation of an image, a set of coding tree units (CTUs) can be generated. Each CTU can include a coding tree block of luminance samples, two corresponding coding tree blocks of chrominance samples, and a syntax structure for coding the samples of the coding tree block. In a monochrome image or an image with three independent color planes, a CTU can include a single coding tree block and a syntax structure for coding the samples of the coding tree block. The coding tree block can be an N×N block of samples. A CTU can also be referred to as a tree block or a largest coding unit (LCU). The CTU of HEVC can be generally similar to the macroblock of other standards such as H.264 / AVC. However, the CTU is not necessarily limited to a specific size and can include one or more coding units (CUs). A slice can include an integer number of CTUs sorted consecutively in raster scan order.
[0066] In HEVC, a quadtree structure represented as a coding tree is used to partition a CTU into CUs to adapt to different local characteristics. At the CU level, it is determined whether to code an image region by inter-frame (temporal) prediction or by intra-frame (spatial) prediction. A CU can include a coding block of luminance samples of an image, two corresponding coding blocks of chrominance samples, and a syntax structure for coding the samples of the coding block, where the image has an array of luminance samples, an array of Cb samples, and an array of Cr samples. In a monochrome image or an image with three independent color planes, a CU can include a single coding block and a syntax structure for coding the samples of the coding block. The coding block is an N×N block of samples. In some examples, the size of a CU can be the same as that of a CTU. Each CU is coded with a coding mode, which can be, for example, an intra coding mode or an inter coding mode. Other coding modes can also be used. The encoder 200 receives video data. The encoder 200 can encode each CTU in a slice of the image of the video data. As part of encoding the CTU, the prediction processing unit 260 of the encoder 200 or other processing units (including but not limited to Figure 2 the units of the encoder 200 shown in
[0067] The syntax data in the bitstream can also define the size of the CTU. A slice includes a plurality of consecutive CTUs arranged in decoding order. A video frame or image can be segmented into one or more slices. As described above, each tree block can be segmented into coding units (CUs) according to a quadtree. Generally, the quadtree data structure includes one node for each CU, and the root node corresponds to the tree block (e.g., CTU). If a CU is segmented into four sub-CUs, the node corresponding to the CU includes 4 child nodes, and each child node corresponds to a sub-CU. The multiple nodes of the quadtree structure include leaf nodes and non-leaf nodes. A leaf node has no child nodes in the tree structure (i.e., the leaf node is not further divided). The non-leaf nodes include the root node of the tree structure. For each corresponding non-root node among the multiple nodes, the corresponding non-root node corresponds to the sub-CU of the CU, and the sub-CU of the CU corresponds to the parent node in the tree structure of the corresponding non-root node. Each corresponding non-leaf node has one or more child nodes in the tree structure.
[0068] Each node of the quadtree data structure can provide syntax data for the corresponding CU. For example, a node in the quadtree can include a partitioning flag indicating whether the CU corresponding to the node has been partitioned into sub-CUs. The syntax elements of the CU can be defined recursively and can depend on whether the CU has been partitioned into sub-CUs. If the CU is not further partitioned, the CU is called a leaf CU. If the block of the CU is further partitioned, the CU can generally be called a non-leaf CU. Each level of partitioning is a quadtree partitioning into four sub-CUs. The black CU is an example of a leaf node (i.e., a block that is not further divided).
[0069] The role of the CU is similar to that of the macroblock in the H.264 standard, except that the CU has no size distinction. For example, a tree block can be divided into four child nodes (also called sub-CUs), and each child node can in turn be the parent node and be divided into another four child nodes. The final undivided child nodes are called the leaf nodes of the quadtree and include the decoding nodes, also called leaf CUs. The syntax data related to the decoded bitstream can define the maximum number of times of partitioning the tree block, called the maximum CU depth, and can also define the minimum size of the decoding node. Correspondingly, the bitstream can also define the smallest coding unit (SCU). The term "block" refers to any one of the CU, PU, or TU in the HEVC context, or a similar data structure in other standard contexts (e.g., the macroblock and its sub-blocks in H.264 / AVC).
[0070] In HEVC, each CU can also be divided into one, two, or four PUs according to the PU partition type. Within a PU, the same prediction process is performed, and relevant information is sent to the decoder in units of PUs. After obtaining the residual block through the prediction process, according to the PU partition type, the CU can be divided into transform units (TUs) according to another quadtree structure similar to the decoding tree used for the CU. A key feature of the HEVC structure is that it has multiple partitioning concepts such as CUs, PUs, and TUs. The PU can be partitioned into a non-square shape. The syntax data related to the CU can also describe, for example, dividing the CU into one or more PUs. The shape of the TU can be square or non-square (e.g., rectangular), and the syntax data related to the CU can describe, for example, dividing the CU into one or more TUs according to the quadtree. When the CU is encoded using the skip mode, direct mode, intra prediction mode, or inter prediction mode, the partitioning mode may be different.
[0071] In VVC (Versatile Video Coding), regardless of the differences in the concepts of PUs and TUs, multiple CU partition shapes are supported. The size of the CU corresponds to the size of the decoding node and can be square or non-square (e.g., rectangular). The size of the CU can range from 4×4 pixels (or 8×8 pixels) to the size of the tree block, with a maximum of 128×128 pixels or larger (e.g., 256×256 pixels).
[0072] After the encoder 200 generates the prediction blocks of the CU (e.g., the luminance, Cb, and Cr prediction blocks), the encoder 200 can generate the residual block of the CU. For example, the encoder 100 can generate the luminance residual block of the CU. Each sample in the luminance residual block of the CU represents the difference between the luminance sample in the predicted luminance block of the CU and the corresponding sample in the original luminance decoding block of the CU. In addition, the encoder 200 can generate the Cb residual block of the CU. Each sample in the Cb residual block of the CU can represent the difference between the Cb sample in the predictive Cb block of the CU and the corresponding sample in the original Cb decoding block of the CU. The encoder 200 can also generate the Cr residual block of the CU. Each sample in the Cr residual block of the CU can represent the difference between the Cr sample in the predicted Cr block of the CU and the corresponding sample in the original Cr decoding block of the CU.
[0073] In some examples, the encoder 200 does not perform a transform on the transform block. In such examples, the encoder 200 can process the residual sample values in the same way as the transform coefficients. Therefore, in the examples where the encoder 200 does not perform a transform, the following discussions about the transform coefficients and coefficient blocks can apply to the transform blocks of the residual samples.
[0074] After generating a coefficient block (e.g., a luminance coefficient block, a Cb coefficient block, or a Cr coefficient block), the encoder 200 may quantize the coefficient block to minimize the amount of data used to represent the coefficient block, thereby performing further compression. Quantization generally refers to the process of reducing a series of values to a single value. After the encoder 200 quantizes the coefficient block, the encoder 200 may perform entropy coding on the syntax elements representing the quantized transform coefficients. For example, the encoder 200 may perform Context-Adaptive Binary Arithmetic Coding or other entropy decoding techniques on the syntax elements representing the quantized transform coefficients.
[0075] The encoder 200 may output a bitstream of the encoded image data 271, which includes a column of bits that form a representation of the decoded image and associated data. Thus, the bitstream includes an encoded representation of the video data.
[0076] In "Block Partitioning Structure for Next Generation Video Coding" by J. An et al. (International Telecommunication Union, COM16-C966, September 2015 (hereinafter referred to as "VCEG Proposal COM16-C966")), a quad-tree-binary-tree (QTBT) partitioning technique is proposed for video coding standards beyond the future HEVC. Simulation results show that the proposed QTBT structure is more efficient than the quad-tree structure used in HEVC. In HEVC, to reduce memory access for motion compensation, the inter prediction for small blocks is restricted. Therefore, bi-directional prediction for 4×8 and 8×4 blocks and inter prediction for 4×4 blocks are not supported. In the QTBT of JEM, these restrictions are removed.
[0077] In QTBT, a CU can be square or rectangular. For example, the coding tree unit (CTU) is first segmented by a quadtree structure. The quadtree leaf nodes can be further segmented by a binary tree structure. There are two types of binary tree segmentation: symmetric horizontal segmentation and symmetric vertical segmentation. In each case, the node is segmented horizontally or vertically from the middle downwards. The binary tree leaf nodes are called coding units (CUs), and this segmentation is used for prediction and transform processing without further segmentation. That is, the CUs, PUs, and TUs have the same block size in the QTBT coding block structure. A CU can be composed of coding blocks (CBs) of different color components. For example, in the case of P and B slices in 4:2:0 chroma format, a CU includes one luma CB and two chroma CBs; or a CU can be composed of CBs of a single component. For example, a CU includes only one luma CB or only two chroma CBs in the case of I slices.
[0078] The following parameters are defined for the QTBT segmentation scheme:
[0079] – CTU size: The size of the root node of the quadtree, which is the same concept as in HEVC.
[0080] – MinQTSize: The minimum allowable size of the quadtree leaf node.
[0081] – MaxBTSize: The maximum allowable size of the binary tree root node.
[0082] – MaxBTDPepth: The maximum allowable depth of the binary tree.
[0083] – MinBTSize: The minimum allowable size of the binary tree leaf node.
[0084] In an example of the QTBT splitting structure, the CTU size is set to 128 × 128 luma samples, where a block with two corresponding 64 × 64 chroma samples, MinQTSize is set to 16 × 16, MaxBTSize is set to 64 × 64, MinBTSize (width and height) is set to 4 × 4, and MaxBTDepth is set to 4. Quadtree splitting is first applied to the CTU to generate quadtree leaf nodes. The size of the quadtree leaf nodes can be from 16 × 16 (i.e., MinQTSize) to 128 × 128 (i.e., CTU size). When the size of the quadtree node is equal to MinQTSize, further quadtree is not considered. If the leaf quadtree node is 128 × 128, its size exceeds MaxBTSize (i.e., 64 × 64), so it is not further divided by the binary tree. Otherwise, the leaf quadtree node can be further split by the binary tree. Therefore, the quadtree leaf node is also the root node of the binary tree, and its binary tree depth is 0. When the binary tree depth reaches MaxBTDepth (i.e., 4), further division is not considered. When the width of the binary tree node is equal to MinBTSize (i.e., 4), further horizontal division is not considered. Similarly, when the height of the binary tree node is equal to MinBTSize, further vertical division is not considered. Prediction and transformation processing are performed on the leaf nodes of the binary tree without further splitting. In JEM, the maximum CTU size is 256 × 256 luma samples. Further processing (e.g., by performing a prediction process and a transformation process) can be performed on the leaf nodes of the binary tree (binary-tree, CU) without further splitting.
[0085] In addition, in the QTBT scheme, luma and chroma have separate QTBT structures. Currently, for P and B slices, the luma and chroma CTBs in a CTU can share the same QTBT structure. However, for I slices, the luma CTB is split into CUs by the QTBT structure, and the chroma CTB can be split into chroma CUs by another QTBT structure. That is, the CUs in I slices consist of decoded blocks of the luma component or decoded blocks of two chroma components, while the CUs in P slices or B slices consist of decoded blocks of all three color components.
[0086] The encoder 200 applies a rate-distortion optimization (RDO) process for the QTBT structure to determine the block splitting.
[0087] In addition, a block partitioning structure called multi-type-tree (MTT) was proposed in the U.S. Patent Application Publication No. 20170208336 to replace the CU structure based on QT, BT, and / or QTBT. The MTT partitioning structure is still a recursive tree structure. In MTT, multiple different partitioning structures (e.g., three or more) are used. For example, according to the MTT technique, at each depth of the tree structure, three or more different partitioning structures can be used for each corresponding non-leaf node of the tree structure. The depth of a node in the tree structure can refer to the length of the path from the node of the tree structure to the root (e.g., the number of divisions). The partitioning structure generally can refer to how many different blocks a block can be divided into. The partitioning structure can be a quadtree partitioning structure that divides a block into four blocks, a binary tree partitioning structure that divides a block into two blocks, or a triple-tree (TT) partitioning structure that divides a block into three blocks. In addition, the triple-tree partitioning structure can divide the block without passing through the center. The partitioning structure can have multiple different partitioning types. The partitioning type can also define the way to divide the block, including symmetric or asymmetric division, uniform or non-uniform division, and / or horizontal or vertical division.
[0088] In MTT, at each depth of the tree structure, the encoder 200 can be used to further divide the subtree, and the further division uses a specific partitioning type of one of the partitioning structures with more than three. For example, the encoder 100 can be used to determine the specific partitioning type according to QT, BT, triple-tree (TT), and other partitioning structures. In one example, the QT partitioning structure can include a square quadtree or a rectangular quadtree partitioning type. The encoder 200 can use the square quadtree partitioning to divide a square block by horizontally and vertically dividing the block into four equal-sized square blocks along the center. Similarly, the encoder 200 can use the rectangular quadtree partitioning to divide a rectangular (e.g., non-square) block by horizontally and vertically dividing the rectangular block into four equal-sized rectangular blocks along the center.
[0089] The BT segmentation structure may include at least one of a horizontal symmetric binary tree, a vertical symmetric binary tree, a horizontal asymmetric binary tree, and a vertical asymmetric binary tree. For the horizontal symmetric binary tree segmentation type, the encoder 200 may be used to horizontally divide a block into two symmetric blocks of the same size along the center. For the vertical symmetric binary tree segmentation type, the encoder 200 may be used to vertically divide a block into two symmetric blocks of the same size along the center. For the horizontal asymmetric binary tree segmentation type, the encoder 200 may be used to horizontally divide a block into two blocks of different sizes. For example, the size of one block may be 1 / 4 of the parent block, and the size of the other block may be 3 / 4 of the parent block, similar to the PART_2N×nU or PART_2N×nD segmentation type. For the vertical asymmetric binary tree segmentation type, the encoder 200 may be used to vertically divide a block into two blocks of different sizes. For example, the size of one block may be 1 / 4 of the parent block, and the size of the other block may be 3 / 4 of the parent block, similar to the PART_nL×2N or PART_nR×2N segmentation type. In other examples, the asymmetric binary tree segmentation type may divide the parent block into parts of different sizes. For example, one sub-block may be 3 / 8 of the parent block, and the other sub-block may be 5 / 8 of the parent block. Of course, such a segmentation type may be vertical or horizontal.
[0090] The TT segmentation structure is different from the QT or BT structure. The TT segmentation structure does not divide the block along the center. The central region of the block remains in the same sub-block. Different from the QT that produces four blocks or the binary tree that produces two blocks, the division according to the TT segmentation structure produces three blocks. Example segmentation types according to the TT segmentation structure include symmetric segmentation types (horizontal and vertical) and asymmetric segmentation types (horizontal and vertical). In addition, according to the TT segmentation structure, the symmetric segmentation type may be uneven / non-uniform or even / uniform. According to the TT segmentation structure, the asymmetric segmentation type is uneven. In one example, the TT segmentation structure may include at least one of the following segmentation types: horizontal even symmetric ternary tree segmentation type, vertical even symmetric ternary tree segmentation type, horizontal uneven symmetric ternary tree segmentation type, vertical uneven symmetric ternary tree segmentation type, horizontal uneven asymmetric ternary tree segmentation type, or vertical uneven asymmetric ternary tree segmentation type.
[0091] Generally, the uneven symmetric ternary tree segmentation type is a segmentation type that is symmetric about the center line of the block, but among them, the size of at least one of the three obtained blocks is different from the sizes of the other two blocks. In a preferred example, the sizes of the two side blocks are both 1 / 4 of the block, while the size of the central block is 1 / 2 of the block. The uniform symmetric ternary tree segmentation type is a segmentation type that is symmetric about the center line of the block, and the obtained blocks have the same size. This segmentation is possible if the height or width of the block (depending on vertical or horizontal segmentation) is a multiple of 3. The uneven asymmetric ternary tree segmentation type refers to a segmentation type that is not symmetric about the center line of the block, and among them, the size of at least one of the obtained blocks is different from the sizes of the other two blocks.
[0092] In an example of dividing a block (e.g., at a subtree node) into an asymmetric ternary tree segmentation type, the encoder 200 and / or the decoder 300 can apply a restriction such that the sizes of two of the three divided parts are the same. This restriction can correspond to a restriction that the encoder 200 must abide by when encoding video data. In addition, in some examples, the encoder 200 and the decoder 300 can apply a restriction, so that when dividing according to the asymmetric ternary tree segmentation type, the sum of the areas of two of the divided parts is equal to the area of the remaining one divided part.
[0093] In some examples, the encoder 200 can be used to select a segmentation type from all the above-mentioned segmentation types for each of the QT, BT, and TT segmentation structures. In other examples, the encoder 200 can be used to determine a segmentation type only from a subset of the foregoing segmentation types. For example, a subset of the above-mentioned segmentation types (or other segmentation types) can be used for a specific block size or a specific depth of the quadtree structure. The subset of the supported segmentation types can be indicated (signaled) in the bitstream for use by the decoder 200, or can be predefined so that the encoder 200 and the decoder 300 can determine the subset without any indication.
[0094] In other examples, for all depths in all CTUs, the number of supported segmentation types can be fixed. That is, the encoder 200 and the decoder 300 can be preconfigured to use the same number of segmentation types for any depth of the CTU. In other examples, the number of supported segmentation types can vary and can depend on the depth, slice type, or other previously decoded information. In one example, at depth 0 or depth 1 of the tree structure, only the QT segmentation structure is used. At depths greater than 1, each of the QT, BT, and TT segmentation structures can be used.
[0095] In some examples, the encoder 200 and / or the decoder 300 may have preconfigured restrictions on the supported partitioning types to avoid redundant partitioning of a region of a video image or a region of a CTU. In one example, when partitioning a block with an asymmetric partitioning type, the encoder 200 and / or the decoder 300 may not further partition the largest sub-block obtained from the current block partitioning. For example, when partitioning a square block according to an asymmetric partitioning type (similar to the PART_2N×nU partitioning type), the largest sub-block among all sub-blocks (the largest sub-block partitioning type similar to PART_2N×nU) is the marked leaf node and cannot be further partitioned. However, the smaller sub-blocks (similar to the smaller sub-blocks of the PART_2N×nU partitioning type) can be further partitioned.
[0096] As another example, where the supported partitioning types may be restricted to avoid redundant partitioning of a specific region, when partitioning a block with an asymmetric partitioning type, the largest sub-block obtained from the current block cut cannot be further partitioned in the same direction. For example, when a square block has an asymmetric partitioning type (similar to the PART_2N×nU partitioning type), the encoder 200 and / or the decoder 300 may not partition the large sub-block among all sub-blocks (the largest sub-block similar to the PART_2N×nU partitioning type) in the horizontal direction.
[0097] As another example, where the supported partitioning types may be restricted to facilitate further partitioning, when the width / height of a block is not a power of 2 (e.g., when the width and height are not 2, 4, 8, 16, etc.), the encoder 200 and / or the decoder 300 may not partition the block horizontally or vertically.
[0098] The above examples describe how the encoder 200 may perform MTT partitioning. Then, the decoder 300 may also perform the same MTT partitioning as the partitioning performed by the encoder 200. In some examples, the way the encoder 200 partitions the image of the video data may be determined by applying the same set of predefined rules at the decoder 300. However, in many cases, the encoder 200 may determine the specific partitioning structure and partitioning type to use according to the rate-distortion criterion for the specific image of the video data being decoded. Therefore, to enable the decoder 300 to determine the partitioning of a specific image, the encoder 200 may indicate in the encoded bitstream a syntax element that represents the way to partition the image and the CTUs of the image. The decoder 200 may parse such a syntax element and partition the image and CTUs accordingly.
[0099] In one example, the prediction processing unit 260 of the video encoder 200 may be used to perform any combination of the above partitioning techniques, especially for motion estimation, which will be described in detail below.
[0100] Although the size of block 203 is smaller than that of image 201, like image 201, block 203 is also or can also be considered as a two-dimensional array or matrix of samples having intensity values (sample values). In other words, image block 203 can include, for example, an array of samples (e.g., a luminance array in the case of a black-and-white image 201), three arrays of samples (e.g., a luminance array and two chrominance arrays in the case of a color image 201), or any other number and / or type of arrays, depending on the color format applied. The number of samples of block 203 in the horizontal and vertical directions (or axes) defines the size of block 203.
[0101] As Figure 2 shown, encoder 200 is used to encode image 201 block by block, for example, performing encoding and prediction on each block 203.
[0102] Residual calculation
[0103] Residual calculation unit 204 is used to calculate residual block 205 based on image block 203 and prediction block 265 (prediction block 265 will be described in detail later), for example, subtracting the sample values of prediction block 265 from the sample values of image block 203 sample by sample (pixel by pixel) to obtain residual block 205 in the sample domain.
[0104] Transformation
[0105] Transformation processing unit 206 is used to transform the sample values of residual block 205, such as discrete cosine transform (DCT) or discrete sine transform (DST), to obtain transform coefficients 207 in the transform domain. Transform coefficients 207 can also be referred to as transform residual coefficients and represent residual block 205 in the transform domain.
[0106] The transform processing unit 206 can be used to apply integer approximations of DCT / DST, such as the transforms specified for HEVC / H.265. Such integer approximations are typically scaled by a certain factor compared to the orthogonal DCT transform. Other scaling factors are used as part of the transform process to maintain the norm of the residual block after forward and inverse transform processing. The scaling factors are typically selected according to certain constraints, such as the power of 2 for shift operations, the bit depth of the transform coefficients, the trade-off between precision and implementation cost, etc. For example, a specific scaling factor for the inverse transform is specified on the decoder 300 side by the inverse transform processing unit 212, etc. (and a corresponding scaling factor for the inverse transform is specified on the encoder 200 side by the inverse transform processing unit 212, etc.), and a corresponding scaling factor for the forward transform can be specified on the encoder 200 side by the transform processing unit 206, etc.
[0107] Quantization
[0108] Quantization unit 208 is used to quantize transform coefficients 207 (e.g., perform scalar quantization or vector quantization) to obtain quantized transform coefficients 209. The quantized transform coefficients 209 can also be referred to as quantized residual coefficients 209. The quantization process can reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient can be rounded down to an m-bit transform coefficient during quantization, where n is greater than m, and the degree of quantization can be modified by adjusting the quantization parameter (QP). For example, for scalar quantization, different scalings can be performed to achieve finer or coarser quantization. The smaller the quantization step size, the finer the quantization; the larger the quantization step size, the coarser the quantization. A suitable quantization step size can be indicated by the quantization parameter (QP). For example, the quantization parameter can be an index applicable to a predefined set of suitable quantization step sizes. For example, a small quantization parameter can correspond to fine quantization (small quantization step size), while a large quantization parameter can correspond to coarse quantization (large quantization step size), and vice versa. The quantization operation can include dividing by the quantization step, and the corresponding dequantization or inverse dequantization operation performed by dequantization unit 210 and the like can include multiplying by the quantization step. According to some standards (such as HEVC), the quantization parameter can be used in embodiments to determine the quantization step size. Generally, the quantization step size can be calculated based on the quantization parameter by a fixed-point approximation of an equation including division. Other scaling factors can be introduced into quantization and dequantization to restore the norm of the residual block. Since scaling is used in the fixed-point approximation of the equation for the quantization step size and the quantization parameter, this norm can be modified. In one exemplary implementation, the scaling in the inverse transform and dequantization can be combined. Alternatively, a custom quantization table can be used and indicated from the encoder to the decoder in the bitstream and the like. Quantization is a lossy operation, and the loss increases with the increase of the quantization step size.
[0109] Dequantization unit 210 is used to perform dequantization of the quantization performed by quantization unit 208 on the quantized coefficients to obtain dequantized coefficients 211. For example, perform an inverse dequantization scheme of the quantization scheme performed by quantization unit 208 according to or using the same quantization step as quantization unit 208. The dequantized coefficients 211 can also be referred to as dequantized residual coefficients 211, which correspond to the transform coefficients 207. However, due to the loss caused by quantization, the dequantized coefficients 211 are usually not exactly the same as the transform coefficients.
[0110] The inverse transform processing unit 212 is configured to perform an inverse transform on the transform performed by the transform processing unit 206, such as an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain an inverse transform block 213 in the sample domain. The inverse transform block 213 may also be referred to as an inverse transform dequantization block 213 or an inverse transform residual block 213.
[0111] The reconstruction unit 214 (e.g., adder 214) is configured to add the inverse transform block 213 (i.e., the reconstruction residual block 213) to the prediction block 265 to obtain a reconstruction block 215 in the sample domain, e.g., by adding the sample values of the reconstruction residual block 213 to the sample values of the prediction block 265.
[0112] Optionally, a buffer unit 216 (or buffer 216), such as a line buffer 216, is configured to buffer or store the reconstruction block 215 and corresponding sample values, e.g., for intra prediction. In other embodiments, the encoder may use the unfiltered reconstruction block and / or the corresponding sample values stored in the buffer unit 216 for any type of estimation and / or prediction, such as intra prediction.
[0113] An embodiment of the encoder 200 may cause the cache unit 216 to be used not only for storing the reconstruction block 215 for intra prediction 254, but also for the loop filter unit 220 ( Figure 2 not shown in), and / or cause, for example, the buffer unit 216 and the decoded image buffer unit 230 to form a buffer. Other embodiments may use the filtered block 221 of the decoded image buffer 230 and / or blocks or samples (blocks or samples not shown in Figure 2 as the input or basis for intra prediction 254.
[0114] The loop filter unit 220 (or loop filter 220) is configured to filter the reconstruction block 215 to obtain a filtered block 221, e.g., to smooth pixel transitions or improve video quality. The loop filter unit 220 represents one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, and other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening or smoothing filter, or a collaborative filter. Although the loop filter unit 220 is shown as an in-loop filter in Figure 2 , in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as a filtered reconstruction block 221. The decoded image buffer 230 may store the reconstructed decoded block after the loop filter unit 220 performs a filtering operation on the reconstructed decoded block.
[0115] For example, an embodiment of the encoder 200 (and correspondingly, the loop filter unit 220) may directly output loop filter parameters (e.g., sample adaptive offset information), or output the loop filter parameters after entropy coding by the entropy coding unit 270 or any other entropy decoding unit, such that, for example, the decoder 300 can receive the same loop filter parameters and apply the same loop filter parameters to decoding.
[0116] The decoded picture buffer (DPB) 230 may be a reference image memory that stores reference image data used by the video encoder 20 to encode video data. The DPB 230 may be formed of any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magneto resistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The DPB 230 and the buffer 216 may be provided by the same memory device or by different memory devices. In some examples, the decoded picture buffer (DPB) 230 is used to store the filtered block 221. The decoded image buffer 230 may also be used to store other previously filtered blocks (e.g., previously reconstructed and filtered blocks 221) of the same current image or different images (e.g., previously reconstructed images), and may provide a complete previously reconstructed (i.e., decoded) image (and corresponding reference blocks and samples) and / or a partially reconstructed current image (and corresponding reference blocks and samples) for, e.g., inter prediction. In some examples, if the reconstructed block 215 is reconstructed but not loop-filtered, the decoded picture buffer (DPB) 230 is used to store the reconstructed block 215.
[0117] The prediction processing unit 260, also referred to as the block prediction processing unit 260, is configured to: receive or obtain a block 203 (e.g., the current block 203 of the current image 201) and reconstructed image data, e.g., reference samples of the same (or current) image from the buffer 216 and / or reference image data 231 of one or more previously decoded images from the decoded picture buffer 230; and process such data for prediction, i.e., provide a prediction block 265, which may be an inter prediction block 245 or an intra prediction block 255.
[0118] The mode selection unit 262 can be used to select a prediction mode (e.g., an intra or inter prediction mode) and / or select a corresponding prediction block 245 or 255 as the prediction block 265 for the calculation of the residual block 205 and the reconstruction of the reconstruction block 215.
[0119] An embodiment of the mode selection unit 262 can be used to select a prediction mode (e.g., from those prediction modes supported by the prediction processing unit 260) that provides the best match, or in other words, provides the smallest residual (a smaller residual is more beneficial for compression for transmission or storage), or provides the smallest indication overhead (a smaller indication overhead is more beneficial for compression for transmission or storage), or considers or balances both. The mode selection unit 262 can be used to determine the prediction mode according to rate - distortion optimization (RDO), i.e., determine a prediction mode that provides the smallest rate - distortion optimization, or the related rate - distortion at least meets the prediction mode selection criteria.
[0120] The prediction processing (e.g., by the prediction processing unit 260) and mode selection (e.g., by the mode selection unit 262) performed by the example encoder 200 are explained in more detail below.
[0121] As described above, the encoder 200 is used to determine or select the best or optimal prediction mode from a set (predetermined) of prediction modes. The set of prediction modes can include, for example, intra prediction modes and / or inter prediction modes.
[0122] The set of intra prediction modes can include 35 different intra prediction modes, such as non - directional modes like DC (or mean) mode and planar mode, or directional modes defined in H.265, etc., or can include 67 different intra prediction modes, such as non - directional modes like DC (or mean) mode and planar mode, or directional modes defined in the currently developing H.266, etc.
[0123] The set (or possible) of inter prediction modes depends on the available reference images (i.e., for example, previously at least partially decoded images stored in the DBP 230) and other inter prediction parameters, such as whether the entire reference image or only a part of the reference image (e.g., the search window region around the region of the current block) is used to search for the best - matching reference block, and / or for example, whether pixel interpolation is applied, such as half - pixel and / or quarter - pixel interpolation, or no pixel interpolation is applied.
[0124] In addition to the above - mentioned prediction modes, skip mode and / or direct mode can also be applied.
[0125] The prediction processing unit 260 can also be used, for example, to divide the block 203 into smaller block parts or sub-blocks by iteratively using quad-tree partitioning (QT), binary partitioning (BT), or triple-tree partitioning (TT), or any combination thereof, and to perform predictions, for example, on each of the block parts or sub-blocks, where the mode selection includes selecting the tree structure of the divided block 203 and the prediction mode applied to each block part or sub-block.
[0126] The inter-frame prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit ( Figure 2 not shown). The motion estimation unit is used to receive or obtain the image block 203 (the current image block 203 of the current image 201) and the decoded image 331, or at least one or more previously reconstructed blocks (e.g., the reconstructed blocks of one or more other / different previously decoded images 331), for motion estimation. For example, the video sequence may include the current image and the previously decoded image 331, or, in other words, the current image and the previously decoded image 331 may be part of the image sequence that constitutes the video sequence or may constitute the sequence. For example, the encoder 200 may be used to select a reference block from a plurality of reference blocks of the same or different images of a plurality of other images and provide the reference image (or reference image index...) and / or the offset (spatial offset) between the position (x coordinate, y coordinate) of the reference block and the position of the current block, as the inter-frame prediction parameter of the motion estimation unit ( Figure 2 not shown). This offset is also referred to as a motion vector (MV). Fusion is an important motion estimation tool used in HEVC and is also used in VVC. To perform fusion estimation, a fusion candidate list is first constructed, where each candidate includes all motion data, which includes information on whether to use one or two reference image lists and the reference index and motion vector of each list. The fusion candidate list is constructed based on the following candidates: 1. Up to four spatial fusion candidates, which are obtained from five spatially adjacent (i.e., neighboring) blocks; 2. One temporal fusion candidate, which is obtained from two temporal, collocated blocks; 3. Other fusion candidates, including combined bidirectional prediction candidates and zero motion vector candidates.
[0127] The intra prediction unit 254 is further configured to determine an intra prediction block 255 according to intra prediction parameters (e.g., the selected intra prediction mode). In any case, after an intra prediction mode is selected for a block, the intra prediction unit 254 is further configured to provide the intra prediction parameters, i.e., information indicating the selected intra prediction mode for the block, to the entropy coding unit 270. In one example, the intra prediction unit 254 may be configured to perform any combination of the intra prediction techniques described below.
[0128] The entropy coding unit 270 is configured to apply an entropy coding algorithm or scheme (e.g., variable length coding (VLC) scheme, context adaptive VLC (CAVLC) scheme, arithmetic coding scheme, context adaptive binary arithmetic coding (CABAC) scheme, syntax-based context-adaptive binary arithmetic coding (SBAC) scheme, probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to the quantized residual coefficients 209, inter prediction parameters, intra prediction parameters, and / or loop filter parameters, either individually or jointly (or not at all), to obtain the encoded image data 21 that can be output from the output terminal 272, e.g., output in the form of an encoded bitstream 21. The encoded bitstream 21 can be sent to the video decoder 30, or archived for subsequent transmission or retrieval by the video decoder 30. The entropy coding unit 270 may also be configured to perform entropy coding on other syntax elements of the current video slice being decoded.
[0129] Other structural variations of the video encoder 200 may be used to encode the video stream. For example, for some blocks or frames, the non-transform-based encoder 200 may directly quantize the residual signal without the transform processing unit 206. In another implementation, the encoder 200 may combine the quantization unit 208 and the inverse quantization unit 210 into a single unit.
[0130] Figure 3 An exemplary video decoder 300 for implementing the techniques of the present invention is shown. The video decoder 300 is configured to receive encoded image data (e.g., an encoded bitstream) 271 encoded, for example, by the encoder 200, to obtain a decoded image 331. During the decoding process, the video decoder 300 receives video data from the video encoder 200, e.g., an encoded video bitstream representing image blocks of an encoded video slice and associated syntax elements.
[0131] In Figure 3 this example, the decoder 300 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., an adder 314), a buffer 316, a loop filter 320, a decoded picture buffer 330, and a prediction processing unit 360. The prediction processing unit 360 may include an inter prediction unit 344, an intra prediction unit 354, and a mode selection unit 362. In some examples, the video decoder 300 may perform a decoding channel that is substantially inverse to the encoding channel described with respect to the video encoder 200 in Figure 2 .
[0132] The entropy decoding unit 304 is configured to perform entropy decoding on the encoded picture data 271 to obtain, for example, quantization coefficients 309 and / or decoded syntax parameters ( Figure 3 not shown in), such as any or all of (decoded) inter prediction parameters, intra prediction parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 is further configured to forward the inter prediction parameters, intra prediction parameters, and / or other syntax elements to the prediction processing unit 360. The video decoder 300 may receive syntax elements at the video slice level and / or the video block level.
[0133] The inverse quantization unit 310 may have the same function as the inverse quantization unit 110, the inverse transform processing unit 312 may have the same function as the inverse transform processing unit 112, the reconstruction unit 314 may have the same function as the reconstruction unit 114, the buffer 316 may have the same function as the buffer 116, the loop filter 320 may have the same function as the loop filter 120, and the decoded picture buffer 330 may have the same function as the decoded picture buffer 130.
[0134] The prediction processing unit 360 may include an inter prediction unit 344 and an intra prediction unit 354, where the inter prediction unit 344 is functionally similar to the inter prediction unit 144, and the intra prediction unit 354 is functionally similar to the intra prediction unit 154. The prediction processing unit 360 is generally configured to perform block prediction and / or obtain a prediction block 365 from the encoded data 21, and receive or obtain (explicitly or implicitly), for example, prediction-related parameters and / or information about the selected prediction mode from the entropy decoding unit 304.
[0135] When a video slice is decoded as an intra-coded (I) slice, the intra prediction unit 354 of the prediction processing unit 360 is configured to: generate a prediction block 365 for an image block of the current video slice according to the indicated intra prediction mode and data of previously decoded blocks in the current frame or picture. When a video frame is decoded as an inter-coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., motion compensation unit) of the prediction processing unit 360 is configured to generate a prediction block 365 for a video block of the current video slice according to motion vectors and other syntax elements received from the entropy decoding unit 304. For inter prediction, the prediction block may be generated from a reference image in one of the reference image lists. The video decoder 300 may construct reference frame lists, list 0 and list 1, according to the reference images stored in the DPB 330 through a default construction technique.
[0136] The prediction processing unit 360 is configured to determine prediction information for a video block of the current video slice by parsing motion vectors and other syntax elements, and use the prediction information to generate a prediction block for the current video block being decoded. For example, the prediction processing unit 360 uses some received syntax elements to determine a prediction mode (e.g., intra or inter prediction) for decoding a video block of the video slice, an inter prediction slice type (e.g., B slice, P slice or GPB slice), construction information for one or more of the reference image lists of the slice, motion vectors for each inter-coded video block of the slice, the inter prediction state for each inter-coded video block of the slice, and other information, to decode the video blocks in the current video slice.
[0137] The inverse quantization unit 310 may be configured to inverse-quantize (i.e., de-quantize) the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 304. The inverse quantization process may include: using the quantization parameter of each video block in the video slice calculated by the video encoder 100 to determine the degree of quantization and the degree of inverse quantization to be applied.
[0138] The inverse transform processing unit 312 is configured to apply an inverse transform to the transform coefficients, e.g., inverse DCT, inverse integer transform or a conceptually similar inverse transform process, to generate a residual block in the pixel domain.
[0139] The reconstruction unit 314 (e.g., adder 314) is configured to add the inverse transform block 313 (i.e., the reconstructed residual block 313) to the prediction block 365 to obtain a reconstructed block 315 in the sample domain, e.g., by adding the sample values of the reconstructed residual block 313 to the sample values of the prediction block 365.
[0140] The loop filter unit 320 (in or after the decoding loop) is used to filter the reconstructed block 315 to obtain a filtered block 321, for example, to smooth the abrupt change of pixels or improve the video quality. In one example, the loop filter unit 320 can be used to perform any combination of the filtering techniques described below. The loop filter unit 320 represents one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, and other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening or smoothing filter, or a collaborative filter. Although the loop filter unit 320 is shown as an in-loop filter in Figure 3 , in other configurations, the loop filter unit 320 can be implemented as a post-loop filter.
[0141] Then, the decoded video block 321 in a given frame or image is stored in the decoded image buffer 330, and the decoded image buffer 330 stores the reference image for subsequent motion compensation.
[0142] The decoder 300 is used to output the decoded image 311 through the output terminal 312, for example, to be presented to or viewed by the user.
[0143] Other variants of the video decoder 300 can be used to decode the compressed bitstream. For example, the decoder 300 can generate an output video stream without the loop filter unit 320. For example, for some blocks or frames, the non-transform-based decoder 300 can directly quantize the residual signal without the inverse transform processing unit 312. In another implementation, the video decoder 300 can combine the inverse quantization unit 310 and the inverse transform processing unit 312 into a single unit.
[0144] Figure 4 Schematic diagram of a network device 400 (e.g., a decoding device) provided for an embodiment of the present invention. The network device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the network device 400 can be a decoder (e.g., Figure 1A the video decoder 30 in Figure 1A ) or an encoder (e.g., Figure 1A the video encoder 20 in Figure 1A ). In one embodiment, the network device 400 can be one or more components of the video decoder 30 in
[0145] The network device 400 includes: an ingress port 410 and a receive unit (Rx) 420 for receiving data; a processor, logic unit, or central processing unit (CPU) 430 for processing data; a transmitter unit (Tx) 440 and an egress port 450 for transmitting data; and a memory 460 for storing data. The network device 400 may further include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the ingress port 410, the receiver unit 420, the transmitter unit 440, and the egress port 450 for the ingress or egress of optical or electrical signals.
[0146] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 430 communicates with the ingress port 410, the receive unit 420, the transmitter unit 440, the egress port 450, and the memory 460. The processor 430 includes a decoding module 470. The decoding module 470 implements the embodiments disclosed above. For example, the decoding module 470 implements, processes, prepares, or provides various decoding operations. Thus, having the decoding module 470 can greatly enhance the functionality of the network device 400 and affect the transition of the network device 400 to different states. Alternatively, the decoding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.
[0147] The memory 460 includes one or more disks, tape drives, and solid-state drives and may be used as an overflow data storage device to store programs when such programs are selected for execution, as well as instructions and data read during program execution. The memory 460 may be volatile and / or non-volatile memory and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0148] Figure 5 A simplified block diagram of a device 500 adopted for an exemplary embodiment, where the device 500 may be used as Figure 1AOne or both of the source device 12 and the destination device 14 in []. The apparatus 500 may implement the techniques of the present invention. The apparatus 500 may be in the form of a computing system including a plurality of computing devices or in the form of a single computing device, such as a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, etc.
[0149] The processor 502 in the apparatus 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or devices capable of manipulating or processing information that exists now or will be developed later. Although the disclosed implementation may be implemented by a single processor (such as the processor 502), the speed and efficiency may be improved by more than one processor.
[0150] The memory 504 in the apparatus 500 may be a read only memory (ROM) device or a random access memory (RAM) device in one implementation. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 accessed by the processor 502 via a bus 512. The memory 504 may further include an operating system 508 and an application program 510, and the application program 510 includes at least one program that causes the processor 502 to execute the methods described herein. For example, the application program 510 may include application programs 1 to N, which further include a video decoding application program that executes the methods described herein. The apparatus 500 may further include other memory in the form of an auxiliary memory 514. For example, the auxiliary memory 514 may be a memory card used with a mobile computing device. Since video communication sessions may include a large amount of information, they may be stored in whole or in part in the auxiliary memory 514 and loaded into the memory 504 as needed for processing.
[0151] Device 500 may also include one or more output devices, such as display 518. In one example, display 518 may be a touch-sensitive display that combines a display with touch-sensitive elements operable to sense touch input. Display 518 may be coupled to processor 502 via bus 512. In addition to or as an alternative to display 518, other output devices may be provided that allow a user to program or otherwise use device 500. When the output device is or includes a display, the display may be implemented in a variety of ways, including by a liquid crystal display (LCD), a cathode ray tube (CRT) display, a plasma display, or a light emitting diode (LED) display, such as an organic LED (OLED) display.
[0152] Device 500 may also include or communicate with an image sensing device 520, such as a camera or any other existing or later developed image sensing device capable of sensing images such as images of a user operating device 500. Image sensing device 520 may be positioned to face the user operating device 500. In one example, the position and optical axis of image sensing device 520 may be such that the field of view includes an area immediately adjacent to display 518 from which display 518 can be seen.
[0153] Device 500 may also include or communicate with a sound sensing device 522, such as a microphone or any other existing or later developed sound sensing device capable of sensing sounds in the vicinity of device 500. Sound sensing device 522 may be positioned to face the user operating device 500 and may be used to receive sounds made by the user while operating device 500, such as speech or other utterances.
[0154] Although Figure 5The processor 502 and the memory 504 of the apparatus 500 are depicted as integrated into a single unit, but other configurations may be used. The operations of the processor 502 may be distributed across multiple machines, each having one or more processors, which may be directly coupled or coupled via a local area network or other network. The memory 504 may be distributed across multiple machines, such as network-based memory or memory in multiple machines that perform the operations of the apparatus 500. Although described herein as a single bus, the bus 512 of the apparatus 500 may consist of multiple buses. Additionally, the secondary memory 514 may be directly coupled to other components of the apparatus 500 or may be accessible via a network, and may include a single integrated unit (such as a memory card) or multiple units (such as multiple memory cards). Thus, the apparatus 500 may be implemented in a variety of configurations.
[0155] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If the functionality is implemented in software, the functionality may be stored or transmitted as one or more instructions or code in a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may include a tangible medium such as a computer-readable storage medium, corresponding data storage medium, etc., or a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this manner, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium, such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0156] Video compression techniques such as motion compensation, intra prediction, and loop filters have proven to be effective and are thus applied in various video coding standards such as H.264 / AVC and H.265 / HEVC. For example, intra prediction may be performed on an I-frame or I-slice when no reference image is available, or when the current block or image is not coded using inter prediction. The reference samples for intra prediction are typically obtained from previously coded (or reconstructed) neighboring blocks in the same image. For example, both H.264 / AVC and H.265 / HEVC use the boundary samples of neighboring blocks as references for intra prediction. Multiple different intra prediction modes are used to cover different texture or structural features. In each mode, a different method for obtaining the prediction signal is used. For example, as Figure 6 shown, H.265 / HEVC supports a total of 35 intra prediction modes.
[0157] Description of the Intra Prediction Algorithm in H.265 / HEVC
[0158] For intra prediction, the decoded boundary samples of adjacent blocks are used as references. The encoder selects the best luma intra prediction mode for each block from 35 options: 33 directional prediction modes, 1 DC mode, and 1 planar mode. Figure 6 The mapping relationship between the intra prediction direction and the intra prediction mode number is specified in. It should be noted that in the latest video coding technology, such as in VVC (Versatile Video Coding), 65 or more intra prediction modes have been developed, which can capture any edge direction presented in natural videos.
[0159] As Figure 7 shown, the block "CUR" is the current block to be predicted, and the grayscale samples along the boundaries of adjacent constructed blocks (left and above the current block) are used as reference samples. The prediction signal can be obtained by mapping the reference samples according to a specific method indicated by the intra prediction mode.
[0160] Refer to Figure 7 , where W is the width of the current block and H is the height of the current block. W1 is the number of top reference samples. H1 is the number of remaining reference samples. Generally, W1 > W and H1 > H, that is, the top reference samples also include the upper-right reference (adjacent) samples, and the left reference samples also include the lower-left reference (adjacent) samples. For example, H1 = 2×H, W1 = 2×W, or H1 = H + W, W1 = W + H.
[0161] The reference samples are not always available. For example, as Figure 8 shown, after the availability check process, there are W2 samples available at the top or above the current block, and H2 samples available on the left side of the current block. There are W3 samples unavailable at the top and H3 samples unavailable on the left.
[0162] Before obtaining the prediction signal, these unavailable samples need to be replaced (or filled) with the available samples. For example, by scanning the reference samples in a clockwise direction and replacing the unavailable samples with the values of the latest available samples. If the lower part of the left reference sample is unavailable, it is replaced with the value of the nearest available reference sample.
[0163] Reference sample availability check
[0164] The reference sample availability check refers to checking whether the reference samples are available. For example, if the reference samples have been reconstructed, then the reference samples are available.
[0165] In the existing methods, the availability check process is usually completed by checking the luma samples, that is, even for chrominance component blocks, the reference sample availability is obtained by checking the availability of the corresponding luma samples.
[0166] In the method proposed by the present invention, for a block, the reference sample availability check process is performed by checking the samples of its own components. The components here can be the Y component, the Cb component, or the Cr component. According to the provided method, for a block, the reference sample availability check process is performed by checking the samples of its corresponding components. The components here can be the luminance component or the chrominance component, and the chrominance component can include both the Cb component and the Cr component. The Cb component and the Cr component here are called chrominance components, and the Cb component and the Cr component are not distinguished during the availability detection process.
[0167] The present invention proposes a set of methods, focusing on the following two aspects.
[0168] According to the first aspect of the present invention, for a block, the reference sample availability check process is completed by checking the samples of its own components. The components here can be the Y component, the Cb component, or the Cr component.
[0169] For a Y block, the availability of the reference sample of the Y block is obtained by checking whether the reference sample is available (for example, by checking whether the reference sample has been reconstructed).
[0170] For a Cb block, the availability of the reference sample of the Cb block is obtained by checking whether the reference sample is available (for example, by checking whether the reference sample has been reconstructed).
[0171] For a Cr block, the availability of the reference sample of the Cr block is obtained by checking whether the reference sample is available (for example, by checking whether the reference sample has been reconstructed).
[0172] According to the second aspect of the present invention, for a block, the reference sample availability check process is performed by checking the samples of its corresponding components. Here, the components can be the luminance component or the chrominance component. The Y component is called the luminance component. Both the Cb component and the Cr component are called chrominance components. Here, the Cb component and the Cr component are not distinguished during the availability check process.
[0173] For a luminance block, the availability of the reference sample of the luminance block is obtained by checking whether the reference sample is available (for example, by checking whether the reference sample has been reconstructed).
[0174] For a chrominance block, the availability of the reference sample of the chrominance block is obtained by checking whether the reference sample is available (for example, by checking whether the reference sample has been reconstructed).
[0175] In the existing solution, for a block, regardless of which component the block belongs to, the reference sample availability check process is completed by checking the luminance samples. In contrast, in the method proposed by the present invention, for blocks belonging to different components, the reference sample availability check process is performed by checking different component samples.
[0176] It should be noted here that the method proposed in the present invention is used to obtain the availability of reference samples for blocks decoded in the intra prediction mode. This method can be executed by the intra prediction modules 254 or 354 respectively as shown in Figure 2 and Figure 3 . Therefore, this method is applicable to both the decoding end and the encoding end. The process of checking the availability of reference samples for blocks in the encoder and decoder is the same.
[0177] Figure 9 FIG. is a simplified flowchart of a method for obtaining a reconstructed signal provided by an exemplary embodiment of the present invention. Referring to Figure 9 , for a block decoded in the intra prediction mode, in order to obtain a reconstructed block (or signal, or sample), the method includes first obtaining a prediction (or prediction signal, or sample) of the block (901). Thereafter, the method includes obtaining the residual of the block (or residual signal) (902). Then, the method includes obtaining a reconstructed block by adding the residual to the prediction of the current block (903).
[0178] Figure 10 FIG. is a simplified flowchart of obtaining a prediction (or prediction signal) provided by an exemplary embodiment of the present invention. Referring to Figure 10 , in order to obtain a prediction signal, the method first includes obtaining the intra prediction mode of the current block, such as a planar mode or a DC mode, etc. (1001). Thereafter, the method includes obtaining the availability of reference samples of the components of the current block (1002). In one embodiment, the component includes a Y component, a Cb component, or a Cr component. In another embodiment, the component includes a luminance component or a chrominance component. The method can obtain the availability of reference samples by checking the availability of Y samples in adjacent Y blocks, where the adjacent Y blocks include reference samples; or by checking the availability of Cb samples in adjacent Cb blocks, where the adjacent Cb blocks include reference samples; or by checking the availability of Cr samples in adjacent Cr blocks, where the adjacent Cr blocks include reference samples.
[0179] When the method determines that there are unavailable reference samples (yes in 1003), the method includes replacing or filling the unavailable reference samples with available reference samples (1004). Thereafter, the method includes obtaining a prediction of the block according to the intra prediction mode and the replaced reference samples (1005). In 1003, when the method determines that there are no unavailable reference samples (no in 1003), step 1005 of the method is executed, and this step includes obtaining a prediction of the current block according to the intra prediction mode and the available reference samples. In 1006, the method includes reconstructing the current block according to the prediction.
[0180] It should be noted here that the embodiments of the present invention described below relate to the process of obtaining a prediction signal and improving the process of checking the availability of reference samples.
[0181] Embodiment 1
[0182] In this embodiment, the availability check process is performed by checking the samples of its own components.
[0183] Reference Figure 10 , for a block decoded in an intra prediction mode, in order to obtain a prediction signal, the method includes:
[0184] Step 1: Obtain the intra prediction mode (1001)
[0185] The intra prediction mode is obtained by parsing the syntax related to the intra prediction mode in the bitstream. For example, for a Y block, it is necessary to parse intra_luma_mpm_flag, intra_luma_mpm_idx or intra_luma_mpm_remainder. For a Cb / Cr block, it is necessary to parse the intra_chroma_pred_mode signal. Thereafter, the parsed syntax can be used to obtain the intra prediction mode of the current block.
[0186] Step 2: Perform the reference sample availability check process (1002)
[0187] This step includes checking the availability of the reference samples of the current block.
[0188] In one embodiment, when the block is a Y block, the reference sample availability is obtained by checking the reference samples of the Y component. Alternatively, the reference sample availability is obtained by checking adjacent Y samples (for example, by checking the boundary samples of adjacent Y blocks).
[0189] In one embodiment, when the block is a Cb block, the reference sample availability is obtained by checking the reference samples of the Cb component. Alternatively, the reference sample availability is obtained by checking adjacent Cb samples (for example, by checking the boundary samples of adjacent Cb blocks).
[0190] In one embodiment, when the block is a Cr block, the reference sample availability is obtained by checking the reference samples of the Cr component. Alternatively, the reference sample availability is obtained by checking adjacent Cr samples (for example, by checking the boundary samples of adjacent Cr blocks).
[0191] Step 3: Determine whether there are unavailable reference samples (1003)
[0192] In 1003, the method includes determining whether there are unavailable reference samples. When the method determines that there are unavailable reference samples, step 4 (1004) of the method is executed; otherwise, step 5 (1005) of the method is executed.
[0193] Step 4: Reference sample replacement (1004)
[0194] Reference sample replacement means obtaining the sample value of an unavailable reference sample with the sample value of an available reference sample. In one embodiment, the reference samples are scanned in a clockwise direction and the latest available sample value is used to replace the unavailable sample. If the lower part of the left reference sample is unavailable, it is replaced with the value of the nearest available reference sample.
[0195] Step 5: Obtain the prediction signal (1005)
[0196] After obtaining the reference samples and the intra-prediction mode, the prediction signal can be obtained by mapping the reference samples to the current block, and the mapping method is represented by the intra-prediction mode.
[0197] After obtaining the prediction signal of the current block, the reconstructed signal of the current block can be obtained by adding the residual signal to the obtained prediction signal ( Figure 9 of 903).
[0198] It should be noted here that after reconstructing the current block, the samples in the block area can be marked as available, as Figure 11 shown. That is to say, when the current block is a Y block, all Y samples at the positions covered by the block area are available; when the current block is a Cb block, all Cb samples at the positions covered by the block area are available; when the current block is a Cr block, all Cr samples at the positions covered by the block area are available.
[0199] An example of the specifications of these sections is as follows:
[0200] Reference sample availability annotation process
[0201] The inputs to this process are:
[0202] – Sample position (xTbCmp, yTbCmp), specifying the upper-left sample of the current transform block relative to the upper-left sample of the current image;
[0203] – Variable refIdx, specifying the intra-prediction reference clue index,
[0204] – Variable refW, specifying the reference sample width,
[0205] – Variable refH, specifying the reference sample height,
[0206] – The variable cIdx specifies the color component of the current block.
[0207] For in-sample prediction, the output of the process is the reference sample refUnfilt[x][y], where x = -1 - refIdx, y = -1 - refIdx..refH - 1 and x = -refIdx..refW - 1, y = -1 - refIdx.
[0208] refW + refH + 1+(2 * refIdx) The adjacent sample refUnfilt[x][y] is the sample constructed before intra-loop filtering, where x = -1 - refIdx, y = -1 - refIdx..refH - 1 and x = -refIdx..refW - 1, y = -1 - refIdx, and is obtained as follows:
[0209] – The adjacent position (xNbCmp, yNbCmp) is represented by the following parameters:
[0210] (xNbCmp, yNbCmp) = (xTbCmp + x, yTbCmp + y) (310)
[0211] – Call the adjacent block availability acquisition process specified in Clause 6.4.4, set the current sample position (xCurr, yCurr) to (xTbCmp, yTbCmp), the adjacent sample position to (xNbCmp, yNbCmp), set checkPredModeY to false, use cIdx as the input, and assign the output to availableN.
[0212] – Each sample refUnfilt[x][y] is obtained as follows:
[0213] – If availableN is false, mark the sample refUnfilt[x][y] as "not available for intra-frame prediction".
[0214] – Otherwise, mark the sample refUnfilt[x][y] as "available for intra-frame prediction", and assign the sample at the position (xNbCmp, yNbCmp) to refUnfilt[x][y].
[0215] Adjacent block availability acquisition process
[0216] The inputs of this process are:
[0217] – The sample position (xCurr, yCurr) of the top-left sample of the current block relative to the top-left sample of the current image,
[0218] – The sample positions (xNbCmp, yNbCmp) covered by the neighboring block relative to the top - left luminance sample of the current image.
[0219] – The variable checkPredModeY specifies whether the availability depends on the prediction mode.
[0220] – The variable cIdx specifies the color component of the current block.
[0221] The output of the process is the availability of the neighboring block covering the position (xNbCmp, yNbCmp), denoted as availableN.
[0222] The current luminance position (xTbY, yTbY) and the neighboring luminance position (xNbY, yNbY) are obtained as follows:
[0223] (xTbY,yTbY) = (cIdx == 0)? (xCurr,yCurr) :
[0224] (xCurr * SubWidthC,yCurr * SubHeightC)
[0225] (xNbY,yNbY) = (cIdx == 0)? (xNbCmp,yNbCmp) :
[0226] (xNbCmp * SubWidthC,yNbCmp * SubHeightC)
[0227] The availability availableN of the neighboring block is obtained as follows:
[0228] – If one or more of the following conditions are true, then set availableN to false.
[0229] – xNbCmp is less than 0.
[0230] – yNbCmp is less than 0.
[0231] – xNbY is greater than or equal to pic_width_in_luma_samples.
[0232] – yNbY is greater than or equal to pic_height_in_luma_samples.
[0233] – IsAvailable[cIdx][xNbCmp][yNbCmp] is false.
[0234] – The neighboring block is included in a different strip from the current block.
[0235] – The neighboring block is included in a different tile from the current block.
[0236] – The entropy_coding_sync_enabled_flag is equal to 1 and (xNbY >> CtbLog2SizeY) is greater than or equal to (xTbY >>
[0237] CtbLog2SizeY) + 1.
[0238] – Otherwise, set availableN to true.
[0239] Set availableN to false when the following conditions are all satisfied:
[0240] – checkPredModeY is true.
[0241] – availableN is true.
[0242] – CuPredMode[0][xNbY][yNbY] is not equal to CuPredMode[0][xTbY][yTbY].
[0243] IsAvailable[cIdx][x][y] is used to store the availability information of the samples at the sample position (x, y) for each component cIdx.
[0244] For cIdx = 0, 0 <= x <= pic_width_in_luma_samples, 0 <= y <= pic_height_in_luma_samples;
[0245] cIdx = 1 or 2. 0 <= x <= pic_width_in_chroma_samples, 0 <= y <= pic_height_in_chroma_samples;
[0246] The inputs to this process are:
[0247] – The position (xCurr, yCurr) specifies the position of the top - left sample of the current block relative to the top - left sample of the current image component,
[0248] – The variables nCurrSw and nCurrSh respectively specify the width and height of the current block.
[0249] – The variable cIdx specifies the color component of the current block,
[0250] – The (nCurrSw) x (nCurrSh) array predSamples specifies the predicted samples of the current block.
[0251] – The (nCurrSw) x (nCurrSh) array resSamples specifies the residual samples for the current block.
[0252] The output of this process is the reconstructed image sample array recSamples.
[0253] Depending on the value of the color component cIdx, the following assignments are made:
[0254] – If cIdx = 0, recSamples corresponds to the reconstructed image sample array S L 。
[0255] – Otherwise, if cIdx equals 1, set tuCbfChroma to be equal to tu_cbf_cb[xCurr][yCurr], and recSamples
[0256] corresponds to the reconstructed chroma sample array S Cb 。
[0257] – Otherwise (cIdx equals 2), set tuCbfChroma to be equal to tu_cbf_cr[xCurr][yCurr], and recSamples corresponds to the reconstructed chroma sample array S Cr 。
[0258] Depending on the value of pic_lmcs_enabled_flag, the following applies:
[0259] – If pic_lmcs_enabled_flag equals 0, the (nCurrSw) x (nCurrSh) block of the reconstructed samples recSamples at position (xCurr, yCurr) is obtained as follows, where i = 0..nCurrSw - 1, j = 0..nCurrSh - 1:
[0260] recSamples[xCurr + i][yCurr + j] = Clip1(predSamples[i][j] + resSamples[i][j]) (1195)
[0261] – Otherwise (pic_lmcs_enabled_flag equals 1), the following applies:
[0262] – When cIdx equals 0, the following applies:
[0263] – Invoke the image reconstruction of the mapping process for luminance samples specified in Clause 8.7.5.2, where the luminance position (xCurr,
[0264] It takes the current chrominance position (xCurr, yCurr), block width nCurrSw and height nCurrSh, predicted luminance sample array predSamples, and residual luminance sample array resSamples as inputs, and the output is the reconstructed luminance sample array recSamples.
[0265] – Otherwise (cIdx is greater than 0), call the image reconstruction of the luminance-related chrominance residual scaling process for chrominance samples specified in Clause 8.7.5.3, where the chrominance position (xCurr, yCurr), transform block width nCurrSw and height nCurrSh, decoding block flag of the current chrominance transform block tuCbfChroma, predicted chrominance sample array predSamples, and residual chrominance sample array resSamples are used as inputs, and the output is the reconstructed chrominance sample array recSamples.
[0266] Perform the following assignments, where i = 0..nCurrSw - 1 and j = 0..nCurrSh - 1:
[0267] xVb = (xCurr + i) % ((cIdx == 0)? IbcBufWidthY : IbcBufWidthC)(1196)
[0268] yVb = (yCurr + j) % ((cIdx == 0)? CtbSizeY : (CtbSizeY / subHeightC))(1197)
[0269] IbcVirBuf[cIdx][xVb][yVb] = recSamples[xCurr + i][yCurr + j](1198)
[0270] IsAvailable[cIdx][xCurr + i][yCurr + j] = TRUE(1199)
[0271] Another example of the specifications of these sections is shown below:
[0272] Reference sample availability annotation process
[0273] The inputs to this process are:
[0274] – Sample position (xTbCmp, yTbCmp), specifying the upper-left sample of the current transform block relative to the upper-left sample of the current image;
[0275] – Variable refIdx, specifying the intra-prediction reference cue index,
[0276] – Variable refW, which specifies the reference sample width,
[0277] – Variable refH, which specifies the reference sample height,
[0278] – Variable cIdx, which specifies the color component of the current block.
[0279] For intra-sample prediction, the output of the process is the reference sample refUnfilt[x][y], where x = -1 - refIdx, y = -1 - refIdx..refH - 1 and x = -refIdx..refW - 1, y = -1 - refIdx.
[0280] The refW + refH + 1+(2 * refIdx) adjacent samples refUnfilt[x][y] are the samples constructed before in-loop filtering, where x = -1 - refIdx, y = -1 - refIdx..refH - 1 and x = -refIdx..refW - 1, y = -1 - refIdx, and are obtained as follows:
[0281] – The adjacent positions (xNbCmp, yNbCmp) are specified by the following parameters:
[0282] (xNbCmp, yNbCmp) = (xTbCmp + x, yTbCmp + y) (310)
[0283] – Call the adjacent block availability obtaining process specified in Clause 6.4.4, set the current sample position (xCurr, yCurr) to (xTbCmp, yTbCmp), the adjacent sample position to (xNbCmp, yNbCmp), set checkPredModeY to false, use cIdx as the input, and assign the output to availableN.
[0284] – Each sample refUnfilt[x][y] is obtained as follows:
[0285] – If availableN is false, mark the sample refUnfilt[x][y] as "not available for intra prediction".
[0286] – Otherwise, mark the sample refUnfilt[x][y] as "available for intra prediction", and assign the sample at the position (xNbCmp, yNbCmp) to refUnfilt[x][y].
[0287] Adjacent block availability obtaining process
[0288] The inputs to this process are:
[0289] – The sample position (xCurr, yCurr) of the top - left sample of the current block relative to the top - left sample of the current image,
[0290] – The sample position (xNbCmp, yNbCmp) covered by the adjacent block relative to the top - left luma sample of the current image,
[0291] – The variable checkPredModeY specifies whether the availability depends on the prediction mode.
[0292] – The variable cIdx specifies the color component of the current block.
[0293] The output of the process is the availability of the adjacent block covering the position (xNbCmp, yNbCmp), denoted as availableN.
[0294] The current luma position (xTbY, yTbY) and the adjacent luma position (xNbY, yNbY) are obtained as follows:
[0295] (xTbY,yTbY) = (cIdx == 0)? (xCurr,yCurr):
[0296] (xNbY,yNbY) = (cIdx == 0)? (xNbCmp,yNbCmp):(XXX)
[0297] (xNbCmp*SubWidthC,yNbCmp*SubHeightC)
[0298] The availability availableN of the adjacent block is obtained as follows:
[0299] – If one or more of the following conditions are true, set availableN to false.
[0300] – xNbCmp is less than 0.
[0301] – yNbCmp is less than 0.
[0302] – xNbY is greater than or equal to pic_width_in_luma_samples.
[0303] – yNbY is greater than or equal to pic_height_in_luma_samples.
[0304] – IsAvailable[cIdx][xNbCmp][yNbCmp] is false.
[0305] – The adjacent block is included in a different strip from the current block.
[0306] – The adjacent block is included in a different partition from the current block.
[0307] – entropy_coding_sync_enabled_flag is equal to 1 and (xNbY >> CtbLog2SizeY) is greater than or equal to (xTbY >>
[0308] CtbLog2SizeY) + 1.
[0309] – Otherwise, set availableN to true.
[0310] Set availableN to false when the following conditions are all met:
[0311] – checkPredModeY is true.
[0312] – availableN is true.
[0313] – CuPredMode[0][xNbY][yNbY] is not equal to CuPredMode[0][xTbY][yTbY].
[0314] The inputs to the process are:
[0315] – The position (xCurr, yCurr) specifies the position of the top - left sample of the current block relative to the top - left sample of the current image component,
[0316] – The variables nCurrSw and nCurrSh respectively specify the width and height of the current block.
[0317] – The variable cIdx specifies the color component of the current block,
[0318] – The (nCurrSw) x (nCurrSh) array predSamples specifies the predicted samples of the current block.
[0319] – The (nCurrSw) x (nCurrSh) array resSamples specifies the residual samples of the current block.
[0320] The output of the process is the reconstructed image sample array recSamples.
[0321] According to the value of the color component cIdx, the following assignments are made:
[0322] – If cIdx = 0, recSamples corresponds to the reconstructed image sample array S L .
[0323] – Otherwise, if cIdx is equal to 1, set tuCbfChroma to be equal to tu_cbf_cb[xCurr][yCurr], and recSamples corresponds to the reconstructed chroma sample array S Cb .
[0324] – Otherwise (cIdx is equal to 2), set tuCbfChroma to be equal to tu_cbf_cr[xCurr][yCurr], and recSamples corresponds to the reconstructed chroma sample array S Cr .
[0325] Depending on the value of pic_lmcs_enabled_flag, the following cases apply:
[0326] – If pic_lmcs_enabled_flag is equal to 0, the (nCurrSw)x(nCurrSh) block of the reconstructed samples recSamples at position (xCurr, yCurr) is obtained as follows, where i = 0..nCurrSw-1, j = 0..nCurrSh-1:
[0327] recSamples[ xCurr + i ][ yCurr + j ] = Clip1(predSamples[ i ][ j ] +resSamples[ i ][ j ]) (1195)
[0328] – Otherwise (pic_lmcs_enabled_flag is equal to 1), the following cases apply:
[0329] – When cIdx is equal to 0, the following cases apply:
[0330] – Invoke the image reconstruction of the mapping process for the luma samples specified in Clause 8.7.5.2, where the luma position (xCurr,
[0331] yCurr), block width nCurrSw and height nCurrSh, predicted luma sample array predSamples, residual luma sample array resSamples are used as inputs, and the output is the reconstructed luma sample array recSamples.
[0332] – Otherwise (if cIdx is greater than 0), call the image reconstruction of the luminance-related chrominance residual scaling process for the chrominance samples specified in Clause 8.7.5.3, where the chrominance position (xCurr, yCurr), the transform block width nCurrSw and height nCurrSh, the decoding block flag of the current chrominance transform block tuCbfChroma, the predicted chrominance sample array predSamples, and the residual chrominance sample array resSamples are used as inputs, and the output is the reconstructed chrominance sample array recSamples.
[0333] Perform the following assignments, where i = 0..nCurrSw - 1 and j = 0..nCurrSh - 1:
[0334] xVb = (xCurr + i) % ((cIdx == 0)? IbcBufWidthY : IbcBufWidthC)(1196)
[0335] yVb = (yCurr + j) % ((cIdx == 0)? CtbSizeY : (CtbSizeY / subHeightC))(1197)
[0336] IbcVirBuf[cIdx][xVb][yVb] = recSamples[xCurr + i][yCurr + j](1198)
[0337] IsAvailable[cIdx][(xCurr + i)*((cIdx == 0)? 1 : SubWidthC)][(yCurr + j)*((cIdx == 0)? 1 : SubHeightC)] = TRUE(1199)
[0339] Figure 12 Schematic diagram showing that samples in the current block in the unit N×N provided for the exemplary embodiment of the present invention are marked as available. In one embodiment, refer to Figure 12, after reconstructing the current block, the samples in the unit of size N×N are marked as available, for example, N = 4 or N = 2. Alternatively, for the Y component block, N = 4, and for the Cb / Cr component block, N = 2. Any sample in an "available" unit is regarded as "available". That is, if a unit is marked as "available", any sample within the unit area can be marked as "available". That is to say, when the current unit is a Y unit, all Y samples at the positions covered by the unit area are available; when the current unit is a Cb unit, all Cb samples at the positions covered by the unit area are available; when the current block is a Cr unit, all Cr samples at the positions covered by the unit area are available.
[0340] In one embodiment, to check whether a reference sample is available, the process first obtains the position or index of the unit to which the sample belongs. When it is determined that the unit is available, the reference sample is regarded as available or marked as available.
[0341] Figure 13 Schematic diagram showing that the samples at the right boundary and bottom boundary of the current block provided by the exemplary embodiment of the present invention are marked as available. In one embodiment, refer to Figure 13 , after reconstructing the current block, only the right boundary samples and bottom boundary samples (which will be used as reference samples for other blocks) are marked as available.
[0342] Figure 14 Schematic diagram showing that the samples at the right boundary and bottom boundary of the current block in unit N provided by the exemplary embodiment of the present invention are marked as available. In one embodiment, refer to Figure 14 , after reconstructing the current block, only the right boundary samples and bottom boundary samples in the unit of size N (which will be used as reference samples for other blocks) are marked as available, for example, N = 4 or N = 2. Alternatively, for the Y component block, N = 4; for the Cb / Cr component block, N = 2. Any sample in an "available" unit is regarded as "available". That is, if a unit is marked as "available", any sample within the unit area can be marked as "available". That is to say, when the current unit is a Y unit, all Y samples at the positions covered by the unit area are available; when the current unit is a Cb unit, all Cb samples at the positions covered by the unit area are available; when the current block is a Cr unit, all Cr samples at the positions covered by the unit area are available.
[0343] To check whether a reference sample is available, in one embodiment, the process first obtains the position or index of the unit to which the sample belongs. When it is determined that the unit is available, the reference sample is marked as available.
[0344] Figure 15Schematic diagram provided for an exemplary embodiment of the present invention, where samples at the right boundary and bottom boundary of the current block in a unit N×N are marked as available. In one embodiment, referring to Figure 15 , after reconstructing the current block, only the right boundary samples and bottom boundary samples (which will serve as reference samples for other blocks) in a unit of size N×N are marked as available, for example, N = 4 or N = 2. Alternatively, for Y-component blocks, N = 4; for Cb / Cr-component blocks, N = 2. Any sample in an "available" unit is considered "available". That is, when a unit is marked as "available", any sample within the unit area can be marked as "available". That is to say, when the current unit is a Y unit, all Y samples at the positions covered by the unit area are available; when the current unit is a Cb unit, all Cb samples at the positions covered by the unit area are available; when the current block is a Cr unit, all Cr samples at the positions covered by the unit area are available.
[0345] To check whether a reference sample is available, in one embodiment, the process or method first obtains the position or index of the unit to which the sample belongs. When the process or method determines that the unit is "available", the process or method marks the reference sample as available.
[0346] It should be noted here that according to the embodiments of the present invention, the "available" information of samples or units for each component (a total of 3 components, Y component, Cb component, and Cr component) is stored in the memory.
[0347] It should be noted here that even if the block is decoded into an inter prediction mode, after reconstructing the block, the "available" marking method can also be applied.
[0348] Embodiment 2
[0349] In this embodiment, the availability check process is performed by checking the samples of the corresponding component.
[0350] The difference between Embodiment 2 and Embodiment 1 is that in the availability detection process, the Cb component and the Cr component are not distinguished. For a block, the reference sample availability check process is completed by checking the samples of its corresponding component, as Figure 16 shown. Here, the component can be a luminance component or a chrominance component. The Y component is called the luminance component. Both the Cb component and the Cr component are called chrominance components.
[0351] The difference between Embodiment 2 and Embodiment 1 is only in Step 2, and other steps are the same as those in Embodiment 1. Step 2 of Embodiment 2 is described in detail below.
[0352] Step 2: Perform the reference sample availability check process, including checking the availability of the reference samples of the current block ( Figure 10 1002 in
[0353] In one embodiment, when the block is a luminance block, the availability of reference samples is obtained by checking the reference luminance samples.
[0354] In one embodiment, when the block is a chrominance block, the availability of reference samples is obtained by checking the reference chrominance samples.
[0355] Reference Figure 10 , after checking the availability of reference samples (1002), the method includes determining whether there are unavailable reference samples (1003). When the method determines that there are unavailable reference samples (yes in 1003), step 1004 of the method is executed; otherwise, step 1005 of the method is executed.
[0356] It should be noted here that the marking method discussed in Embodiment 1 can be directly used in Embodiment 2. The only difference lies in the chrominance component. Only after reconstructing the Cb component block and the Cr component block can the chrominance samples in the block area be marked as "available".
[0357] For example, for a luminance block, when the block is reconstructed, the samples in the block area can be marked as available. For a chrominance block, after reconstructing the Cb and Cr blocks, the samples in the block area can be marked as "available". That is, when the current block is a luminance block, all luminance samples at the positions covered by the block area are marked as available. When the current block is a chrominance block, after reconstructing the Cb and Cr blocks, all chrominance samples at the positions covered by the block area are marked as available.
[0358] Other marking methods in Embodiment 1 can also be used similarly.
[0359] It should be noted here that in the embodiments of the present invention, the "availability" information of samples or units of each component (a total of 2 components, the luminance component and the chrominance component) is stored in the memory. That is, the Cb component and the Cr component will share the same "availability" or "available" information.
[0360] It should be noted here that even if the block is decoded into an inter prediction mode, after reconstructing the block, the "available" marking method can still be used.
[0361] The applications of the encoding method and the decoding method shown in the above embodiments and the systems using these methods will be described below.
[0362] Figure 17A block diagram of a content supply system 3100 for implementing a content distribution service. The content supply system 3100 includes a capture device 3102, a terminal device 3106, and optionally includes a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the above-mentioned communication channel 13. The communication link 3104 includes but is not limited to WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination thereof, etc.
[0363] The capture device 3102 generates data and may encode the data by an encoding method as shown in the above embodiments. Alternatively, the capture device 3102 may distribute the data to a streaming server (not shown in the figure), and the server encodes the data and sends the encoded data to the terminal device 3106. The capture device 3102 includes but is not limited to a camera, a smartphone or a tablet computer, a computer or a laptop, a video conferencing system, a PDA, a vehicle-mounted device, or any combination thereof, etc. For example, the capture device 3102 may include the source device 12 as described above. When the data includes video, the video encoder 20 included in the capture device 3102 may actually perform video encoding processing. When the data includes audio (i.e., voice), the audio encoder included in the capture device 3102 may actually perform audio encoding processing. For some actual scenarios, the capture device 3102 distributes by multiplexing the encoded video and audio data together. For other practical scenarios, such as in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and the encoded video data to the terminal device 3106 separately.
[0364] In the content providing system 3100, the terminal device 310 receives and reproduces the encoded data. The terminal device 3106 can be a device with data receiving and restoring capabilities, such as a smart phone or a tablet computer 3108, a computer or a laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof, or a device capable of decoding the above-mentioned encoded data. For example, the terminal device 3106 can include the target device 14 as described above. When the encoded data includes video, the video decoder 30 included in the terminal device preferentially performs video decoding. When the encoded data includes audio, the audio decoder included in the terminal device preferentially performs audio decoding processing.
[0365] For a terminal device with a display, such as a smart phone or a tablet computer 3108, a computer or a laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can feed the decoded data to the display of the terminal device. For a terminal device without a display (such as an STB 3116, a video conferencing system 3118, or a video surveillance system 3120), an external display 3126 is connected to receive and display the decoded data.
[0366] When performing encoding or decoding, each device in the system can use an image encoding device or an image decoding device as shown in the above embodiments.
[0367] Figure 18A diagram of the structure of an example of the terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol proceeding unit 3202 analyzes the transport protocol of the stream. The protocols include but are not limited to Real Time Streaming Protocol (RTSP), HyperText Transfer Protocol (HTTP), HTTP Live streaming protocol (HLS), MPEG-DASH, Real-time Transport protocol (RTP), RealTime Messaging Protocol (RTMP), or any combination thereof, etc.
[0368] After the protocol processing unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can divide the multiplexed data into encoded audio data and encoded video data. As described above, for some actual scenarios, such as in a video conferencing system, the encoded audio data and the encoded video data are not demultiplexed. In this case, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.
[0369] Through the demultiplexing process, a video elementary stream (ES), an audio ES, and optional subtitles are generated. The video decoder 3206 includes the video decoder 30 as described in the above embodiments, decodes the video ES by the decoding method as shown in the above embodiments to generate video frames, and feeds the data to the synchronization unit 3212. The audio decoder 3208 decodes the audio ES to generate audio frames and feeds the data to the synchronization unit 3212. Alternatively, before feeding the video frames to the synchronization unit 3212, the video frames can be stored in a buffer ( Figure 18 not shown). Similarly, before feeding the audio frames to the synchronization unit 3212, the audio frames can be stored in a buffer ( Figure 18 not shown).
[0370] The synchronization unit 3212 synchronizes the video frames and the audio frames and provides the video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of the video and audio information. Information can be decoded in the syntax using timestamps regarding the representation of the decoded audio and video data and timestamps regarding the delivery of the data stream itself.
[0371] If the stream includes subtitles, the subtitle decoder 3210 decodes the subtitles, synchronizes the subtitles with the video frames and audio frames, and provides the video / audio / subtitles to the video / audio / subtitle display 3216.
[0372] The present invention is not limited to the above system, and the image encoding device or image decoding device in the above embodiments can be incorporated into other systems, such as an automotive system.
[0373] According to the present invention, the described methods and processes can be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, the method or process can be executed by instructions or program code stored in a computer-readable medium and executed by a hardware processing unit.
[0374] In the proposed method, the available information of each sample is stored in each component so that the available information can be provided more accurately during the intra prediction process.
[0375] By way of example and not limitation, a computer-readable storage medium can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection can be termed a computer-readable medium. For example, if the instructions are transmitted from a website, server, or other remote source via coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, microwave, etc.), then the medium's definition includes coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave, etc.). However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transitory tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, disks generally reproduce data magnetically, while discs reproduce data optically with a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0376] Instructions or program code can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described herein can be provided in dedicated hardware and / or software modules for encoding and decoding, or incorporated in a combined codec. Further, these techniques can be fully implemented in one or more circuits or logic elements.
[0377] The techniques of the present invention can be implemented in a variety of devices or apparatuses, including wireless handheld telephones, integrated circuits (ICs), or a group of ICs (such as a chipset). The present invention describes various components, modules, or units to emphasize the functional aspects of the devices for performing the disclosed techniques, but these components, modules, or units do not necessarily require implementation by different hardware units. Instead, as described above, the various units can be combined in a codec hardware unit, or provided by a collection of interoperating hardware units including one or more processors as described above in conjunction with appropriate software and / or firmware.
[0378] Although several embodiments have been provided in the present invention, it should be understood that the disclosed systems and methods can be implemented in many other specific forms without departing from the spirit or scope of the present invention. These examples are considered illustrative rather than restrictive and are not intended to be limited to the details given herein. For example, various elements or components can be combined or integrated in another system, or certain features can be omitted or not implemented.
[0379] Moreover, without departing from the scope of the present invention, techniques, systems, subsystems, and methods described and shown as separate or discrete in various embodiments can be combined or integrated with other systems, modules, techniques, or methods. Other items shown or discussed as being coupled, directly coupled, or communicating with each other can be indirectly coupled or communicating through some interface, device, or intermediate component by electrical, mechanical, or other means. Other examples of alterations, substitutions, and changes can be determined by those skilled in the art and can be made without departing from the spirit and scope disclosed herein.
Claims
1. An intra-picture prediction apparatus, characterized in that, Comprising: A memory including instructions; One or more processors communicating with the memory, wherein the one or more processors execute the instructions to: Obtain the availability of adjacent luminance blocks, the availability of the adjacent luminance blocks indicating whether reference samples of the luminance blocks are available; Obtain the availability of adjacent chrominance blocks, the availability of the adjacent chrominance blocks indicating whether reference samples of the chrominance blocks are available; The luminance block and the chrominance block belong to the same coding tree unit (CTU), wherein the luminance block is obtained by one partitioning method and the chrominance block is obtained by another partitioning method; Obtain the prediction of the luminance block according to the availability of the adjacent luminance blocks, and obtain the prediction of the chrominance block according to the availability of the adjacent chrominance blocks.
2. The device according to claim 1, characterized in that The chrominance block includes a Cb component or a Cr component, or the chrominance block includes a Cb component and a Cr component.
3. The device according to claim 2, characterized in that, The one or more processors execute the instructions to: Obtain the availability of the adjacent chrominance blocks by checking the availability of Cb samples in adjacent Cb blocks, where the adjacent Cb blocks include the reference samples of the chrominance blocks; or Obtain the availability of the adjacent chrominance blocks by checking the availability of Cr samples in adjacent Cr blocks, where the adjacent Cr blocks include the reference samples of the chrominance blocks.
4. The device according to claim 2 or 3, characterized in that, All Cb samples in the reconstructed block are marked as available; or all Cr samples in the reconstructed block are marked as available.
5. The device according to any one of claims 1 to 3, characterized in that The block region includes a plurality of units, each unit having a unit region, and all samples in the unit region are marked as available.
6. The apparatus according to claim 5, wherein: When the current unit is a Cb unit, all Cb samples at the positions covered by the unit region are marked as available; or When the current unit is a Cr unit, all Cr samples at the positions covered by the unit region are marked as available.
7. The device according to any one of claims 1 to 6, characterized in that, Save the availability of the adjacent luminance blocks and save the availability of the adjacent chrominance blocks.
8. The device according to any one of claims 1 to 6, characterized in that, The chrominance block consists of a Cb component and a Cr component, and the Cb component and the Cr component share the same availability information.
9. A method for in-frame image prediction performed by an encoder or a decoder, characterized in that, Comprising: Obtain the availability of adjacent luminance blocks, the availability of the adjacent luminance blocks indicating whether reference samples of the luminance blocks are available; Obtain the availability of adjacent chrominance blocks, the availability of the adjacent chrominance blocks indicating whether reference samples of the chrominance blocks are available; The luminance block and the chrominance block belong to the same coding tree unit (CTU), wherein the luminance block is obtained by one partitioning method and the chrominance block is obtained by another partitioning method; Obtain the prediction of the luminance block according to the availability of the adjacent luminance blocks, and obtain the prediction of the chrominance block according to the availability of the adjacent chrominance blocks.
10. The method according to claim 9, wherein The chrominance block includes a Cb component or a Cr component, or the chrominance block includes a Cb component and a Cr component.
11. The method according to claim 10, wherein: Obtain the availability of the adjacent chrominance blocks by checking the availability of Cb samples in adjacent Cb blocks, where the adjacent Cb blocks include the reference samples of the chrominance blocks; or The availability of the adjacent chrominance block is obtained by checking the availability of Cr samples within an adjacent Cr block, where the adjacent Cr block includes reference samples of the chrominance block.
12. The method according to claim 10 or 11, characterized in that, All Cb samples in the reconstructed block are marked as available; or, all Cr samples in the reconstructed block are marked as available.
13. The method according to any one of claims 9 to 11, characterized in that The block region includes a plurality of units, each unit having a unit region.
14. The method according to claim 13, wherein: When the current unit is a Cb unit, all Cb samples at positions covered by the unit region are marked as available; or, When the current unit is a Cr unit, all Cr samples at positions covered by the unit region are marked as available.
15. The method according to claim 13, characterized in that, In a unit N×1 or 1×N, only the right boundary samples and the bottom boundary samples are marked as available.
16. The method according to claim 13, wherein In a unit N×N, only the right boundary samples and the bottom boundary samples are marked as available.
17. The method according to any one of claims 9 to 17, characterized in that, Save the availability of the adjacent luma block and save the availability of the adjacent chrominance block.
18. The method according to any one of claims 9 to 17, characterized in that The chrominance block consists of a Cb component and a Cr component, and the Cb component and the Cr component share the same availability information.
19. The method according to claim 13, wherein After reconstructing a unit in the chrominance block, all samples in the unit region of the unit are marked as available.
20. The method according to any one of claims 9 to 11, characterized in that, After reconstructing the chrominance block, the samples at the right boundary and the bottom boundary of the chrominance block are marked as available.
21. The method according to any one of claims 9 to 11, characterized in that After reconstructing the chrominance block, in a unit N×1 or 1×N, only the right boundary samples and the bottom boundary samples are marked as available.
22. The method according to any one of claims 9 to 11, characterized in that, After reconstructing the chrominance block, in a unit N×N, only the right boundary samples and the bottom boundary samples are marked as available.
23. A computer program product, characterized in that, Includes program code for performing the method according to any one of claims 9 to 22.
Citation Information
Patent Citations
Multi-type-tree framework for video coding
US20170208336A1