Intra prediction method and apparatus

By acquiring and replacing the availability of reference samples, an intra-frame prediction method is realized that improves the video compression ratio without affecting image quality, thus solving the problem of insufficient compression ratio in existing technologies.

CN120091135BActive Publication Date: 2026-05-05HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2019-11-21
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Given limited network resources and the demand for higher video quality, existing video compression and decompression technologies struggle to improve compression ratios without compromising image quality.

Method used

By obtaining the intra-prediction mode of the current block, the availability of reference samples is obtained, and unavailable reference samples are replaced with available reference samples to achieve more accurate intra-prediction.

Benefits of technology

It improves the accuracy of information in the intra-frame prediction process and enhances the video compression ratio without affecting image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091135B_ABST
    Figure CN120091135B_ABST
Patent Text Reader

Abstract

The application provides an intra prediction method and device for reconstructing unavailable reference samples. The method comprises: obtaining an intra prediction mode of a current block; obtaining availability of reference samples of a component of the current block; replacing unavailable reference samples with available reference samples; obtaining a prediction of the current block according to the intra prediction mode and the replaced reference samples; and reconstructing the current block according to the prediction. The component comprises a Cb component or a Cr component. Alternatively, the component comprises a chroma component. Since the method obtains availability of reference samples in each component, more accurate available information can be provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The original application has the application number 201980069241.2 and the original application date is November 21, 2019. The entire contents of the original application are incorporated herein by reference.

[0002] Related applications cross-application

[0003] This application claims priority to U.S. Provisional Patent Application No. 62 / 770,736, filed November 21, 2018, entitled “Intra Prediction Method and Device,” the disclosure of which is incorporated herein by reference. Technical Field

[0004] The embodiments of the present invention generally relate to the field of video decoding, and more particularly to the field of intra-frame prediction methods and devices. Background Technology

[0005] Even with shorter videos, a large amount of video data needs to be described, which can be challenging when the data needs to be transmitted over bandwidth-constrained communication networks or otherwise. Therefore, video data is typically compressed before transmission over modern telecommunications networks. Video size can also be an issue when storing video on storage devices due to potentially limited memory resources. Video compression devices typically use software and / or hardware at the source side to decode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device used to decode the video data. Given limited network resources and the growing demand for higher video quality, there is a need to improve compression and decompression techniques that can increase compression ratios with minimal impact on image quality. Summary of the Invention

[0006] Embodiments of the present invention provide intra-frame prediction apparatus and methods for encoding and decoding images. These embodiments should not be construed as limiting the examples set forth herein.

[0007] The method provided by a first aspect of the present invention includes: obtaining an intra-prediction mode of a current block; obtaining the availability of reference samples for components of the current block; replacing unavailable reference samples with available reference samples; obtaining a prediction of the current block based on the intra-prediction mode and the replaced reference samples; and reconstructing the current block based on the prediction. In one embodiment, the component includes a Y component, a Cb component, or a Cr component. In another embodiment, the component includes a luma component or a chroma component.

[0008] Because the first aspect of the invention obtains the availability of reference samples in each component, the available information can be provided more accurately.

[0009] According to one implementation of the first aspect, all Cb samples in the reconstruction block are marked as available; or, all Cr samples in the reconstruction block are marked as available.

[0010] Because this implementation of the first aspect stores available information for each sample in each component, it can provide more accurate available information during intra-frame prediction.

[0011] According to a second aspect of the invention, the decoder includes processing circuitry for performing the steps of the method described above.

[0012] According to a third aspect of the invention, the encoder includes processing circuitry for performing the steps of the method described above.

[0013] According to a fourth aspect of the present invention, a computer program product includes program code that, when executed by a processor, performs the above-described method.

[0014] According to a fifth aspect of the invention, a decoder for intra-frame prediction includes one or more processing units and a non-transient computer-readable storage medium coupled to the one or more processing units and storing program instructions, wherein the one or more processing units execute the program instructions to perform the method described above.

[0015] According to a sixth aspect of the invention, an encoder for intra-frame prediction includes one or more processing units and a non-transient computer-readable storage medium coupled to the one or more processing units and storing program instructions, the one or more processing units executing the program instructions to perform the method described above.

[0016] This invention also provides a decoding device and an encoding device for performing the above-described methods.

[0017] For clarity, any of the above embodiments may be combined with any of the other above embodiments to create new embodiments within the scope of the present invention.

[0018] These and other features will become clearer from the following detailed description taken in conjunction with the accompanying drawings and claims. Attached Figure Description

[0019] To gain a more complete understanding of the invention, reference is made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein the same reference numerals denote the same parts.

[0020] Figure 1A This is a block diagram of an example decoding system that can implement an embodiment of the present invention.

[0021] Figure 1B This is a block diagram of another example decoding system that can implement an embodiment of the present invention.

[0022] Figure 2 This is a block diagram of an example video encoder that can implement an embodiment of the present invention.

[0023] Figure 3 This is a block diagram illustrating an example of a video decoder that can implement an embodiment of the present invention.

[0024] Figure 4 A schematic diagram of a network device provided for an exemplary embodiment of the present invention.

[0025] Figure 5 Provided as an exemplary embodiment of the present invention, which can be used as Figure 1A A simplified block diagram of one or both of the source device and the target device.

[0026] Figure 6 This describes an intra-frame prediction algorithm for H.265 / HEVC that can be used in the description of this invention.

[0027] Figure 7 A schematic diagram of a reference sample provided for an exemplary embodiment of the present invention.

[0028] Figure 8 A schematic diagram of available and unavailable reference samples provided for exemplary embodiments of the present invention.

[0029] Figure 9 A simplified flowchart for obtaining the reconstructed signal is provided for an exemplary embodiment of the present invention.

[0030] Figure 10 A simplified flowchart for obtaining a prediction signal is provided for an exemplary embodiment of the present invention.

[0031] Figure 11 The sample in the current block provided for an exemplary embodiment of the present invention is marked as available.

[0032] Figure 12 This is a schematic diagram showing that samples in the current block of an N×N cell are marked as available, provided as an exemplary embodiment of the present invention.

[0033] Figure 13 Samples at the right and bottom boundaries of the current block provided for an exemplary embodiment of the present invention are marked as available schematic diagrams.

[0034] Figure 14 The samples at the right and bottom boundaries of the current block in cell N, provided for an exemplary embodiment of the present invention, are marked as available schematic diagrams.

[0035] Figure 15 The samples at the right and bottom boundaries of the current block in cell N×N provided as an exemplary embodiment of the present invention are marked as available.

[0036] Figure 16 A schematic diagram of a method for labeling samples of chromaticity components provided as an exemplary embodiment of the present invention.

[0037] Figure 17 A block diagram of an exemplary structure for a content delivery system 3100 that implements a content distribution service.

[0038] Figure 18 A block diagram showing the structure of an example terminal device. Detailed Implementation

[0039] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using various techniques, whether currently known or existing. The invention should not be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but can be modified within the full scope of the appended claims and their equivalents.

[0040] Figure 1A This is a schematic block diagram of an example decoding system 10 that can use bidirectional prediction techniques. Figure 1A As shown, the decoding system 10 includes a source device 12 that provides encoded video data, and a target device 14 that decodes the encoded video data. Specifically, the source device 12 can provide video data to the target device 14 via a computer-readable medium 16. The source device 12 and the target device 14 can include any of a variety of devices, including desktop computers, laptop computers, tablet computers, set-top boxes, mobile phones (e.g., smartphones, smart tablets), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source device 12 and the target device 14 can be used for wireless communication.

[0041] Target device 14 can receive encoded video data to be decoded via computer-readable medium 16. Computer-readable medium 16 can include any type of medium or device capable of moving encoded video data from source device 12 to target device 14. In one example, computer-readable medium 16 can include a communication medium enabling source device 12 to transmit encoded video data directly to target device 14 in real time. The encoded video data can be modulated according to a communication standard (e.g., a wireless communication protocol) and transmitted to target device 14. The communication medium can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, wide area network, or global network (e.g., the Internet). The communication medium can include a router, switch, base station, or any other device that may facilitate communication from source device 12 to target device 14.

[0042] In some examples, encoded data can be output from output interface 22 to a storage device. Similarly, encoded data can be accessed from a storage device via an input interface. The storage device can include any of a variety of distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, digital video disks (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, the storage device can correspond to a file server or other intermediate storage device capable of storing encoded video generated by source device 12. Target device 14 can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing and sending the encoded video data to target device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. Target device 14 can access the encoded video data via any standard data connection, including an internet connection. This can include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., digital subscriber line (DSL), cable modems, etc.), or combinations thereof, suitable for accessing encoded video data stored on a file server. Transmission of the encoded video data from the storage device can be streaming, downloading, or a combination thereof.

[0043] The technologies of this invention are not necessarily limited to wireless applications or setups. These technologies can be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission (e.g., dynamic adaptive streaming over HTTP (DASH)), digital video encoded into data storage media, decoding digital video stored in data storage media, or other applications. In some examples, the decoding system 10 can be used to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.

[0044] exist Figure 1AIn one example, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Target device 14 includes an input interface 28, a video decoder 30, and a display device 32. According to the invention, the video encoder 200 of source device 12 and / or the video decoder 300 of target device 14 can employ bidirectional prediction technology. In other examples, the source device and target device may include other components or devices. For example, source device 12 may receive video data from an external video source (e.g., an external camera). Similarly, target device 14 may be connected to an external display device, rather than including an integrated display device.

[0045] Figure 1A The decoding system 10 shown is merely an example. Bidirectional prediction techniques can be implemented by any digital video encoding and / or decoding device. Although the techniques of this invention are typically implemented by video decoding devices, these techniques can also be implemented by video encoders / decoders, commonly referred to as "codecs / decoders (CODECs)". Furthermore, the techniques of this invention can also be implemented by a video preprocessor. The video encoder and / or decoder can be a graphics processing unit (GPU) or a similar device.

[0046] Source device 12 and target device 14 are merely examples of such decoding devices, where source device 12 generates decoded video data and sends it to target device 14. In some examples, source device 12 and target device 14 can operate in a substantially symmetrical manner, such that each of source device 12 and target device 14 includes both a video coded component and a decoded component. Therefore, decoding system 10 can support one-way or two-way video transmission between video devices 12, 14, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0047] The video source 18 of source device 12 may include a video capture device (e.g., a video camera), a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. Alternatively, video source 18 may generate computer graphics-based data as source video, or a combination of real-time video, archived video, and computer-generated video.

[0048] In some cases, when the video source 18 is a video camera, the source device 12 and the target device 14 can form a camera phone or video phone. However, as described above, the techniques described in this invention are generally applicable to video decoding and can be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video can be encoded by the video encoder 20. The encoded video information can then be output from the output interface 22 to the computer-readable medium 16.

[0049] Computer-readable medium 16 may include transient media, such as wireless broadcasting or wired network transmissions, or storage media (i.e., non-transient storage media), such as hard disks, flash drives, optical discs, digital video discs, Blu-ray discs, or other computer-readable media. In some examples, a network server (not shown) may receive encoded video data from source device 12 and, for example, provide the encoded video data to target device 14 via network transmission. Similarly, a computing device in a media production facility (e.g., an optical disc stamping facility) may receive encoded video data from source device 12 and produce an optical disc including the encoded video data. Therefore, in various examples, computer-readable medium 16 can be understood to include one or more computer-readable media of various forms.

[0050] The input interface 28 of the target device 14 receives information from the computer-readable medium 16. The information on the computer-readable medium 16 may include syntax information defined by the video encoder 20, which is also used by the video decoder 30. This syntax information includes description blocks and other decoding units (e.g., features of groups of images (GOPs) and / or syntax elements processed). The display device 32 displays the decoded video data to the user and may include any of a variety of display devices, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light emitting diode (OLED) display, or other types of display devices.

[0051] The video encoder 200 and video decoder 300 can operate according to video decoding standards (such as the currently developing High Efficiency Video Coding (HEVC) standard) and can conform to the HEVC Test Model (HM). Alternatively, the video encoder 200 and video decoder 300 can operate according to other proprietary or industry standards, such as the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) H.264 standard, also known as the Motion Picture Expert Group (MPEG)-4, Part 10, Advanced Video Coding (AVC), H.265 / HEVC, or extensions of such standards. However, the technology of this invention is not limited to any particular decoding standard. Other examples of video decoding standards include MPEG-2 and ITU-T H.263. Although not explicitly stated... Figure 1A As shown, but in some respects, the video encoder 200 and video decoder 300 can be integrated with the audio encoder and decoder respectively, and can include appropriate multiplexer-demultiplexer (MUX-DEMUX) units or other hardware and software to encode audio and video in a common data stream or separate data streams. Where applicable, the MUX-DEMUX unit can conform to the ITU H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP).

[0052] The video encoder 200 and video decoder 300 can be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When these technologies are partially implemented in software, the device can store software instructions in a suitable non-transitory computer-readable medium and execute these instructions in hardware by one or more processors to perform the techniques of the present invention. Each of the video encoder 200 and video decoder 300 can be included in one or more encoders or decoders, wherein any encoder or decoder can be integrated as part of a combined encoder / decoder (codec) in the respective device. Devices such as the video encoder 200 and / or video decoder 300 can include integrated circuits, microprocessors, and / or wireless communication devices, such as cellular phones.

[0053] Figure 1B Provided for an exemplary embodiment, including Figure 2 encoder 200 and / or Figure 3 A schematic block diagram of an example video decoding system 40 with decoder 300. System 40 can implement the techniques of the present invention, such as fusion estimation in inter-frame prediction. In the illustrated implementation, video decoding system 40 may include imaging device 41, video encoder 20, video decoder 300 (and / or a video decoder implemented by logic circuitry 47 of processing unit 46), antenna 42, one or more processors 43, one or more memories 44, and / or display device 45.

[0054] As shown in the figure, the imaging device 41, antenna 42, processing unit 46, logic circuit 47, video encoder 20, video decoder 30, processor 43, memory 44, and / or display device 45 can communicate with each other. As discussed, although both video encoder 20 and video decoder 30 are shown, in various practical scenarios, the video decoding system 40 may include only video encoder 20 or only video decoder 30.

[0055] As shown in the figure, in some examples, the video decoding system 40 may include an antenna 42. For example, the antenna 42 may be used to transmit or receive encoded bitstreams of video data. Furthermore, in some examples, the video decoding system 40 may include a display device 45. The display device 45 may be used to present video data. As shown in the figure, in some examples, the logic circuit 47 may be implemented by a processing unit 46. The processing unit 46 may include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. The video decoding system 40 may also include an optional processor 43, which may similarly include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. In some examples, the logic circuit 54 may be implemented by hardware, video decoding-specific hardware, etc., and the processor 43 may implement general-purpose software, an operating system, etc. Additionally, the memory 44 may be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting example, memory 44 may be implemented using cache memory. In some examples, logic circuitry 47 may access memory 44 (e.g., for implementing an image buffer). In other examples, logic circuitry 47 and / or processing unit 46 may include memory (e.g., cache, etc.) for implementing image buffers, etc.

[0056] In some examples, the video encoder 200 implemented via logic circuitry may include an image buffer (e.g., implemented via processing unit 46 or memory 44) and a graphics processing unit (e.g., implemented via processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the video encoder 200 implemented via logic circuitry 47 to achieve a combination Figure 2 The various modules discussed and / or any other encoder systems or subsystems described herein. Logic circuits can be used to perform the various operations described herein.

[0057] The video decoder 300 can be implemented in a similar manner to that implemented via logic circuit 47 to achieve integration. Figure 3The various modules discussed in the decoder 300 and / or any other decoder system or subsystem described herein. In some examples, the video decoder 300, which can be implemented via logic circuitry, may include an image buffer (e.g., implemented via processing unit 46 or memory 44) and a graphics processing unit (e.g., implemented via processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the video decoder 300 implemented via logic circuitry 47 to achieve a combination Figure 3 The various modules discussed and / or any other decoder systems or subsystems described herein.

[0058] In some examples, the antenna 42 of the video decoding system 40 can be used to receive an encoded bitstream of video data. As discussed, the encoded bitstream may include data related to the encoding of the video frames discussed herein, indicators, index values, mode selection data, etc., such as data related to decoding segmentation (e.g., transform coefficients or quantization transform coefficients, optional indicators (as discussed), and / or data defining the decoding segmentation). The video decoding system 40 may also include a video decoder 300 coupled to the antenna 42, which is used to decode the encoded bitstream. The display device 45 is used to present the video frames.

[0059] Figure 2 This is a block diagram illustrating an example of a video encoder 200 that can implement the technology of this application. The video encoder 200 can perform intra-frame and inter-frame decoding on video blocks within a video slice. Intra-frame decoding reduces or eliminates spatial redundancy in a given video frame or image through spatial prediction. Inter-frame decoding reduces or eliminates temporal redundancy in adjacent frames or images of a video sequence through temporal prediction. Intra-frame mode (I-mode) can refer to any of several spatially based decoding modes. Inter-frame mode (e.g., one-way prediction (P-mode) or two-way prediction (B-mode)) can refer to any of several time-based decoding modes.

[0060] Figure 2 This is a schematic / conceptual block diagram of an exemplary video encoder 200 used to implement the technology of the present invention. Figure 2In the example, the video encoder 200 includes a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a buffer 216, a loop filtering unit 220, a decoded picture buffer (DPB) 230, a prediction processing unit 260, and an entropy coding unit 270. The prediction processing unit 260 may include inter-frame estimation 242, inter-frame prediction unit 244, intra-frame estimation unit 252, intra-frame prediction unit 254, and a mode selection unit 262. The inter-frame prediction unit 244 may also include a motion compensation unit (not shown). According to the hybrid video codec, Figure 2 The video encoder 200 shown can also be called a hybrid video encoder or a video encoder.

[0061] For example, the residual calculation unit 204, transform processing unit 206, quantization unit 208, prediction processing unit 260, and entropy coding unit 270 constitute the forward signal path of the encoder 200, while, for example, the inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, buffer 216, loop filter 220, decoder buffer (DPB) 230, and prediction processing unit 260 constitute the reverse signal path of the encoder, where the reverse signal path of the encoder corresponds to the decoder (see...). Figure 3 The signal path of the decoder 300.

[0062] For example, encoder 200 is used to receive image 201 or block 203 of image 201 through input terminal 202, where image 201 is, for example, an image sequence that constitutes a video or video sequence. Image block 203 may also be referred to as current image block or image block to be decoded, and image 201 is referred to as current image or image to be decoded (especially in video decoding, to distinguish the current image from other images, such as images previously encoded and / or decoded in the same video sequence (i.e., a video sequence that also includes the current image).

[0063] segmentation

[0064] Embodiments of encoder 200 may include a segmentation unit ( Figure 2 (Not shown in the image), the segmentation unit is used to divide image 201 into multiple blocks (e.g., blocks like block 203), typically into multiple non-overlapping blocks. The segmentation unit can be used to apply the same block size to all images in the video sequence and the corresponding raster with a defined block size, or to change the block size between images or subsets or groups of images and segment each image into the corresponding block.

[0065] In HEVC and other video decoding specifications, a set of coding tree units (CTUs) is generated to produce a coded representation of an image. Each CTU can include a coding tree block for luminance samples, two corresponding coding tree blocks for chrominance samples, and a syntax structure for decoding the samples of the coding tree block. In monochrome images or images with three independent color planes, a CTU can include a single coding tree block and a syntax structure for decoding the samples of the coding tree block. The coding tree block can be an N×N sample block. A CTU can also be called a tree block or the largest coding unit (LCU). HEVC's CTU can be broadly similar to macroblocks in other standards such as H.264 / AVC. However, a CTU is not necessarily limited to a specific size and can include one or more coding units (CUs). A strip can include an integer number of CTUs ordered consecutively in raster scan order.

[0066] In HEVC, the CTU is divided into CUs using a quadtree structure represented as a decoding tree to accommodate different local features. At the CU level, a decision is made whether to decode the image region via inter-frame (temporal) prediction or intra-frame (spatial) prediction. A CU may include a decoding block of luminance samples, two corresponding decoding blocks of chrominance samples, and a syntax structure for decoding the samples of the decoding block, where the image has luminance sample arrays, Cb sample arrays, and Cr sample arrays. In monochrome images or images with three independent color planes, a CU may include a single decoding block and a syntax structure for decoding the samples of the decoding block. The decoding block is an N×N sample block. In some examples, the size of the CU may be the same as the CTU. Each CU is decoded using a decoding mode, which may be, for example, an intra-frame decoding mode or an inter-frame decoding mode. Other decoding modes may also be used. Encoder 200 receives video data. Encoder 200 may encode each CTU in a strip of the video data image. As part of encoding the CTU, the prediction processing unit 260 or other processing units of the encoder 200 (including but not limited to) Figure 2 The encoder 200 shown can perform segmentation to divide the CTU's CTB into smaller blocks 203. The smaller blocks can be the decoder blocks of the CU.

[0067] The syntax data in the bitstream can also define the size of the CTU. A stripe consists of multiple consecutive CTUs arranged in decoding order. Video frames or images / pictures can be segmented into one or more stripes. As mentioned earlier, each tree block can be segmented into coding units (CUs) according to a quadtree. Typically, a quadtree data structure includes one node per CU, with the root node corresponding to a tree block (e.g., a CTU). If a CU is divided into four sub-CUs, the node corresponding to that CU includes four child nodes, each child node corresponding to a sub-CU. The multiple nodes of the quadtree structure include leaf nodes and non-leaf nodes. Leaf nodes have no child nodes in the tree structure (i.e., leaf nodes are not further subdivided). Non-leaf nodes include the root node of the tree structure. For each corresponding non-root node among the multiple nodes, the corresponding non-root node corresponds to a child CU of the CU, and the child CU of the CU corresponds to the parent node of the corresponding non-root node in the tree structure. Each corresponding non-leaf node has one or more child nodes in the tree structure.

[0068] Each node in a quadtree data structure can provide syntax data for its corresponding CU. For example, a node in a quadtree can include a partition flag indicating whether the CU corresponding to that node has been partitioned into a subCU. The syntax elements of a CU can be defined recursively and can depend on whether the CU has been partitioned into subCUs. If no further partitioning of a CU is performed, the CU is called a leaf CU. If further partitioning of the CU's blocks is performed, the CU is generally called a non-leaf CU. Each level of partitioning is a quadtree partition, dividing the CU into four subCUs. Black CUs are examples of leaf nodes (i.e., blocks that have not been further partitioned).

[0069] The role of a CU is similar to that of a macroblock in the H.264 standard, except that CUs do not distinguish between sizes. For example, a tree block can be divided into four child nodes (also called child CUs), and each child node can become a parent node and be divided into four more child nodes. The final undivided child nodes are called leaf nodes of the quadtree, including the decoding nodes, also called leaf CUs. The syntax data associated with the decoded bitstream can be defined as the maximum number of times the tree block can be divided, called the maximum CU depth, and the minimum size of the decoding node can also be defined. Correspondingly, the bitstream can also define the smallest coding unit (SCU). The term "block" refers to any of the CU, PU, ​​or TU in the HEVC context, or a similar data structure in other standard contexts (e.g., macroblocks and their subblocks in H.264 / AVC).

[0070] In HEVC, each CU can be further divided into one, two, or four PUs based on the PU partitioning type. Within a PU, the same prediction process is performed, and relevant information is sent to the decoder on a PU-by-PU basis. After obtaining the residual block through the prediction process, the CU can be partitioned into transform units (TUs) according to the PU partitioning type, based on other quadtree structures similar to the decoder tree used for the CU. A key feature of the HEVC structure is that it has multiple partitioning concepts such as CU, PU, ​​and TU. PUs can be partitioned into non-square shapes. Syntax data associated with the CU can also describe, for example, partitioning the CU into one or more PUs. The shape of a TU can be square or non-square (e.g., rectangular), and syntax data associated with the CU can describe, for example, partitioning the CU into one or more TUs according to a quadtree. The partitioning mode may differ depending on whether the CU is encoded using skip mode, direct mode, intra-prediction mode, or inter-prediction mode.

[0071] VVC (Versatile Video Coding) does not distinguish between PU and TU concepts and supports various CU segmentation shapes. The size of the CU corresponds to the size of the decoding node and can be square or non-square (e.g., rectangular). The size of the CU can range from 4×4 pixels (or 8×8 pixels) to the size of a tree block, with a maximum of 128×128 pixels or larger (e.g., 256×256 pixels).

[0072] After encoder 200 generates prediction blocks (e.g., luminance, Cb, and Cr prediction blocks) for the CU, encoder 200 can generate residual blocks for the CU. For example, encoder 100 can generate luminance residual blocks for the CU. Each sample in the luminance residual block of the CU represents the difference between a luminance sample in the predicted luminance block of the CU and a corresponding sample in the original luminance decoded block of the CU. Furthermore, encoder 200 can generate Cb residual blocks for the CU. Each sample in the Cb residual block of the CU can represent the difference between a Cb sample in the predicted Cb block of the CU and a corresponding sample in the original Cb decoded block of the CU. Encoder 200 can also generate Cr residual blocks for the CU. Each sample in the Cr residual block of the CU can represent the difference between a Cr sample in the predicted Cr block of the CU and a corresponding sample in the original Cr decoded block of the CU.

[0073] In some examples, encoder 200 does not perform a transformation on the transform block. In such examples, encoder 200 can process the residual sample values ​​in the same way as the transform coefficients. Therefore, in examples where encoder 200 does not perform a transformation, the following discussion about transform coefficients and coefficient blocks can be applied to the transform block of the residual samples.

[0074] After generating coefficient blocks (e.g., luminance coefficient blocks, Cb coefficient blocks, or Cr coefficient blocks), encoder 200 can quantize the coefficient blocks to minimize the amount of data used to represent them, thereby enabling further compression. Quantization typically refers to the process of simplifying a series of values ​​into a single value. After quantizing the coefficient blocks, encoder 200 can entropy encode the syntax elements representing the quantized transform coefficients. For example, encoder 200 can perform context-adaptive binary arithmetic coding or other entropy decoding techniques on the syntax elements representing the quantized transform coefficients.

[0075] Encoder 200 can output a bitstream of encoded image data 271, which includes a column of bits that form a representation of the decoded image and related data. Therefore, the bitstream includes an encoded representation of video data.

[0076] In J. An et al.'s "Block Partitioning Structure for Next-Generation Video Decoding" (International Telecommunication Union, COM16-C966, September 2015 (hereinafter referred to as "VCEG Proposal COM16-C966")), a quad-tree-binary-tree (QTBT) partitioning technique was proposed for future video decoding standards beyond HEVC. Simulation results show that the proposed QTBT structure is more efficient than the quad-tree structure used in HEVC. In HEVC, to reduce memory access for motion compensation, inter-frame prediction for small blocks is restricted; therefore, bidirectional prediction for 4×8 blocks and 8×4 blocks, as well as inter-frame prediction for 4×4 blocks, are not supported. These restrictions are removed in JEM's QTBT.

[0077] In QTBT, CUs can be squares or rectangles. For example, the decoding tree unit (CTU) is first segmented using a quadtree structure. The leaf nodes of the quadtree can be further segmented using a binary tree structure. There are two types of binary tree segmentation: symmetrical horizontal segmentation and symmetrical vertical segmentation. In each case, the node is segmented horizontally or vertically from the center downwards. The leaf nodes of the binary tree are called decoding units (CUs), and this segment is used for prediction and transform processing without further segmentation. That is, CUs, PUs, and TUs have the same block size in the QTBT decoding block structure. CUs can consist of code blocks (CBs) of different color components; for example, in the case of P and B stripes in a 4:2:0 chroma format, a CU includes one luma CB and two chroma CBs; or a CU can consist of CBs of a single component, for example, a CU includes only one luma CB or only two chroma CBs in the case of I stripes.

[0078] Define the following parameters for the QTBT segmentation scheme:

[0079] –CTU size: The size of the root node of the quadtree, the same concept as in HEVC.

[0080] –MinQTSize: The minimum allowed size of a quadtree leaf node.

[0081] –MaxBTSize: The maximum allowed size of the root node of a binary tree.

[0082] –MaxBTDPepth: The maximum allowed binary tree depth.

[0083] –MinBTSize: The minimum allowed size of a binary leaf node.

[0084] In one example of a QTBT segmentation structure, the CTU size is set to 128 × 128 luminance samples, with two corresponding blocks of 64 × 64 chrominance samples. The MinQTSize is set to 16 × 16, the MaxBTSize to 64 × 64, the MinBTSize (width and height) to 4 × 4, and the MaxBTDepth to 4. Quadtree segmentation is first applied to the CTU to generate quadtree leaf nodes. The size of the quadtree leaf nodes can range from 16 × 16 (i.e., MinQTSize) to 128 × 128 (i.e., the CTU size). When the size of a quadtree node equals MinQTSize, no further quadtree segmentation is considered. If a leaf quadtree node is 128 × 128, its size exceeds MaxBTSize (i.e., 64 × 64), so no further segmentation via a binary tree is performed. Otherwise, the leaf quadtree node can be further segmented via a binary tree. Therefore, the quadtree leaf node is also the root node of a binary tree with a depth of 0. When the binary tree depth reaches MaxBTDepth (i.e., 4), further partitioning is no longer considered. When the width of a binary tree node equals MinBTSize (i.e., 4), further horizontal partitioning is no longer considered. Similarly, when the height of a binary tree node equals MinBTSize, further vertical partitioning is no longer considered. Prediction and transformation processing is performed on the leaf nodes of the binary tree without further partitioning. In JEM, the maximum CTU size is 256×256 luminance samples. Further processing (e.g., by performing prediction and transformation processes) can be performed on the leaf nodes of the binary-tree (CU) without further partitioning.

[0085] Furthermore, in the QTBT scheme, luminance and chrominance have separate QTBT structures. Currently, for P and B stripes, the luminance and chrominance CTBs within a single CTU can share the same QTBT structure. However, for I stripes, the luminance CTB is divided into CUs using QTBT structures, and the chrominance CTB can be divided into chrominance CUs using other QTBT structures. That is, the CUs in I stripes consist of decoding blocks for the luminance component or decoding blocks for the two chrominance components, while the CUs in P or B stripes consist of decoding blocks for all three color components.

[0086] Encoder 200 applies a rate-distortion optimization (RDO) process to the QTBT structure to determine block segmentation.

[0087] Furthermore, U.S. Patent Application Publication No. 20170208336 proposes a block partitioning structure called multi-type-tree (MTT) to replace CU structures based on QT, BT, and / or QTBT. The MTT partitioning structure remains a recursive tree structure. In MTT, multiple different partitioning structures (e.g., three or more) are used. For example, according to MTT technology, at each depth of the tree structure, three or more different partitioning structures can be used for each corresponding non-leaf node of the tree structure. The depth of a node in the tree structure can refer to the length of the path from the node to the root of the tree structure (e.g., the number of partitions). A partitioning structure can generally refer to how many different blocks a block can be divided into. A partitioning structure can be a quadtree partitioning structure that divides a block into four blocks, a binary tree partitioning structure that divides a block into two blocks, or a ternary tree partitioning structure that divides a block into three blocks. Furthermore, a ternary tree partitioning structure can partition blocks without a center. There can be multiple different partitioning types. The segmentation type can also define how the block is segmented, including symmetrical or asymmetrical segmentation, uniform or non-uniform segmentation, and / or horizontal or vertical segmentation.

[0088] In MTT, at each depth of the tree structure, encoder 200 can be used to further partition the subtree using a specific partition type from one of three or more partition structures. For example, encoder 100 can be used to determine a specific partition type based on QT, BT, triple-tree (TT), and other partition structures. In one example, a QT partition structure may include a square quadtree or a rectangular quadtree partition type. Encoder 200 can use a square quadtree partition to partition a square block by dividing the block horizontally and vertically along its center into four equal-sized square blocks. Similarly, encoder 200 can use a rectangular quadtree partition to partition a rectangular (e.g., non-square) block by dividing the rectangular block horizontally and vertically along its center into four equal-sized rectangular blocks.

[0089] The BT partitioning structure can include at least one of horizontally symmetric binary trees, vertically symmetric binary trees, horizontally asymmetric binary trees, and vertically asymmetric binary trees. For horizontally symmetric binary tree partitioning, encoder 200 can be used to horizontally divide a block into two symmetric blocks of the same size along its center. For vertically symmetric binary tree partitioning, encoder 200 can be used to vertically divide a block into two symmetric blocks of the same size along its center. For horizontally asymmetric binary tree partitioning, encoder 200 can be used to horizontally divide a block into two blocks of different sizes. For example, one block may be 1 / 4 the size of its parent block, and the other block may be 3 / 4 the size of its parent block, similar to the PART_2N×nU or PART_2N×nD partitioning types. For vertically asymmetric binary tree partitioning, encoder 200 can be used to vertically divide a block into two blocks of different sizes. For example, one block may be 1 / 4 the size of its parent block, and the other block may be 3 / 4 the size of its parent block, similar to the PART_nL×2N or PART_nR×2N partitioning types. In other examples, asymmetric binary tree partitioning types can divide a parent block into parts of different sizes. For example, one child block could be 3 / 8 of the parent block, and another child block could be 5 / 8 of the parent block. Of course, such partitioning types can be vertical or horizontal.

[0090] Unlike QT or BT structures, the TT partitioning structure does not divide blocks along the center. The central region of a block remains within the same sub-block. Unlike QT, which produces four blocks, or a binary tree, which produces two blocks, partitioning according to the TT partitioning structure produces three blocks. Example partitioning types according to the TT partitioning structure include symmetrical partitioning types (horizontal and vertical) and asymmetrical partitioning types (horizontal and vertical). Furthermore, according to the TT partitioning structure, symmetrical partitioning types can be either uneven or non-uniform (uniform / uniform). According to the TT partitioning structure, asymmetrical partitioning types are non-uniform. In one example, the TT partitioning structure may include at least one of the following partitioning types: horizontally uniform symmetrical ternary tree partitioning type, vertically uniform symmetrical ternary tree partitioning type, horizontally non-uniform symmetrical ternary tree partitioning type, vertically non-uniform symmetrical ternary tree partitioning type, horizontally non-uniform asymmetrical ternary tree partitioning type, or vertically non-uniform asymmetrical ternary tree partitioning type.

[0091] Typically, a non-uniform symmetric ternary tree partition is a partition symmetrical about the center line of a block, but in which at least one of the resulting three blocks is different in size from the other two. In a preferred example, the two side blocks are each 1 / 4 the size of the block, while the center block is 1 / 2 the size of the block. A uniform symmetric ternary tree partition is a partition symmetrical about the center line of a block, and all resulting blocks are the same size. This partition is possible if the height or width of the block (depending on whether it is a vertical or horizontal partition) is a multiple of 3. A non-uniform asymmetric ternary tree partition is a partition asymmetrical about the center line of a block, in which at least one of the resulting blocks is different in size from the other two.

[0092] In examples where blocks (e.g., at subtree nodes) are partitioned into an asymmetric ternary tree segmentation type, encoder 200 and / or decoder 300 may apply a constraint that two of the three partitions are the same size. This constraint may correspond to a constraint that encoder 200 must adhere to when encoding video data. Furthermore, in some examples, encoder 200 and decoder 300 may apply a constraint such that, when partitioning according to the asymmetric ternary tree segmentation type, the sum of the areas of two partitions equals the area of ​​the remaining partition.

[0093] In some examples, encoder 200 can be used to select a segmentation type from all the aforementioned segmentation types for each of the QT, BT, and TT segmentation structures. In other examples, encoder 200 can be used to determine a segmentation type only from a subset of the aforementioned segmentation types. For example, a subset of the aforementioned segmentation types (or other segmentation types) can be used for a specific block size or a specific depth of a quadtree structure. The subset of supported segmentation types can be signaled in the bitstream for use by decoder 200, or it can be predefined so that encoder 200 and decoder 300 can determine the subset without any signal.

[0094] In other examples, the number of supported segmentation types can be fixed for all depths across all CTUs. That is, encoder 200 and decoder 300 can be pre-configured to use the same number of segmentation types for any depth of the CTU. In other examples, the number of supported segmentation types can vary and may depend on depth, stripe type, or other previously decoded information. In one example, at depth 0 or depth 1 of the tree structure, only the QT segmentation structure is used. At depths greater than 1, each of the QT, BT, and TT segmentation structures can be used.

[0095] In some examples, encoder 200 and / or decoder 300 can pre-configure restrictions on the supported segmentation types to avoid duplicate segmentation of a region of the video image or a region of the CTU. In one example, when dividing a block using an asymmetric segmentation type, encoder 200 and / or decoder 300 may not further subdivide the largest sub-block derived from the current block. For example, when dividing a square block according to an asymmetric segmentation type (similar to the PART_2N×nU segmentation type), the largest sub-block among all sub-blocks (similar to the largest sub-block segmentation type in PART_2N×nU) is the marked leaf node and cannot be further subdivided. However, smaller sub-blocks (similar to smaller sub-blocks in the PART_2N×nU segmentation type) can be further subdivided.

[0096] As another example, the supported segmentation types can be restricted to avoid repeated segmentation of specific regions. When dividing a block with an asymmetric segmentation type, it is not possible to further subdivide the largest sub-block obtained from the current block along the same direction. For example, when a square block is of an asymmetric segmentation type (similar to the PART_2N×nU segmentation type), the encoder 200 and / or decoder 300 may not subdivide the largest sub-block (similar to the largest sub-block of the PART_2N×nU segmentation type) among all sub-blocks along the horizontal direction.

[0097] As another example, the supported segmentation types can be restricted to facilitate further segmentation, where the encoder 200 and / or decoder 300 can segment the block neither horizontally nor vertically when the width / height of the block is not a power of 2 (e.g., when the width and height are not 2, 4, 8, 16, etc.).

[0098] The above example describes how encoder 200 can perform MTT segmentation. Decoder 300 can then perform the same MTT segmentation as encoder 200. In some examples, the segmentation method of encoder 200 for video data can be determined by applying the same set of predefined rules at decoder 300. However, in many cases, encoder 200 can determine the specific segmentation structure and type to use based on the rate-distortion criterion for the specific image of the video data being decoded. Therefore, in order for decoder 300 to determine the segmentation of a specific image, encoder 200 can indicate syntax elements in the encoded bitstream that represent how the image and its CTUs should be segmented. Decoder 200 can parse such syntax elements and segment the image and CTUs accordingly.

[0099] In one example, the prediction processing unit 260 of the video encoder 200 can be used to perform any combination of the segmentation techniques described above, particularly for motion estimation, which will be described in detail below.

[0100] Although block 203 is smaller than image 201, like image 201, block 203 is, or can be considered, a two-dimensional array or matrix of samples with intensity values ​​(sample values). In other words, image block 203 may include, for example, a sample array (e.g., a luminance array in the case of black and white image 201), three sample arrays (e.g., a luminance array and two chrominance arrays in the case of color image 201), or any other number and / or type of array, depending on the applied color format. The number of samples in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203.

[0101] like Figure 2 The encoder 200 shown is used to encode the image 201 block by block, for example, to perform encoding and prediction for each block 203.

[0102] Residual calculation

[0103] The residual calculation unit 204 is used to calculate the residual block 205 based on the image block 203 and the prediction block 265 (the prediction block 265 will be described in detail later). For example, the residual block 205 in the sample domain is obtained by subtracting the sample value of the prediction block 265 from the sample value of the image block 203 on a sample-by-sample (pixel-by-pixel) basis.

[0104] Transformation

[0105] The transform processing unit 206 is used to transform the sample values ​​of the residual block 205, such as by discrete cosine transform (DCT) or discrete sine transform (DST), to obtain transform coefficients 207 in the transform domain. The transform coefficients 207 can also be called transform residual coefficients and represent the residual block 205 in the transform domain.

[0106] Transform processing unit 206 can be used to apply integer approximations of DCT / DST, such as those specified for HEVC / H.265. Compared to orthogonal DCT transforms, such integer approximations are typically scaled by a certain factor. Other scaling factors are used as part of the transform process to maintain the norm of the residual block after forward and backward transforms. The scaling factor is usually selected based on certain constraints, such as the scaling factor being a power of 2 of the shift operation, the bit depth of the transform coefficients, and a trade-off between precision and implementation cost. For example, a specific scaling factor for the inverse transform can be specified on the decoder 300 side via inverse transform processing unit 212 (and a scaling factor can be specified for the corresponding inverse transform on the encoder 200 side via inverse transform processing unit 212), and a corresponding scaling factor for the forward transform can be specified on the encoder 200 side via transform processing unit 206.

[0107] Quantification

[0108] Quantization unit 208 is used to quantize the transform coefficients 207 (e.g., scalar quantization or vector quantization) to obtain quantized transform coefficients 209. Quantized transform coefficients 209 can also be referred to as quantized residual coefficients 209. The quantization process can reduce the bit depth associated with some or all of the transform coefficients 207. For example, n-bit transform coefficients can be rounded down to m-bit transform coefficients during quantization, where n is greater than m. The degree of quantization can be modified by adjusting the quantization parameter (QP). For example, for scalar quantization, different scaling can be used to achieve finer or coarser quantization. Smaller quantization step sizes result in finer quantization; larger quantization step sizes result in coarser quantization. A suitable quantization step size can be indicated by the quantization parameter (QP). For example, the quantization parameter can be an index applicable to a predefined set of suitable quantization step sizes. For example, a small quantization parameter corresponds to fine quantization (small quantization step size), while a large quantization parameter corresponds to coarse quantization (large quantization step size), and vice versa. Quantization operations may include division by the quantization step size, while corresponding dequantization or inverse dequantization operations performed by the dequantization unit 210, etc., may include multiplication by the quantization step size. According to some standards (e.g., HEVC), the quantization step size can be determined using quantization parameters in embodiments. Typically, the quantization step size can be calculated from the quantization parameters using a fixed-point approximation of the equations involving division. Other scaling factors can be introduced into quantization and dequantization to recover the norm of the residual block, as scaling is used in the fixed-point approximation of the equations for the quantization step size and quantization parameters, thus modifying the norm. In one exemplary implementation, scaling in the inverse transform and dequantization can be incorporated. Alternatively, a custom quantization table can be used and indicated from the encoder to the decoder in the bitstream, etc. Quantization is a lossy operation, and the loss increases with the quantization step size.

[0109] The dequantization unit 210 performs dequantization on the quantization coefficients performed by the quantization unit 208 to obtain the dequantization coefficients 211. For example, it performs a dequantization scheme based on or using the same quantization step size as the quantization unit 208, which is the same quantization scheme performed by the quantization unit 208. The dequantization coefficients 211 can also be called dequantization residual coefficients 211, which correspond to the transform coefficients 207. However, due to the loss caused by quantization, the dequantization coefficients 211 are usually not exactly the same as the transform coefficients.

[0110] The inverse transform processing unit 212 is used to perform an inverse transform on the transform performed by the transform processing unit 206, such as an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain an inverse transform block 213 in the sample domain. The inverse transform block 213 can also be called an inverse transform dequantization block 213 or an inverse transform residual block 213.

[0111] Reconstruction unit 214 (e.g., summer 214) is used to add inverse transform block 213 (i.e., reconstruction residual block 213) to prediction block 265 to obtain reconstruction block 215 in the sample domain, for example, by adding the sample values ​​of reconstruction residual block 213 to the sample values ​​of prediction block 265.

[0112] Optionally, for example, buffer unit 216 (or buffer 216) of line buffer 216 may be used to buffer or store reconstructed block 215 and corresponding sample values ​​for, for example, intra-frame prediction. In other embodiments, the encoder may use unfiltered reconstructed blocks and / or corresponding sample values ​​stored in buffer unit 216 for any type of estimation and / or prediction, such as intra-frame prediction.

[0113] An embodiment of encoder 200 may allow buffer unit 216 to be used not only to store reconstructed blocks 215 for intra-frame prediction 254, but also for loop filter unit 220. Figure 2 (not shown in the image), and / or make, for example, buffer unit 216 and decoded image buffer unit 230 constitute a buffer. Other embodiments may use filter blocks 221 and / or blocks or samples (blocks or samples in the image buffer 230) of the decoded image buffer 230. Figure 2 (Not shown in the image) serves as the input or basis for intra-frame prediction 254.

[0114] Loop filter unit 220 (or loop filter 220) is used to filter the reconstructed block 215 to obtain filtered block 221, for example, to smooth pixel transitions or improve video quality. Loop filter unit 220 represents one or more loop filters, such as deblocking filters, sample-adaptive offset (SAO) filters, and other filters, such as bilateral filters, adaptive loop filters (ALF), sharpening or smoothing filters, or cooperative filters. Although in Figure 2 The middle loop filter unit 220 is shown as an in-loop filter, but in other configurations, the loop filter unit 220 can be implemented as a rear loop filter. The filter block 221 can also be referred to as the filter reconstruction block 221. The decoded image buffer 230 can store the reconstructed decoded block after the loop filter unit 220 performs the filtering operation on the reconstructed decoded block.

[0115] For example, an embodiment of encoder 200 (correspondingly, loop filter unit 220) can directly output loop filter parameters (e.g., sampled adaptive offset information), or output loop filter parameters after entropy encoding by entropy encoding unit 270 or any other entropy decoding unit, so that, for example, decoder 300 can receive the same loop filter parameters and apply the same loop filter parameters to decoding.

[0116] The decoded picture buffer (DPB) 230 can be a reference picture memory that stores reference picture data used by the video encoder 20 to encode video data. The DPB 230 can be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magneto resistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The DPB 230 and buffer 216 can be provided by the same memory device or by different memory devices. In some examples, the decoded picture buffer (DPB) 230 is used to store filter block 221. The decoded picture buffer 230 can also be used to store other previously filtered blocks (e.g., previously reconstructed and filtered block 221) of the same current image or different images (e.g., previously reconstructed images), and can provide the complete previously reconstructed (i.e., decoded) image (along with the corresponding reference block and samples) and / or the partially reconstructed current image (along with the corresponding reference block and samples) for, for example, inter-frame prediction. In some examples, if reconstructed block 215 is reconstructed but no in-loop filtering is performed, the decoded picture buffer (DPB) 230 is used to store reconstructed block 215.

[0117] The prediction processing unit 260, also known as the block prediction processing unit 260, is used to: receive or acquire block 203 (e.g., current block 203 of current image 201) and reconstruct image data, such as reference samples of the same (or current) image from buffer 216 and / or reference image data 231 of one or more previously decoded images from decoded image buffer 230; and to process such data for prediction, i.e., to provide prediction block 265, which may be inter-frame prediction block 245 or intra-frame prediction block 255.

[0118] The mode selection unit 262 can be used to select a prediction mode (e.g., intra-frame or inter-frame prediction mode) and / or select the corresponding prediction block 245 or 255 as prediction block 265 for the calculation of residual block 205 and reconstruction of reconstruction block 215.

[0119] Embodiments of the mode selection unit 262 can be used to select a prediction mode (e.g., from those prediction modes supported by the prediction processing unit 260) that provides an optimal match, or in other words, provides the minimum residual (minimum residual is more favorable for compression for transmission or storage), or provides the minimum indication overhead (minimum indication overhead is more favorable for compression for transmission or storage), or considers or balances both. The mode selection unit 262 can be used to determine a prediction mode based on rate distortion optimization (RDO), i.e., to determine a prediction mode that provides minimum rate distortion optimization, or whose associated rate distortion at least satisfies the prediction mode selection criteria.

[0120] The prediction processing (e.g., by prediction processing unit 260) and mode selection (e.g., by mode selection unit 262) performed by example encoder 200 are explained in more detail below.

[0121] As described above, encoder 200 is used to determine or select the best or optimal prediction mode from a set of (predetermined) prediction modes. This set of prediction modes may include, for example, intra-frame prediction modes and / or inter-frame prediction modes.

[0122] This set of intra-prediction modes may include 35 different intra-prediction modes, such as non-directional modes like DC (or mean) mode and planar mode, or directional modes defined in H.265, or may include 67 different intra-prediction modes, such as non-directional modes like DC (or mean) mode and planar mode, or directional modes defined in H.266 which is under development.

[0123] The set (or possible) inter-frame prediction modes depend on the available reference image (i.e., a previously at least partially decoded image stored in the DBP 230) and other inter-frame prediction parameters, such as whether the entire reference image or only a portion of the reference image (e.g., the search window region around the current block) is used to search for the best-matching reference block, and / or whether pixel interpolation is applied, such as half-pixel and / or quarter-pixel interpolation, or no pixel interpolation is applied.

[0124] In addition to the prediction modes mentioned above, skip mode and / or direct mode can also be applied.

[0125] The prediction processing unit 260 can also be used, for example, to divide block 203 into smaller block portions or sub-blocks by iteratively using quad-tree partitioning (QT), binary partitioning (BT), or triple-tree partitioning (TT), or any combination thereof, and to perform a prediction, for example, on each of the block portions or sub-blocks, wherein mode selection includes selecting the tree structure of the partitioned block 203 and the prediction mode applied to each block portion or sub-block.

[0126] The inter-frame prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit. Figure 2 (Not shown in the image). The motion estimation unit is used to receive or acquire image block 203 (current image block 203 of current image 201) and decoded image 331, or at least one or more previously reconstructed blocks (e.g., reconstructed blocks of one or more other / different previously decoded images 331) for motion estimation. For example, a video sequence may include the current image and previously decoded image 331, or, in other words, the current image and previously decoded image 331 may be part of an image sequence constituting the video sequence or may constitute the sequence. For example, encoder 200 may be used to select a reference block from multiple reference blocks of the same or different images of multiple other images and provide the offset (spatial offset) between the position (x-coordinate, y-coordinate) of the reference image (or reference image index...) and / or the position of the reference block (x-coordinate, y-coordinate) and the position of the current block, as a motion estimation unit ( Figure 2 The inter-frame prediction parameters (not shown in the image) are used. This offset is also called the motion vector (MV). Fusion is an important motion estimation tool used in HEVC, and it is also used in VVC. To perform fusion estimation, a fusion candidate list is first constructed, where each candidate includes all motion data, including information on whether one or two reference image lists are used, as well as the reference index and motion vector of each list. The fusion candidate list is constructed based on the following candidates: 1. Up to four spatial fusion candidates, which are obtained from five spatially adjacent (i.e., neighboring) blocks; 2. One temporal fusion candidate, which is obtained from two temporally juxtaposed blocks; 3. Other fusion candidates, including combined bidirectional prediction candidates and zero motion vector candidates.

[0127] Intra-prediction unit 254 is further configured to determine intra-prediction block 255 based on intra-prediction parameters (e.g., the selected intra-prediction mode). In any case, after selecting an intra-prediction mode for the block, intra-prediction unit 254 is also configured to provide intra-prediction parameters, i.e., information indicating the intra-prediction mode selected for the block, to entropy coding unit 270. In one example, intra-prediction unit 254 may be used to perform any combination of intra-prediction techniques described below.

[0128] Entropy coding unit 270 is used to apply entropy coding algorithms or schemes (e.g., variable length coding (VLC), context adaptive VLC (CAVLC), arithmetic decoding, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) decoding, or other entropy coding methods or techniques) individually or jointly (or not at all) to the quantization residual coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, and / or loop filter parameters to obtain encoded image data 21 that can be output by output terminal 272, for example, in the form of an encoded bitstream 21. The encoded bitstream 21 can be sent to video decoder 30 or archived for subsequent transmission or retrieval by video decoder 30. Entropy coding unit 270 can also be used to entropy code other syntax elements of the current video strip being decoded.

[0129] Other structural variations of the video encoder 200 can be used to encode video streams. For example, for certain blocks or frames, the non-transform-based encoder 200 can directly quantize the residual signal without the transform processing unit 206. In another implementation, the encoder 200 can combine the quantization unit 208 and the inverse quantization unit 210 into a single unit.

[0130] Figure 3 An exemplary video decoder 300 is provided for implementing the technology of the present invention. The video decoder 300 is used to receive, for example, encoded image data (e.g., an encoded bitstream) 271 encoded by encoder 200 to obtain a decoded image 331. During the decoding process, the video decoder 300 receives video data from the video encoder 200, such as an encoded video bitstream representing image blocks and associated syntax elements of an encoded video stripe.

[0131] exist Figure 3 In one example, the decoder 300 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., a summer 314), a buffer 316, a loop filter 320, a decoded image buffer 330, and a prediction processing unit 360. The prediction processing unit 360 may include an inter-frame prediction unit 344, an intra-frame prediction unit 354, and a mode selection unit 362. In some examples, the video decoder 300 can perform... Figure 2 The encoding channels described in the video encoder 200 are largely inverse decoding channels.

[0132] Entropy decoding unit 304 is used to perform entropy decoding on encoded image data 271 to obtain, for example, quantization coefficients 309 and / or decoded decoding parameters. Figure 3 (Not shown in the image), such as any or all of the (decoded) inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 is also used to forward the inter-frame prediction parameters, intra-frame prediction parameters, and / or other syntax elements to the prediction processing unit 360. The video decoder 300 can receive syntax elements at the video stripe level and / or the video block level.

[0133] The inverse quantization unit 310 can have the same function as the inverse quantization unit 110, the inverse transform processing unit 312 can have the same function as the inverse transform processing unit 112, the reconstruction unit 314 can have the same function as the reconstruction unit 114, the buffer 316 can have the same function as the buffer 116, the loop filter 320 can have the same function as the loop filter 120, and the decoded image buffer 330 can have the same function as the decoded image buffer 130.

[0134] The prediction processing unit 360 may include an inter-frame prediction unit 344 and an intra-frame prediction unit 354, wherein the inter-frame prediction unit 344 is functionally similar to the inter-frame prediction unit 144, and the intra-frame prediction unit 354 is functionally similar to the intra-frame prediction unit 154. The prediction processing unit 360 is typically used to perform block prediction and / or obtain prediction blocks 365 from the coded data 21, and to receive or obtain (explicitly or implicitly) prediction-related parameters and / or information on the selected prediction mode from the entropy decoding unit 304.

[0135] When the video stripe is decoded as an intra-frame decoded (I) stripe, the intra-frame prediction unit 354 of the prediction processing unit 360 is used to generate a prediction block 365 for the image block of the current video stripe based on the indicated intra-frame prediction mode and data from previously decoded blocks from the current frame or image. When the video frame is decoded as an inter-frame decoded (i.e., B or P) stripe, the inter-frame prediction unit 344 (e.g., motion compensation unit) of the prediction processing unit 360 is used to generate a prediction block 365 for the video block of the current video stripe based on motion vectors and other syntax elements received from the entropy decoding unit 304. For inter-frame prediction, the prediction block can be generated from a reference image in a list of reference images. The video decoder 300 can construct the reference frame list, list 0 and list 1, based on the reference images stored in the DPB 330 using a default construction technique.

[0136] The prediction processing unit 360 is used to determine the prediction information of video blocks in the current video slice by parsing motion vectors and other syntax elements, and to use the prediction information to generate prediction blocks for the current video block being decoded. For example, the prediction processing unit 360 uses some received syntax elements to determine the prediction mode (e.g., intra-frame or inter-frame prediction) used to decode the video blocks in the video slice, the inter-frame prediction slice type (e.g., B-slice, P-slice, or GPB-slice), the construction information of one or more of the reference image list of the slice, the motion vector of each inter-frame coded video block of the slice, the inter-frame prediction state of each inter-frame decoded video block of the slice, and other information to decode the video blocks in the current video slice.

[0137] The dequantization unit 310 can be used to dequantize (i.e., dequantize) the quantization transform coefficients provided in the bitstream and decoded by the entropy decoding unit 304. The dequantization process may include: using the quantization parameters of each video block in the video strip calculated by the video encoder 100 to determine the degree of quantization and dequantization to be applied.

[0138] The inverse transform processing unit 312 is used to apply an inverse transform to the transform coefficients, such as an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, in order to generate a residual block in the pixel domain.

[0139] Reconstruction unit 314 (e.g., summer 314) is used to add inverse transform block 313 (i.e., reconstruction residual block 313) to prediction block 365 to obtain reconstruction block 315 in the sample domain, for example, by adding the sample values ​​of reconstruction residual block 313 to the sample values ​​of prediction block 365.

[0140] Loop filtering unit 320 (in or after the decoding loop) is used to filter the reconstructed block 315 to obtain filtered block 321, for example, to smooth pixel abrupt changes or improve video quality. In one example, loop filtering unit 320 can be used to perform any combination of the filtering techniques described below. Loop filter unit 320 represents one or more loop filters, such as deblocking filters, sample-adaptive offset (SAO) filters, and other filters such as bilateral filters, adaptive loop filters (ALF), sharpening or smoothing filters, or cooperative filters. Although in Figure 3 The middle loop filter unit 320 is shown as an in-loop filter, but in other configurations, the loop filter unit 320 can be implemented as a rear loop filter.

[0141] Then, the decoded video block 321 in a given frame or image is stored in the decoded image buffer 330, which stores a reference image for subsequent motion compensation.

[0142] The decoder 300 is used, for example, to output a decoded image 311 via output terminal 312 for presentation to a user or for viewing by a user.

[0143] Other variations of the video decoder 300 can be used to decode compressed bitstreams. For example, the decoder 300 can generate an output video stream without the loop filtering unit 320. For example, for certain blocks or frames, the non-transform-based decoder 300 can directly quantize the residual signal without the inverse transform processing unit 312. In another implementation, the video decoder 300 can combine the inverse quantization unit 310 and the inverse transform processing unit 312 into a single unit.

[0144] Figure 4 This is a schematic diagram of a network device 400 (e.g., a decoding device) provided according to one embodiment of the present invention. The network device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the network device 400 may be a decoder (e.g., Figure 1A The video decoder 30 or encoder (e.g., in the video decoder 30) or encoder (e.g. Figure 1A (Video encoder 20 in the image). In one embodiment, the network device 400 may be as described above. Figure 1A Video decoder 30 or Figure 1A One or more components of the video encoder 20 in the video encoder 20.

[0145] Network device 400 includes: an input port 410 and a receiver unit (Rx) 420 for receiving data; a processor, logic unit, or central processing unit (CPU) 430 for processing data; a transmitter unit (Tx) 440 and an output port 450 for transmitting data; and a memory 460 for storing data. Network device 400 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the input port 410, receiver unit 420, transmitter unit 440, and output port 450 for the input or output of optical or electrical signals.

[0146] Processor 430 is implemented in both hardware and software. Processor 430 can be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. Processor 430 communicates with ingress port 410, receiver unit 420, transmitter unit 440, egress port 450, and memory 460. Processor 430 includes a decoding module 470. Decoding module 470 implements the embodiments disclosed above. For example, decoding module 470 implements, processes, prepares, or provides various decoding operations. Therefore, having decoding module 470 can significantly improve the functionality of network device 400 and affect the transitions of network device 400 to different states. Alternatively, decoding module 470 can be implemented as instructions stored in memory 460 and executed by processor 430.

[0147] Memory 460 includes one or more disks, tape drives, and solid-state drives, and can be used as an overflow data storage device to store programs when such programs are selected for execution, as well as instructions and data read during program execution. Memory 460 can be volatile and / or non-volatile memory, and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).

[0148] Figure 5 A simplified block diagram of a device 500 through an exemplary embodiment, which can be used as... Figure 1AThe device 500 can implement the technology of the present invention, either the source device 12 or the target device 14. The device 500 can be in the form of a computing system comprising multiple computing devices, or it can be in the form of a single computing device, such as a mobile phone, tablet computer, laptop computer, desktop computer, etc.

[0149] The processor 502 in device 500 may be a central processing unit. Alternatively, processor 502 may be any other type of device or multiple devices capable of manipulating or processing information that is present or developed in the future. Although the disclosed implementation may be implemented with a single processor (e.g., processor 502), speed and efficiency may be improved by using more than one processor.

[0150] In one implementation, the memory 504 in device 500 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as memory 504. Memory 504 may include code and data 506 accessed by processor 502 via bus 512. Memory 504 may also include an operating system 508 and an application program 510, which includes at least one program that causes processor 502 to perform the methods described herein. For example, application program 510 may include applications 1 to N, which also include a video decoding application that performs the methods described herein. Device 500 may also include other memory in the form of auxiliary memory 514, for example, auxiliary memory 514 may be a memory card used with a mobile computing device. Since video communication sessions may include a large amount of information, it may be stored, in whole or in part, in auxiliary memory 514 and loaded into memory 504 as needed for processing.

[0151] Device 500 may also include one or more output devices, such as display 518. In one example, display 518 may be a touch-sensitive display that combines a display with a touch-sensitive element operable to sense touch input. Display 518 may be coupled to processor 502 via bus 512. In addition to or as an alternative to display 518, other output devices may be provided that allow the user to program or otherwise use device 500. When the output device is or includes a display, the display may be implemented in various ways, including via a liquid crystal display (LCD), a cathode ray tube (CRT) display, a plasma display, or a light-emitting diode (LED) display, such as an organic LED (OLED) display.

[0152] The device 500 may also include an image sensing device 520 or communicate with an image sensing device 520, such as a camera or any existing or later-developed image sensing device 520 capable of sensing images, such as images of a user operating the device 500. The image sensing device 520 may be positioned toward a user operating the device 500. In one example, the position and optical axis of the image sensing device 520 may allow the field of view to include an area directly adjacent to the display 518 from which the display 518 can be seen.

[0153] The device 500 may also include or communicate with a sound sensing device 522, such as a microphone or any other existing or later-developed sound sensing device capable of sensing sounds in the vicinity of the device 500. The sound sensing device 522 may be positioned toward a user operating the device 500 and may be used to receive sounds, such as voice or other words, emitted by the user while operating the device 500.

[0154] although Figure 5The processor 502 and memory 504 of device 500 are depicted as integrated into a single unit, but other configurations can be used. The operation of processor 502 can be distributed across multiple machines (each with one or more processors), which can be directly coupled or coupled via a local area network or other network. Memory 504 can be distributed across multiple machines, such as network-based memory or memory across multiple machines performing the operations of device 500. Although described herein as a single bus, bus 512 of device 500 can consist of multiple buses. Furthermore, auxiliary memory 514 can be directly coupled to other components of device 500 or accessed via a network, and can comprise a single integrated unit (e.g., a memory card) or multiple units (e.g., multiple memory cards). Therefore, device 500 can be implemented in a variety of configurations.

[0155] In one or more examples, the described functionality can be implemented in hardware, software, firmware, or any combination thereof. If the functionality is implemented in software, it can be stored or transmitted as one or more instructions or code in a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium can include a computer-readable storage medium, a tangible medium such as a data storage medium, or a communication medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. In this manner, a computer-readable medium can generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium, such as a signal or carrier wave. A data storage medium can be any available medium accessible by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described herein. A computer program product can include a computer-readable medium.

[0156] Video compression techniques such as motion compensation, intra-frame prediction, and loop filtering have proven effective and are therefore applied in various video decoding standards such as H.264 / AVC and H.265 / HEVC. For example, intra-frame prediction can be performed on I-frames or I-strips when no reference image is available, or when inter-frame prediction decoding is not used for the current block or image. Reference samples for intra-frame prediction are typically obtained from previously decoded (or reconstructed) neighboring blocks in the same image. For example, both H.264 / AVC and H.265 / HEVC use boundary samples of neighboring blocks as references for intra-frame prediction. Several different intra-frame prediction modes are used to cover different texture or structural features. In each mode, a different method for obtaining the prediction signal is used. For example, ... Figure 6 As shown, H.265 / HEVC supports a total of 35 intra-frame prediction modes.

[0157] Description of Intra-Frame Prediction Algorithms in H.265 / HEVC

[0158] For intra-frame prediction, the decoding boundary samples of adjacent blocks are used as a reference. The encoder selects the optimal luminance intra-frame prediction mode for each block from 35 options: 33 directional prediction modes, 1 DC mode, and 1 planar mode. Figure 6 The document defines the mapping relationship between intra-frame prediction directions and intra-frame prediction mode numbers. It should be noted that in the latest video decoding technologies, such as VVC (versatile video coding), 65 or more intra-frame prediction modes have been developed, capable of capturing arbitrary edge directions in natural video rendering.

[0159] like Figure 7 As shown, block "CUR" is the current block to be predicted, and grayscale samples along the boundaries of adjacent building blocks (to the left and above the current block) serve as reference samples. The predicted signal can be obtained by mapping the reference samples according to a specific method indicated by the intra-frame prediction mode.

[0160] refer to Figure 7 W is the width of the current block, and H is the height of the current block. W1 is the number of top reference samples. H1 is the number of remaining reference samples. Typically, W1 > W, H1 > H, meaning that the top reference samples also include the top-right (adjacent) reference samples, and the left reference samples also include the bottom-left (adjacent) reference samples. For example, H1 = 2 × H, W1 = 2 × W, or H1 = H + W, W1 = W + H.

[0161] Reference samples are not always available. For example, such as Figure 8 As shown, after the availability check process, W2 samples are available at the top or above the current block, and H2 samples are available to the left of the current block. W3 samples are unavailable at the top, and H3 samples are unavailable to the left.

[0162] Before obtaining the predicted signal, these unavailable samples need to be replaced (or filled) with available samples. For example, the reference samples are scanned clockwise, and the unavailable samples are replaced with the latest available sample values. If the lower part of the left reference sample is unavailable, it is replaced with the value of the nearest available reference sample.

[0163] Reference Sample Availability Check

[0164] Reference sample availability check refers to checking whether a reference sample is available. For example, if the reference sample has been reconstructed, then the reference sample is available.

[0165] In existing methods, the availability check process is usually completed by checking the luminance samples. That is, even for chrominance component blocks, the availability of the reference sample is obtained by checking the availability of the corresponding luminance samples.

[0166] In the method proposed in this invention, for a block, a reference sample availability check process is performed by examining samples of its own components. These components can be Y components, Cb components, or Cr components. According to the provided method, for a block, a reference sample availability check process is performed by examining samples of its corresponding components. These components can be luminance components or chrominance components, and the chrominance component can simultaneously include Cb and Cr components. The Cb and Cr components are referred to as chrominance components, and their distinction is not made during the availability detection process.

[0167] This invention proposes a method that focuses on the following two aspects.

[0168] According to a first aspect of the invention, for a block, the reference sample availability check process is performed by examining samples of its own components. These components can be Y components, Cb components, or Cr components.

[0169] For block Y, the availability of the reference sample for block Y is determined by checking whether the reference sample is available (e.g., by checking whether the reference sample has been reconstructed).

[0170] For block Cb, the availability of the reference sample for block Cb is determined by checking whether the reference sample is available (e.g., by checking whether the reference sample has been reconstructed).

[0171] For Cr blocks, the availability of the reference sample for the Cr block is determined by checking whether the reference sample is available (e.g., by checking whether the reference sample has been reconstructed).

[0172] According to a second aspect of the invention, for a block, a reference sample availability check process is performed by examining samples of its corresponding components. Here, a component can be a luminance component or a chromaticity component. The Y component is called the luminance component. Both the Cb component and the Cr component are called chromaticity components. Here, the Cb component and the Cr component are not distinguished during the availability check process.

[0173] For a luminance block, the availability of the reference sample for the luminance block is determined by checking whether the reference sample is available (e.g., by checking whether the reference sample has been reconstructed).

[0174] For a chroma block, the availability of the reference sample for the chroma block is determined by checking whether the reference sample is available (e.g., by checking whether the reference sample has been reconstructed).

[0175] In existing solutions, for each block, regardless of which component it belongs to, the reference sample availability check is performed by examining luminance samples. In contrast, in the method proposed in this invention, for blocks belonging to different components, the reference sample availability check is performed by examining samples from different components.

[0176] It should be noted here that the method proposed in this invention is used to obtain the availability of reference samples for blocks decoded into intra-frame prediction modes. This method can be implemented by, respectively, as follows: Figure 2 and Figure 3 The intra-prediction module 254 or 354 shown is executed. Therefore, this method is applicable at both the decoder and encoder ends. The availability check process for reference samples of a block is the same in both the encoder and decoder.

[0177] Figure 9 A simplified flowchart of a method for obtaining a reconstructed signal, provided as an exemplary embodiment of the present invention. (See reference...) Figure 9 For a block decoded to an intra-frame prediction mode, in order to obtain a reconstructed block (or signal, or sample), the method includes first acquiring the prediction (or prediction signal, or sample) of the block (901). Subsequently, the method includes acquiring the residual (or residual signal) of the block (902). Then, the method includes obtaining the reconstructed block by adding the residual to the prediction of the current block (903).

[0178] Figure 10 This is a simplified flowchart illustrating the process of obtaining a prediction (or prediction signal) according to an exemplary embodiment of the present invention. (See reference) Figure 10 To obtain the predicted signal, the method first includes acquiring the intra-prediction mode of the current block, such as a planar mode or a DC mode (1001). Subsequently, the method includes acquiring the availability of reference samples for the components of the current block (1002). In one embodiment, the component includes a Y component, a Cb component, or a Cr component. In another embodiment, the component includes a luma component or a chroma component. The method can acquire the availability of reference samples by checking the availability of Y samples within adjacent Y blocks, where adjacent Y blocks include reference samples; or by checking the availability of Cb samples within adjacent Cb blocks, where adjacent Cb blocks include reference samples; or by checking the availability of Cr samples within adjacent Cr blocks, where adjacent Cr blocks include reference samples.

[0179] When the method determines that an unavailable reference sample exists (yes in 1003), the method includes replacing or filling the unavailable reference sample with an available reference sample (1004). Thereafter, the method includes obtaining a prediction of the block based on the intra-prediction mode and the replaced reference sample (1005). In 1003, when the method determines that no unavailable reference sample exists (no in 1003), step 1005 of the method is performed, which includes obtaining a prediction of the current block based on the intra-prediction mode and the available reference sample. In 1006, the method includes reconstructing the current block based on the prediction.

[0180] It should be noted that the embodiments of the present invention described below relate to the process of obtaining a predicted signal and the process of improving the availability check of a reference sample.

[0181] Example 1

[0182] In this embodiment, the availability check process is performed by examining samples of its own components.

[0183] refer to Figure 10 For blocks decoded into intra-frame prediction mode, in order to obtain the prediction signal, the method includes:

[0184] Step 1: Obtain the intra-frame prediction mode (1001)

[0185] Intra-prediction modes are obtained by parsing the intra-prediction mode-related syntax in the bitstream. For example, for a Y block, it is necessary to parse intra_luma_mpm_flag, intra_luma_mpm_idx, or intra_luma_mpm_remainder. For a Cb / Cr block, it is necessary to parse the intra_chroma_pred_mode signal. The parsed syntax can then be used to obtain the intra-prediction mode for the current block.

[0186] Step 2: Perform the reference sample availability check process (1002)

[0187] This step includes checking the availability of a reference sample for the current block.

[0188] In one embodiment, when the block is a Y block, reference sample availability is obtained by examining reference samples of the Y component. Alternatively, reference sample availability is obtained by examining adjacent Y samples (e.g., by examining boundary samples of adjacent Y blocks).

[0189] In one embodiment, when the block is a Cb block, reference sample availability is obtained by checking reference samples of the Cb components. Alternatively, reference sample availability is obtained by checking adjacent Cb samples (e.g., by checking boundary samples of adjacent Cb blocks).

[0190] In one embodiment, when the block is a Cr block, reference sample availability is obtained by examining reference samples of the Cr component. Alternatively, reference sample availability is obtained by examining adjacent Cr samples (e.g., by examining boundary samples of adjacent Cr blocks).

[0191] Step 3: Determine if there are any unusable reference samples (1003)

[0192] In step 1003, the method includes determining whether an unavailable reference sample exists. If the method determines that an unavailable reference sample exists, step 4 (1004) of the method is executed; otherwise, step 5 (1005) of the method is executed.

[0193] Step 4: Replace the reference sample (1004)

[0194] Reference sample replacement refers to obtaining the sample value of an unavailable reference sample using the available reference sample value. In one embodiment, the reference samples are scanned clockwise and the unavailable sample is replaced with the latest available sample value. If the lower part of the left reference sample is unavailable, it is replaced with the value of the nearest available reference sample.

[0195] Step 5: Obtain the predicted signal (1005)

[0196] After obtaining the reference sample and the intra-prediction mode, the prediction signal can be obtained by mapping the reference sample to the current block. The mapping method is represented by the intra-prediction mode.

[0197] After obtaining the prediction signal for the current block, the reconstructed signal for the current block can be obtained by adding the residual signal to the obtained prediction signal. Figure 9 (903).

[0198] It should be noted here that after reconstructing the current block, samples within the block region can be marked as available, such as... Figure 11 As shown. That is, when the current block is a Y block, all Y samples in the area covered by the block region are available; when the current block is a Cb block, all Cb samples in the area covered by the block region are available; and when the current block is a Cr block, all Cr samples in the area covered by the block region are available.

[0199] An example of the specifications for these chapters is shown below:

[0200] Reference Sample Usability Labeling Process

[0201] The input to this process is:

[0202] –Sample position (xTbCmp, yTbCmp), specifies the top-left sample of the current transform block relative to the top-left sample of the current image;

[0203] – The variable refIdx specifies the intra-frame prediction reference line index.

[0204] – The variable refW specifies the width of the reference sample.

[0205] – The variable refH specifies the height of the reference sample.

[0206] – The variable cIdx specifies the color component of the current block.

[0207] For in-sample prediction, the output of this process is the reference sample refUnfilt[x][y], where x = -1 - refIdx, y = -1 - refIdx..refH-1 and x = -refIdx..refW-1, y = -1 - refIdx.

[0208] The adjacent samples refW+refH+1+(2*refIdx) refUnfilt[x][y] are the samples constructed before the in-loop filtering process, where x=-1-refIdx, y=-1-refIdx..refH-1 and x=-refIdx..refW-1, y=-1-refIdx, obtained as follows:

[0209] – Adjacent locations (xNbCmp, yNbCmp) are represented by the following parameters:

[0210] (xNbCmp, yNbCmp) = (xTbCmp + x, yTbCmp + y) (310)

[0211] – Call the adjacent block availability acquisition procedure specified in Clause 6.4.4, set the current sample position (xCurr, yCurr) to (xTbCmp, yTbCmp), the adjacent sample position to (xNbCmp, yNbCmp), set checkPredModeY to false, take cIdx as input, and assign the output value availableN.

[0212] – The refUnfilt[x][y] for each sample is obtained as follows:

[0213] – If availableN is false, then mark the sample refUnfilt[x][y] as “not available for intra-frame prediction”.

[0214] Otherwise, mark the sample refUnfilt[x][y] as "available for intra-frame prediction" and assign the sample at position (xNbCmp, yNbCmp) to refUnfilt[x][y].

[0215] process of obtaining the availability of adjacent blocks

[0216] The input to this process is:

[0217] - The sample position (xCurr, yCurr) of the top-left sample of the current block relative to the top-left sample of the current image.

[0218] – The sample positions (xNbCmp, yNbCmp) covered by the top-left brightness sample of the current image relative to the adjacent block.

[0219] The variable checkPredModeY specifies whether availability depends on the prediction mode.

[0220] – The variable cIdx specifies the color component of the current block.

[0221] The output of this process is the availability of the neighboring blocks covering the location (xNbCmp, yNbCmp), denoted as availableN.

[0222] The current brightness position (xTbY, yTbY) and adjacent brightness positions (xNbY, yNbY) are obtained as follows:

[0223] (xTbY,yTbY)=(cIdx==0)? (xCurr,yCurr):

[0224] (xCurr*SubWidthC,yCurr*SubHeightC)

[0225] (xNbY,yNbY)=(cIdx==0)? (xNbCmp,yNbCmp):

[0226] (xNbCmp*SubWidthC,yNbCmp*SubHeightC)

[0227] The availability N of adjacent blocks is obtained as follows:

[0228] - Set availableN to false if one or more of the following conditions are true.

[0229] –xNbCmp is less than 0.

[0230] –yNbCmp is less than 0.

[0231] –xNbY is greater than or equal to pic_width_in_luma_samples.

[0232] –yNbY is greater than or equal to pic_height_in_luma_samples.

[0233] –IsAvailable[cIdx][xNbCmp][yNbCmp] is false.

[0234] – Adjacent blocks are included in stripes that are different from the current block.

[0235] – Adjacent blocks are included in tiles that are different from the current block.

[0236] –entropy_coding_sync_enabled_flag equals 1 and (xNbY>>CtbLog2SizeY) is greater than or equal to (xTbY>>

[0237] CtbLog2SizeY)+1.

[0238] Otherwise, set availableN to true.

[0239] AvailableN is set to false when all of the following conditions are met:

[0240] –checkPredModeY is true.

[0241] –availableN is true.

[0242] –CuPredMode[0][xNbY][yNbY] is not equal toCuPredMode[0][xTbY][yTbY].

[0243] IsAvailable[cIdx][x][y] is used to store the availability information of the sample at sample position (x, y) for each component cIdx.

[0244] For cIdx=0, 0<=x<=pic_width_in_luma_samples, 0<=y<=pic_height_in_luma_samples;

[0245] cIdx=1 or 2. 0<=x<=pic_width_in_chroma_samples, 0<=y<=pic_height_in_chroma_samples;

[0246] The input to this process is:

[0247] – Position (xCurr, yCurr) specifies the position of the top-left sample of the current block relative to the top-left sample of the current image component.

[0248] – The variables nCurrSw and nCurrSh specify the width and height of the current block, respectively.

[0249] – The variable cIdx specifies the color component of the current block.

[0250] The array `predSamples` (`-(nCurrSw)x(nCurrSh)` specifies the predicted samples for the current block.

[0251] The array `resSamples` (`-(nCurrSw)x(nCurrSh)` specifies the residual samples for the current block.

[0252] The output of this process is a reconstructed image sample array recSamples.

[0253] Based on the value of the color component cIdx, the following values ​​are assigned:

[0254] – If cIdx = 0, recSamples corresponds to the reconstructed image sample array S L .

[0255] Otherwise, if cIdx equals 1, then set tuCbfChroma to equal tu_cbf_cb[xCurr][yCurr], recSamples

[0256] The corresponding reconstructed chromaticity sample array S Cb .

[0257] Otherwise (cIdx equals 2), set tuCbfChroma to equal tu_cbf_cr[xCurr][yCurr], recSamples corresponds to the reconstructed chromaticity sample array S. Cr .

[0258] Based on the value of pic_lmcs_enabled_flag, the following applies:

[0259] – If pic_lmcs_enabled_flag equals 0, then the (nCurrSw)x(nCurrSh) block of the reconstructed sample recSamples at position (xCurr, yCurr) is obtained as follows, where i = 0..nCurrSw-1, j = 0..nCurrSh-1:

[0260] recSamples[ xCurr + i ][ yCurr + j ] = Clip1(predSamples[ i ][ j ] +resSamples[ i ][ j ]) (1195)

[0261] – Otherwise (pic_lmcs_enabled_flag equals 1), the following applies:

[0262] – When cIdx equals 0, the following applies:

[0263] – Image reconstruction that invokes the mapping procedure for the luminance samples specified in Clause 8.7.5.2, where the luminance location (xCurr,

[0264] The input consists of the block width nCurrSw and height nCurrSh, the predicted luminance sample array predSamples, and the residual luminance sample array resSamples, and the output is the reconstructed luminance sample array recSamples.

[0265] - Otherwise (cIdx is greater than 0), the image reconstruction of the luminance-related chromaticity residual scaling process of the chromaticity samples specified in Clause 8.7.5.3 is invoked, where the chromaticity position (xCurr, yCurr), transform block width nCurrSw and height nCurrSh, decoded block flag of the current chromaticity transform block tuCbfChroma, predicted chromaticity sample array predSamples, and residual chromaticity sample array resSamples are taken as inputs, and the output is the reconstructed chromaticity sample array recSamples.

[0266] Perform the following assignments, where i = 0..nCurrSw-1, j = 0..nCurrSh-1:

[0267] xVb = (xCurr + i) % ( (cIdx = = 0) ? IbcBufWidthY : IbcBufWidthC)(1196)

[0268] yVb = (yCurr + j) % ( (cIdx = = 0) ? CtbSizeY : (CtbSizeY / subHeightC) ) (1197)

[0269] IbcVirBuf[ cIdx ][ xVb ][ yVb ] = recSamples[ xCurr + i ][ yCurr + j] (1198)

[0270] IsAvailable[cIdx][xCurr + i][yCurr + j] = TRUE (1199)

[0271] Another example of the specifications for these chapters is shown below:

[0272] Reference Sample Usability Labeling Process

[0273] The input to this process is:

[0274] –Sample position (xTbCmp, yTbCmp), specifies the top-left sample of the current transform block relative to the top-left sample of the current image;

[0275] – The variable refIdx specifies the intra-frame prediction reference line index.

[0276] – The variable refW specifies the width of the reference sample.

[0277] – The variable refH specifies the height of the reference sample.

[0278] – The variable cIdx specifies the color component of the current block.

[0279] For in-sample prediction, the output of this process is the reference sample refUnfilt[x][y], where x = -1 - refIdx, y = -1 - refIdx..refH-1 and x = -refIdx..refW-1, y = -1 - refIdx.

[0280] The adjacent samples refW+refH+1+(2*refIdx) refUnfilt[x][y] are the samples constructed before the in-loop filtering process, where x=-1-refIdx, y=-1-refIdx..refH-1 and x=-refIdx..refW-1, y=-1-refIdx, obtained as follows:

[0281] – Adjacent locations (xNbCmp, yNbCmp) are specified by the following parameters:

[0282] (xNbCmp, yNbCmp) = (xTbCmp + x, yTbCmp + y) (310)

[0283] – Call the adjacent block availability acquisition procedure specified in Clause 6.4.4, set the current sample position (xCurr, yCurr) to (xTbCmp, yTbCmp), the adjacent sample position to (xNbCmp, yNbCmp), set checkPredModeY to false, take cIdx as input, and assign the output value availableN.

[0284] – The refUnfilt[x][y] for each sample is obtained as follows:

[0285] – If availableN is false, then mark the sample refUnfilt[x][y] as “not available for intra-frame prediction”.

[0286] Otherwise, mark the sample refUnfilt[x][y] as "available for intra-frame prediction" and assign the sample at position (xNbCmp, yNbCmp) to refUnfilt[x][y].

[0287] process of obtaining the availability of adjacent blocks

[0288] The input to this process is:

[0289] - The sample position (xCurr, yCurr) of the top-left sample of the current block relative to the top-left sample of the current image.

[0290] – The sample positions (xNbCmp, yNbCmp) covered by the top-left brightness sample of the current image relative to the adjacent block.

[0291] The variable checkPredModeY specifies whether availability depends on the prediction mode.

[0292] – The variable cIdx specifies the color component of the current block.

[0293] The output of this process is the availability of the neighboring blocks covering the location (xNbCmp, yNbCmp), denoted as availableN.

[0294] The current brightness position (xTbY, yTbY) and adjacent brightness positions (xNbY, yNbY) are obtained as follows:

[0295] (xTbY,yTbY)=(cIdx==0)? (xCurr,yCurr):

[0296] (xNbY,yNbY)=(cIdx==0)? (xNbCmp,yNbCmp):(XXX)

[0297] (xNbCmp*SubWidthC,yNbCmp*SubHeightC)

[0298] The availability N of adjacent blocks is obtained as follows:

[0299] - Set availableN to false if one or more of the following conditions are true.

[0300] –xNbCmp is less than 0.

[0301] –yNbCmp is less than 0.

[0302] –xNbY is greater than or equal to pic_width_in_luma_samples.

[0303] –yNbY is greater than or equal to pic_height_in_luma_samples.

[0304] –IsAvailable[cIdx][xNbCmp][yNbCmp] is false.

[0305] – Adjacent blocks are included in stripes that are different from the current block.

[0306] – Adjacent blocks are included in blocks that are different from the current block.

[0307] –entropy_coding_sync_enabled_flag equals 1 and (xNbY>>CtbLog2SizeY) is greater than or equal to (xTbY>>

[0308] CtbLog2SizeY)+1.

[0309] Otherwise, set availableN to true.

[0310] AvailableN is set to false when all of the following conditions are met:

[0311] –checkPredModeY is true.

[0312] –availableN is true.

[0313] –CuPredMode[0][xNbY][yNbY] is not equal toCuPredMode[0][xTbY][yTbY].

[0314] The input to this process is:

[0315] – Position (xCurr, yCurr) specifies the position of the top-left sample of the current block relative to the top-left sample of the current image component.

[0316] – The variables nCurrSw and nCurrSh specify the width and height of the current block, respectively.

[0317] – The variable cIdx specifies the color component of the current block.

[0318] The array `predSamples` (`-(nCurrSw)x(nCurrSh)` specifies the predicted samples for the current block.

[0319] The array `resSamples` (`-(nCurrSw)x(nCurrSh)` specifies the residual samples for the current block.

[0320] The output of this process is a reconstructed image sample array recSamples.

[0321] Based on the value of the color component cIdx, the following values ​​are assigned:

[0322] – If cIdx = 0, recSamples corresponds to the reconstructed image sample array S L .

[0323] Otherwise, if cIdx equals 1, then set tuCbfChroma to equal tu_cbf_cb[xCurr][yCurr], where recSamples corresponds to the reconstructed chromaticity sample array S. Cb .

[0324] Otherwise (cIdx equals 2), set tuCbfChroma to equal tu_cbf_cr[xCurr][yCurr], recSamples corresponds to the reconstructed chromaticity sample array S. Cr .

[0325] Based on the value of pic_lmcs_enabled_flag, the following applies:

[0326] – If pic_lmcs_enabled_flag equals 0, then the (nCurrSw)x(nCurrSh) block of the reconstructed sample recSamples at position (xCurr, yCurr) is obtained as follows, where i = 0..nCurrSw-1, j = 0..nCurrSh-1:

[0327] recSamples[ xCurr + i ][ yCurr + j ] = Clip1(predSamples[ i ][ j ] +resSamples[ i ][ j ]) (1195)

[0328] – Otherwise (pic_lmcs_enabled_flag equals 1), the following applies:

[0329] – When cIdx equals 0, the following applies:

[0330] – Image reconstruction that invokes the mapping procedure for the luminance samples specified in Clause 8.7.5.2, where the luminance location (xCurr,

[0331] The input consists of the block width nCurrSw and height nCurrSh, the predicted luminance sample array predSamples, and the residual luminance sample array resSamples, and the output is the reconstructed luminance sample array recSamples.

[0332] - Otherwise (cIdx is greater than 0), the image reconstruction of the luminance-related chromaticity residual scaling process of the chromaticity samples specified in Clause 8.7.5.3 is invoked, where the chromaticity position (xCurr, yCurr), transform block width nCurrSw and height nCurrSh, decoded block flag of the current chromaticity transform block tuCbfChroma, predicted chromaticity sample array predSamples, and residual chromaticity sample array resSamples are taken as inputs, and the output is the reconstructed chromaticity sample array recSamples.

[0333] Perform the following assignments, where i = 0..nCurrSw-1, j = 0..nCurrSh-1:

[0334] xVb = (xCurr + i) % ( (cIdx = = 0) ? IbcBufWidthY : IbcBufWidthC)(1196)

[0335] yVb = (yCurr + j) % ( (cIdx = = 0) ? CtbSizeY : (CtbSizeY / subHeightC) ) (1197)

[0336] IbcVirBuf[ cIdx ][ xVb ][ yVb ] = recSamples[ xCurr + i ][ yCurr + j] (1198)

[0337] IsAvailable[cIdx][(xCurr+i)*((cIdx==0)?1:SubWidthC)][(yCurr+j)*((cIdx==0)?1:

[0338] SubHeightC) ] = TRUE (1199)

[0339] Figure 12 This is a schematic diagram illustrating the marking of samples as available in the current block within an N×N unit, provided as an exemplary embodiment of the present invention. In one embodiment, reference... Figure 12After reconstructing the current block, samples within an N×N cell are marked as available, for example, N=4 or N=2. Alternatively, N=4 for the Y component block and N=2 for the Cb / Cr component block. Any sample in a "available" cell is considered "available." That is, if a cell is marked as "available," then any sample within that cell region can also be marked as "available." In other words, when the current cell is a Y cell, all Y samples at the locations covered by the cell region are available; when the current cell is a Cb cell, all Cb samples at the locations covered by the cell region are available; and when the current block is a Cr cell, all Cr samples at the locations covered by the cell region are available.

[0340] In one embodiment, in order to check whether a reference sample is available, the process first obtains the location or index of the cell to which the sample belongs, and when the cell is determined to be available, the reference sample is considered available or marked as available.

[0341] Figure 13 Samples at the right and bottom boundaries of the current block, provided for an exemplary embodiment of the invention, are marked as available schematic diagrams. In one embodiment, reference... Figure 13 After the current block is reconstructed, only the right boundary sample and the bottom boundary sample (which will be used as reference samples for other blocks) are marked as available.

[0342] Figure 14 The samples at the right and bottom boundaries of the current block in cell N, provided for an exemplary embodiment of the present invention, are marked as available schematic diagrams. In one embodiment, reference... Figure 14 After reconstructing the current block, only the right and bottom boundary samples (which will serve as reference samples for other blocks) in cells of size N are marked as available, for example, N=4 or N=2. Alternatively, for Y component blocks, N=4; for Cb / Cr component blocks, N=2. Any sample in a "available" cell is considered "available." That is, if a cell is marked as "available," then any sample within that cell can also be marked as "available." In other words, when the current cell is a Y cell, all Y samples at the locations covered by the cell region are available; when the current cell is a Cb cell, all Cb samples at the locations covered by the cell region are available; and when the current block is a Cr cell, all Cr samples at the locations covered by the cell region are available.

[0343] To check whether a reference sample is available, in one embodiment, the process first obtains the location or index of the cell to which the sample belongs, and when the cell is determined to be available, the reference sample is marked as available.

[0344] Figure 15A schematic diagram showing samples at the right and bottom boundaries of the current block in an N×N cell, provided for an exemplary embodiment of the present invention, is provided, indicating that samples are available. In one embodiment, reference... Figure 15 After reconstructing the current block, only the right and bottom boundary samples (which will serve as reference samples for other blocks) in cells of size N×N are marked as available, for example, N=4 or N=2. Alternatively, for Y component blocks, N=4; for Cb / Cr component blocks, N=2. Any sample in a "available" cell is considered "available." That is, when a cell is marked as "available," any sample within the cell region can also be marked as "available." In other words, when the current cell is a Y cell, all Y samples at the location covered by the cell region are available; when the current cell is a Cb cell, all Cb samples at the location covered by the cell region are available; and when the current block is a Cr cell, all Cr samples at the location covered by the cell region are available.

[0345] To check whether a reference sample is available, in one embodiment, the process or method first obtains the location or index of the cell to which the sample belongs, and when the process or method determines that the cell is "available", the process or method marks the reference sample as available.

[0346] It should be noted that, according to an embodiment of the present invention, the sample or unit "availability" information for each component (a total of 3 components: Y component, Cb component, and Cr component) is stored in a memory.

[0347] It should be noted that even if the block is decoded into inter-frame prediction mode, the "available" marking method can still be applied after the block is reconstructed.

[0348] Example 2

[0349] In this embodiment, the availability check process is performed by examining samples of the corresponding components.

[0350] The difference between Example 2 and Example 1 is that, in the usability testing process, the Cb and Cr components are not distinguished. For a block, the usability test of the reference sample is completed by checking samples of its corresponding components, such as... Figure 16 As shown. Here, the components can be either luminance components or chrominance components. The Y component is called the luminance component. The Cb and Cr components are both called chrominance components.

[0351] The difference between Example 2 and Example 1 lies only in step 2; the other steps are the same as in Example 1. Step 2 of Example 2 is described in detail below.

[0352] Step 2: Perform a reference sample availability check process, including checking the reference sample availability of the current block. Figure 10 (1002 in the middle).

[0353] In one embodiment, when the block is a luminance block, the availability of a reference sample is obtained by checking a reference luminance sample.

[0354] In one embodiment, when the block is a chroma block, the availability of a reference sample is obtained by checking a reference chroma sample.

[0355] refer to Figure 10 After checking the availability of the reference sample (1002), the method includes determining whether there is an unavailable reference sample (1003). When the method determines that there is an unavailable reference sample (yes in 1003), step 1004 of the method is executed; otherwise, step 1005 of the method is executed.

[0356] It should be noted that the marking method discussed in Example 1 can be directly used in Example 2. The only difference is in the chromaticity component. Only after reconstructing the Cb component block and the Cr component block can the chromaticity sample in the block area be marked as "available".

[0357] For example, for a luma block, samples within the block region can be marked as available during block reconstruction. For a chroma block, samples within that block region can be marked as available after reconstructing the Cb and Cr blocks. That is, when the current block is a luma block, all luma samples in the locations covered by the block region are marked as available. When the current block is a chroma block, all chroma samples in the locations covered by the block region are marked as available after reconstructing the Cb and Cr blocks.

[0358] Other marking methods in Example 1 can also be used similarly.

[0359] It should be noted that, in the embodiments of the present invention, the "availability" information of samples or units for each component (two components in total, luminance component and chrominance component) is stored in memory. That is, the Cb component and the Cr component will share the same "availability" or "availability" information.

[0360] It should be noted that even if the block is decoded into inter-frame prediction mode, the "available" marking method can still be used after the block is reconstructed.

[0361] The following section describes the application of the encoding and decoding methods shown in the above embodiments, as well as the systems using these methods.

[0362] Figure 17This is a block diagram of a content provisioning system 3100 for implementing a content distribution service. The content provisioning system 3100 includes a capture device 3102, a terminal device 3106, and optionally a display 3126. The capture device 3102 and the terminal device 3106 communicate via a communication link 3104. This communication link may include the aforementioned communication channel 13. The communication link 3104 includes, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination thereof.

[0363] The capture device 3102 generates data and can encode the data using the encoding method shown in the above embodiments. Alternatively, the capture device 3102 can distribute the data to a streaming server (not shown), which encodes the data and sends the encoded data to the terminal device 3106. The capture device 3102 includes, but is not limited to, cameras, smartphones or tablets, computers or laptops, video conferencing systems, PDAs, in-vehicle devices, or any combination thereof. For example, the capture device 3102 may include the source device 12 as described above. When the data includes video, the video encoder 20 included in the capture device 3102 can actually perform video encoding processing. When the data includes audio (i.e., voice), the audio encoder included in the capture device 3102 can actually perform audio encoding processing. In some practical scenarios, the capture device 3102 distributes the encoded video and audio data by multiplexing them together. In other practical scenarios, such as in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and the encoded video data separately to the terminal device 3106.

[0364] In the content delivery system 3100, the terminal device 310 receives and reproduces encoded data. The terminal device 3106 can be a device with data reception and recovery capabilities, such as a smartphone or tablet 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof, or a device capable of decoding the aforementioned encoded data. For example, the terminal device 3106 may include the target device 14 as described above. When the encoded data includes video, the video decoder 30 included in the terminal device prioritizes video decoding. When the encoded data includes audio, the audio decoder included in the terminal device prioritizes audio decoding.

[0365] For terminal devices with displays, such as smartphones or tablets 3108, computers or laptops 3110, network video recorders (NVRs) / digital video recorders (DVRs) 3112, TVs 3114, personal digital assistants (PDAs) 3122, or in-vehicle devices 3124, the terminal device can feed the decoded data to its display. For terminal devices without displays (e.g., STBs 3116, video conferencing systems 3118, or video surveillance systems 3120), an external display 3126 is connected to receive and display the decoded data.

[0366] Each device in this system can use an image encoding device or an image decoding device as shown in the above embodiments when performing encoding or decoding.

[0367] Figure 18This is a diagram illustrating an example structure of terminal device 3106. After terminal device 3106 receives a stream from capture device 3102, protocol proceeding unit 3202 analyzes the stream's transport protocol. Protocols include, but are not limited to, Real-Time Streaming Protocol (RTSP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real-time Transport Protocol (RTP), Real-Time Messaging Protocol (RTMP), or any combination thereof.

[0368] After processing the stream, the protocol processing unit 3202 generates a stream file. The file is then output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As mentioned above, in some practical scenarios, such as in a video conferencing system, the encoded audio data and encoded video data are not multiplexed. In this case, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without going through the demultiplexing unit 3204.

[0369] Through demultiplexing, a video elementary stream (ES), an audio ES, and optional subtitles are generated. Video decoder 3206 includes video decoder 30 as described in the above embodiments, which decodes the video ES to generate video frames using the decoding method shown in the above embodiments, and feeds the data to synchronization unit 3212. Audio decoder 3208 decodes the audio ES to generate audio frames, and feeds the data to synchronization unit 3212. Alternatively, the video frame can be stored in a buffer before being fed to synchronization unit 3212. Figure 18 (Not shown in the image). Similarly, before feeding the audio frame to the synchronization unit 3212, the audio frame can be stored in a buffer (not shown in the image). Figure 18 (Not shown in the text)

[0370] Synchronization unit 3212 synchronizes video and audio frames and provides the video / audio to video / audio display 3214. For example, synchronization unit 3212 synchronizes the presentation of video and audio information. The information can be decoded in the syntax using timestamps about the representation of the decoded audio and video data and timestamps about the transmission of the data stream itself.

[0371] If the stream includes subtitles, the subtitle decoder 3210 decodes the subtitles, synchronizes the subtitles with the video and audio frames, and provides the video / audio / subtitles to the video / audio / subtitle display 3216.

[0372] The present invention is not limited to the above-described system. The image encoding device or image decoding device in the above embodiments can be integrated into other systems, such as automotive systems.

[0373] According to the present invention, the described methods and processes can be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, the method or process can be executed by instructions or program code stored in a computer-readable medium and executed by a hardware processing unit.

[0374] In the proposed method, the available information for each sample is stored in each component, which enables more accurate provision of available information during intra-frame prediction.

[0375] By way of example and not limitation, computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store required program code in the form of instructions or data structures and is accessible by a computer. Furthermore, any connection can be referred to as computer-readable media. For example, if instructions are transmitted from a website, server, or other remote source via coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (e.g., infrared, radio, microwave, etc.), the definition of media includes coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (e.g., infrared, radio, and microwave, etc.). However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient tangible storage media. Disks and optical discs as used herein include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where discs typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0376] Instructions or program code can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein can refer to any of the above-described structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided in dedicated hardware and / or software modules for encoding and decoding, or incorporated into combined codecs. Furthermore, these techniques can be fully implemented in one or more circuit or logic elements.

[0377] The technology of this invention can be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or a set of ICs (e.g., chipsets). This invention describes various components, modules, or units to emphasize functional aspects of the apparatus used to perform the disclosed technology, but these components, modules, or units are not necessarily required to be implemented by different hardware units. Rather, as described above, various units can be combined in a codec hardware unit, or provided by a collection of interoperable hardware units including one or more processors as described above, combined with suitable software and / or firmware.

[0378] While several embodiments have been provided in this invention, it should be understood that the disclosed systems and methods may be implemented in many other specific forms without departing from the spirit or scope of the invention. These examples are to be regarded as illustrative rather than limiting and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.

[0379] Furthermore, without departing from the scope of the invention, the technologies, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments may be combined or integrated with other systems, modules, technologies, or methods. Other items shown or discussed as mutually coupled, directly coupled, or communicating may be indirectly coupled or communicating through some interface, device, or intermediate component via electrical, mechanical, or other means. Other examples of changes, substitutions, and modifications can be determined by those skilled in the art and can be made without departing from the spirit and scope disclosed herein.

Claims

1. An image intra-frame prediction device, characterized in that, include: Memory, the memory including instructions; One or more processors communicate with the memory, wherein the one or more processors execute the instructions to: Get the intra-prediction mode for the current block; The availability N of the reference samples for the cIdx component of the current block is obtained, where cIdx takes values ​​of 0, 1, and 2, where 0 indicates the luminance component and 1 and 2 indicate the chrominance component. When IsAvailable[cIdx][x][y] is false, availableN is false, where IsAvailable[cIdx][x][y] is used to indicate the availability information of the sample at the sample position (x, y) for each cIdx component. Replace unavailable reference samples with available reference samples; The prediction of the current block is obtained based on the intra-frame prediction mode and the replaced reference sample.

2. The apparatus according to claim 1, characterized in that, The chromaticity components include Cb components or Cr components, and the one or more processors execute the instructions to: The availability of the reference sample is obtained by checking the availability of Cb samples within adjacent Cb blocks, wherein the adjacent Cb blocks include the reference sample; or The availability of the reference sample is obtained by checking the availability of Cr samples within adjacent Cr blocks, wherein the adjacent Cr blocks include the reference sample.

3. The apparatus according to claim 1 or 2, characterized in that, All Cb samples in the reconstruction block of the current block are marked as available; or, all Cr samples in the reconstruction block of the current block are marked as available.

4. The apparatus according to claim 1 or 2, characterized in that, The one or more processors execute the instructions to: According to the intra-frame prediction mode, the prediction of the current block is obtained by mapping the replaced reference sample to the current block; The current block is reconstructed by adding the residual to the prediction of the current block; After reconstructing the current block, all samples in the block region, which is the area where the current block is located, are marked as available.

5. The apparatus according to claim 4, characterized in that, The block region comprises multiple units, each unit having a unit region, and all samples in the unit region are marked as available.

6. The apparatus according to claim 5, characterized in that, When the current block is a Cb block, all Cb samples in the locations covered by the block region are marked as available; or... When the current block is a Cr block, all Cr samples in the locations covered by the block region are marked as available.

7. The apparatus according to any one of claims 1 to 2, 5 to 6, characterized in that, The availability of reference samples for each component is preserved.

8. The apparatus according to any one of claims 1 to 2, 5 to 6, characterized in that, The chromaticity component consists of a Cb component and a Cr component, and the Cb component and the Cr component share the same availability information.

9. A method for intra-frame prediction of an image performed by an encoder or decoder, characterized in that, include: Get the intra-prediction mode for the current block; The availability N of the reference samples for the cIdx component of the current block is obtained, where cIdx takes values ​​of 0, 1, and 2, where 0 indicates the luminance component and 1 and 2 indicate the chrominance component. When IsAvailable[cIdx][x][y] is false, availableN is false, where IsAvailable[cIdx][x][y] is used to indicate the availability information of the sample at the sample position (x, y) for each cIdx component. Replace unavailable reference samples with available reference samples; The prediction of the current block is obtained based on the intra-frame prediction mode and the replaced reference sample.

10. The method according to claim 9, characterized in that, The chromaticity components include either the Cb component or the Cr component: The availability of the reference sample is obtained by checking the availability of Cb samples within adjacent Cb blocks, wherein the adjacent Cb blocks include the reference sample; or, The availability of the reference sample is obtained by checking the availability of Cr samples within adjacent Cr blocks, wherein the adjacent Cr blocks include the reference sample.

11. The method according to claim 9 or 10, characterized in that, All Cb samples in the reconstruction block of the current block are marked as available; or, all Cr samples in the reconstruction block of the current block are marked as available.

12. The method according to claim 9 or 10, characterized in that, The method further includes: According to the intra-frame prediction mode, the prediction of the current block is obtained by mapping the replaced reference sample to the current block; The current block is reconstructed by adding the residual to the prediction of the current block; After reconstructing the current block, all samples in the block region, which is the area where the current block is located, are marked as available.

13. The method according to claim 12, characterized in that, The block region comprises multiple units, each of which has a unit region.

14. The method according to claim 13, characterized in that: When the current block is a Cb block, all Cb samples in the locations covered by the block region are marked as available; or... When the current block is a Cr block, all Cr samples in the locations covered by the block region are marked as available.

15. The method according to claim 13, characterized in that, In cell N×1 or 1×N, only the right boundary sample and the bottom boundary sample are marked as available.

16. The method according to claim 13, characterized in that, In the N×N cell, only the right boundary sample and the bottom boundary sample are marked as available.

17. The method according to any one of claims 9 to 10, 13 to 16, characterized in that, The availability of reference samples for each component is preserved.

18. The method according to claim 9, characterized in that, The chromaticity component consists of a Cb component and a Cr component, and the Cb component and the Cr component share the same availability information.

19. The method according to claim 12, characterized in that, After a cell is reconstructed in the current block, all samples in the cell region of that cell are marked as available.

20. The method according to claim 12, characterized in that, After the current block is reconstructed, samples at the right and bottom boundaries of the current block are marked as available.

21. The method according to claim 12, characterized in that, After the current block is reconstructed, only the right boundary sample and the bottom boundary sample are marked as available in cell N×1 or 1×N.

22. The method according to claim 12, characterized in that, After the current block is reconstructed, only the right boundary sample and the bottom boundary sample are marked as available in the N×N cell.

23. A computer program product, characterized in that, Includes program code that, when executed by a processor, causes the processor to implement the method as described in any one of claims 9 to 22.

24. A non-transitory computer-readable storage medium, characterized in that, Stored program instructions that, when executed by a processor, cause the processor to implement the method as described in any one of claims 9 to 22.

Citation Information

Patent Citations

  • Multi-type-tree framework for video coding

    US20170208336A1

  • Adaptive cross component residual prediction

    CN107211124A

  • Methods of reference quantization parameter derivation for signaling of quantization parameter in QUAD-tree plus binary tree structure

    WO2018018486A1