Encoding method, decoding method, encoder, decoder, and storage medium
By determining multiple reference block sets in the light field video and performing weighted prediction, the problems of large data volume and low compression efficiency in light field video are solved, and more efficient coding performance is achieved.
Patent Information
- Application Number
- PCT/CN2024/087823
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-15
- Publication Date
- 2025-10-23
AI Technical Summary
Light field videos contain massive amounts of data, and existing video coding technologies have failed to effectively utilize their characteristics, resulting in low compression efficiency.
Prediction accuracy is improved by determining multiple reference block sets in the reference image and leveraging the spatial correlation of the light field video. This includes determining a first reference block and one or more second reference block sets, and performing weighted processing to generate a prediction value for the current block.
It improves the coding prediction accuracy and compression efficiency of light field video, reduces the amount of data, and is suitable for the characteristics of light field video.
Smart Images

Figure CN2024087823_23102025_PF_FP_ABST
Abstract
Description
Coding method, coder and storage medium TECHNICAL FIELD
[0001] The present application relates to the technical field of video coding, and in particular to a coding method, a coder and a storage medium. BACKGROUND
[0002] Light field video is an image or video captured by a multi-lens camera or a multi-lens camera array, and is currently studied by a moving pictures experts group lenslet video coding (MPEG LVC) working group. The data volume of light field video is huge, which brings challenges to its transmission and compression. Therefore, video coding technology for light field video is particularly important.
[0003] SUMMARY
[0004] Embodiments of the present application provide a coding method, a coder and a storage medium. Each aspect involved in the present application is introduced below.
[0005] In a first aspect, a decoding method is provided, applied to a decoder, including: parsing a code stream, determining a reference image of a current block and a first reference block in the reference image; determining one or more second reference block sets according to the first reference block; determining a prediction value of the current block according to the first reference block and the second reference block set; and determining a reconstructed block of the current block according to the prediction value of the current block.
[0006] In a second aspect, a coding method is provided, applied to an encoder, including: determining a reference image of a current block and a first reference block in the reference image; determining one or more second reference block sets according to the first reference block; determining a prediction value of the current block according to the first reference block and the second reference block set; and determining a residual value of the current block according to the prediction value of the current block.
[0007] In a third aspect, a decoder is provided, including: a decoding module configured to parse a code stream, determine a reference image of a current block and a first reference block in the reference image; a first determining module configured to determine one or more second reference block sets according to the first reference block; a second determining module configured to determine a prediction value of the current block according to the first reference block and the second reference block set; and a third determining module configured to determine a reconstructed block of the current block according to the prediction value of the current block.
[0008] In a fourth aspect, an encoder is provided, comprising: a first determining module configured to determine a reference image of a current block and a first reference block in the reference image; a second determining module configured to determine one or more second reference block sets according to the first reference block; a third determining module configured to determine a prediction value of the current block according to the first reference block and the second reference block sets; and a fourth determining module configured to determine a residual value of the current block according to the prediction value of the current block.
[0009] In a fifth aspect, a decoder is provided, comprising: a memory configured to store a computer program; and a processor configured to execute the method of the first aspect when running the computer program.
[0010] In a sixth aspect, an encoder is provided, comprising: a memory configured to store a computer program; and a processor configured to execute the method of the second aspect when running the computer program.
[0011] In a seventh aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed to implement the method of the first aspect or the second aspect.
[0012] In an eighth aspect, a computer program product is provided, which comprises a computer program, and the computer program is executed to implement the method of the first aspect or the second aspect.
[0013] In a ninth aspect, a non-volatile computer readable storage medium storing a bitstream is provided, and the bitstream is generated by using an encoding method of an encoder, or the bitstream is decoded by using a decoding method of a decoder, wherein the decoding method is the method of the first aspect, and the encoding method is the method of the second aspect.
[0014] The embodiments of the present application are based on the characteristics of light field video, and the current block is predicted according to the plurality of reference blocks in the reference image, which helps to improve the prediction accuracy of the current block. BRIEF DESCRIPTION OF DRAWINGS
[0015] FIG. 1 is a structural schematic diagram of a video encoder to which the embodiments of the present application can be applied.
[0016] FIG. 2 is a structural schematic diagram of a video decoder to which the embodiments of the present application can be applied.
[0017] FIG. 3 is an example diagram of an image in a light field video.
[0018] FIG. 4 is an example diagram of a processing process of a light field video.
[0019] FIG. 5 is a flowchart of a decoding method provided by the embodiments of the present application.
[0020] FIG. 6 is an example diagram of a determination manner of a first reference block according to an embodiment of the present application.
[0021] FIG. 7 is an example diagram of a second reference block according to an embodiment of the present application.
[0022] FIG. 8 is an example diagram of a determination manner of the second reference block in FIG. 7.
[0023] FIG. 9 is a flow diagram of an encoding method according to an embodiment of the present application.
[0024] FIG. 10 is a structural diagram of a decoder according to an embodiment of the present application.
[0025] FIG. 11 is a structural diagram of a decoder according to another embodiment of the present application.
[0026] FIG. 12 is a structural diagram of an encoder according to an embodiment of the present application.
[0027] FIG. 13 is a structural diagram of an encoder according to another embodiment of the present application. DETAILED DESCRIPTION
[0028] FIG. 1 is a schematic block diagram of a video encoder according to an embodiment of the present application.
[0029] It should be understood that the video encoder 100 can be used for lossy compression of images, and can also be used for lossless compression of images. The lossless compression can be visually lossless compression or mathematically lossless compression.
[0030] The video encoder 100 can be applied to image data in YCbCr (YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2 or 4:4:4, Y represents brightness (Luma), Cb (U) represents blue chroma, and Cr (V) represents red chroma. U and V represent chroma (Chroma) for describing color and saturation. For example, in terms of color format, 4:2:0 means that there are 4 brightness components and 2 chroma components (YYYYCbCr) for every 4 pixels, 4:2:2 means that there are 4 brightness components and 4 chroma components (YYYYCbCrCbCr) for every 4 pixels, and 4:4:4 means full pixel display (YYYYCbCrCbCrCbCrCbCr).
[0031] For example, the video encoder 100 reads the video data, partitions, for each picture in the video data, a picture into a number of coding tree units (CTUs), which can be referred to in some examples as "treeblocks," "largest coding units" (LCUs) or "coding tree blocks" (CTBs). Each CTU can be associated with a same size block of pixels within the picture. Each pixel can correspond to a luminance (luma) sample and two chrominance (chroma) samples. Thus, each CTU can be associated with a luma sample block and two chroma sample blocks. A CTU size can be, for example, 128x128, 64x64, 32x32, and the like. A CTU can be further partitioned into coding units (CUs) for coding, which can be square or non-square in shape. CUs can be further partitioned into prediction units (PUs) and transform units (TUs) to separate the processes of prediction, motion compensation, and transform, to provide more flexibility in processing. In one example, a CTU is quad-tree partitioned into CUs, and a CU is quad-tree partitioned into TUs and PUs.
[0032] A video encoder and a video decoder can support various PU sizes. Assuming that a size of a particular CU is 2Nx2N, the video encoder and the video decoder can support a PU size of 2Nx2N or NxN for intra prediction, and support symmetric PUs of 2Nx2N, 2NxN, Nx2N, NxN, or similar sizes for inter prediction. The video encoder and the video decoder can also support asymmetric PUs of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter prediction.
[0033] In some embodiments, as shown in FIG. 1, the video encoder 100 can include a prediction unit 110, a residual unit 120, a transform / quantization unit 130, an inverse transform / quantization unit 140, a reconstruction unit 150, a loop filter unit 160, a decoded picture buffer 170, and an entropy encoding unit 180. It is noted that the video encoder 100 can include more, less, or different functional components.
[0034] Optionally, in this application, a current block can be referred to as a current coding unit (CU) or a current prediction unit (PU), and the like. A prediction block can also be referred to as a predicted image block or an image predicted block, and a reconstructed image block can also be referred to as a reconstructed block or an image reconstructed block.
[0035] In some embodiments, the prediction unit 110 includes an inter prediction unit 111 and an intra prediction unit 112. Since there is a strong correlation between adjacent pixels in one image of a video, the method of using intra prediction in video coding technology eliminates the spatial redundancy between adjacent pixels. Since there is a strong similarity between adjacent images in a video, the method of using inter prediction in video coding technology eliminates the temporal redundancy between adjacent images, thereby improving the coding efficiency.
[0036] The inter prediction unit 111 can be used for inter prediction, which can include motion estimation and motion compensation, can refer to image information of different images, and inter prediction uses motion information to find a reference block in a reference image, generates a prediction block according to the reference block, and is used to eliminate temporal redundancy. Inter prediction uses motion information to find a reference block in a reference image, and generates a prediction block according to the reference block. The motion information includes a reference image list where the reference image is located, a reference image index, and a motion vector. The motion vector can be an integer pixel or a fractional pixel. If the motion vector is a fractional pixel, interpolation filtering needs to be used in the reference image to obtain the required fractional pixel block. Here, the integer pixel or fractional pixel block found in the reference image according to the motion vector is called a reference block. Some technologies directly use the reference block as the prediction block, and some technologies generate the prediction block by processing the reference block. Generating the prediction block by processing the reference block can also be understood as using the reference block as the prediction block and then processing the prediction block to generate a new prediction block.
[0037] The intra prediction unit 112 only refers to the information of the same image to predict the pixel information in the current code image block, and is used to eliminate spatial redundancy.
[0038] There are many intra prediction modes. Taking the H series of international digital video coding standards as an example, the H.264 / AVC standard has 8 angle prediction modes and 1 non-angle prediction mode, and the H.265 / HEVC is extended to 33 angle prediction modes and 2 non-angle prediction modes. The intra prediction mode (IPM) used by HEVC has a planar mode, a DC mode and 33 angle modes, a total of 35 prediction modes. The intra mode used by VVC has a planar mode, a DC mode and 65 angle modes, a total of 67 prediction modes.
[0039] It should be noted that with the increase of angle modes, the intra prediction will be more accurate and more in line with the needs of the development of high-definition and ultra-high-definition digital video.
[0040] Residual unit 120 can generate a residual block for a CU based on the pixel block of the CU and the prediction block of the PUs of the CU. For example, residual unit 120 can generate a residual block for a CU such that each sample in the residual block has a value equal to a difference between a sample in the pixel block of the CU and a corresponding sample in the prediction block of the PUs of the CU.
[0041] Transform / quantization unit 130 can quantize the transform coefficients. Transform / quantization unit 130 can quantize the transform coefficients associated with a TU of a CU based on a quantization parameter (QP) value associated with the CU. Video encoder 100 can adjust the degree of quantization applied to the transform coefficients associated with a CU by adjusting the QP value associated with the CU.
[0042] Inverse transform / quantization unit 140 can apply inverse quantization and inverse transformation, respectively, to the quantized transform coefficients to reconstruct a residual block from the quantized transform coefficients.
[0043] Reconstruction unit 150 can add samples of the reconstructed residual block to corresponding samples of one or more prediction blocks generated by prediction unit 110 to produce a reconstructed image block associated with a TU. By reconstructing the sample blocks of each TU of a CU in this way, video encoder 100 can reconstruct the pixel block of the CU.
[0044] Loop filter unit 160 can be used to process the pixels after inverse transform and inverse quantization, to compensate for distortion information, to provide a better reference for subsequent encoding pixels, for example, a deblocking filter operation can be performed to reduce blocking artifacts associated with the pixel block of the CU.
[0045] In some embodiments, loop filter unit 160 includes a deblocking filter unit to remove blocking artifacts, a sample adaptive offset (SAO) unit to remove ringing artifacts, and an adaptive loop filter (ALF) unit to reduce reconstruction error.
[0046] Decoded picture buffer 170 can store reconstructed pixel blocks. Inter prediction unit 111 can use reference pictures containing reconstructed pixel blocks to perform inter prediction for PUs of other pictures. In addition, intra prediction unit 112 can use reconstructed pixel blocks in decoded picture buffer 170 to perform intra prediction for other PUs in the same picture as the CU.
[0047] Entropy encoding unit 180 can receive quantized transform coefficients from transform / quantization unit 130. Entropy encoding unit 180 can perform one or more entropy encoding operations on the quantized transform coefficients to generate entropy encoded data.
[0048] FIG. 2 is a schematic block diagram of a video decoder according to an embodiment of the present application.
[0049] As shown in FIG. 2, video decoder 200 includes an entropy decoding unit 210, a prediction unit 220, an inverse quantization / unit conversion unit 230, a reconstruction unit 240, a loop filtering unit 250, and a decoded picture buffer 260. It is noted that video decoder 200 can include more, less, or different functional components.
[0050] Video decoder 200 can receive a bitstream. Entropy decoding unit 210 can parse the bitstream to extract syntax elements from the bitstream. As part of parsing the bitstream, entropy decoding unit 210 can entropy decode syntax elements in the bitstream. Prediction unit 220, inverse quantization / unit conversion unit 230, reconstruction unit 240, and loop filtering unit 250 can decode video data according to the syntax elements extracted from the bitstream, i.e., produce decoded video data.
[0051] In some embodiments, prediction unit 220 includes an intra-prediction unit 222 and an inter-prediction unit 221.
[0052] Intra-prediction unit 222 can perform intra-prediction to generate a prediction block for a PU. Intra-prediction unit 222 can use an intra-prediction mode to generate the prediction block for the PU based on blocks of pixels of spatially neighboring PUs. Intra-prediction unit 222 can also determine the intra-prediction mode for the PU according to one or more syntax elements parsed from the bitstream.
[0053] Inter-prediction unit 221 can construct a first reference picture list (List 0) and a second reference picture list (List 1) according to syntax elements parsed from the bitstream. In addition, if the PU is coded using inter-prediction, entropy decoding unit 210 can parse motion information for the PU. Inter-prediction unit 221 can determine one or more reference blocks for the PU according to the motion information for the PU. Inter-prediction unit 221 can generate the prediction block for the PU according to the one or more reference blocks for the PU.
[0054] Inverse quantization / unit conversion unit 230 can inverse quantize (i.e., de-quantize) transform coefficients associated with a TU. Inverse quantization / unit conversion unit 230 can determine a degree of quantization using a QP value associated with a CU of the TU.
[0055] After inverse quantizing the transform coefficients, inverse quantization / unit conversion unit 230 can apply one or more inverse transforms to the inverse quantized transform coefficients in order to produce a residual block associated with the TU.
[0056] Reconstruction unit 240 reconstructs a pixel block of a CU using a residual block associated with a TU of the CU and a prediction block for a PU of the CU. For example, reconstruction unit 240 can add samples of the residual block to corresponding samples of the prediction block to reconstruct the pixel block of the CU, resulting in a reconstructed image block.
[0057] The in-loop filter unit 250 can perform a deblocking filtering operation to reduce blocking artifacts of the pixel blocks associated with the CU.
[0058] The video decoder 200 can store the reconstructed pictures of CUs in a decoded picture buffer 260. The video decoder 200 can use the reconstructed pictures in the decoded picture buffer 260 as reference pictures for subsequent prediction, or transmit the reconstructed pictures to a display device for presentation.
[0059] The basic procedure of video coding is as follows: at the encoding side, a picture is divided into blocks, for a current block, the prediction unit 110 uses intra prediction or inter prediction to generate a prediction block of the current block. The residual unit 120 can calculate a residual block based on the prediction block and the original block of the current block, i.e., the difference between the prediction block and the original block of the current block, which can also be referred to as residual information. The residual block can be processed by the transform / quantization unit 130, which can remove information insensitive to human eyes to eliminate visual redundancy. Optionally, the residual block before being processed by the transform / quantization unit 130 can be referred to as a time-domain residual block, and the residual block after being processed by the transform / quantization unit 130 can be referred to as a frequency-domain residual block or a frequency-domain residual block. The entropy encoding unit 180 receives the quantized transform coefficients output by the transform / quantization unit 130 and entropy encodes the quantized transform coefficients to output a bitstream. For example, the entropy encoding unit 180 can eliminate character redundancy according to a target context model and probability information of a binary code stream.
[0060] At the decoding side, the entropy decoding unit 210 can parse the bitstream to obtain prediction information, a quantized coefficient matrix, etc. of the current block, and the prediction unit 220 uses intra prediction or inter prediction based on the prediction information to generate a prediction block of the current block. The inverse quantization / inverse transform unit 230 uses the quantized coefficient matrix obtained from the bitstream to perform inverse quantization and inverse transform on the quantized coefficient matrix to obtain a residual block. The reconstruction unit 240 adds the prediction block and the residual block to obtain a reconstructed block. The reconstructed blocks constitute a reconstructed picture, and the in-loop filter unit 250 performs in-loop filtering on the reconstructed picture based on the picture or based on the block to obtain a decoded picture. The encoding side also needs to perform similar operations to obtain a decoded picture. The decoded picture can also be referred to as a reconstructed picture, and the reconstructed picture can be used as a reference picture for subsequent inter prediction.
[0061] It should be noted that the block division information determined at the encoding side, as well as the mode information or parameter information of prediction, transform, quantization, entropy encoding, in-loop filtering, etc. are carried in the bitstream when necessary. The decoding side determines the same block division information, prediction, transform, quantization, entropy encoding, in-loop filtering, etc. mode information or parameter information by analyzing the bitstream and based on the existing information, so as to ensure that the decoded picture obtained at the encoding side is the same as the decoded picture obtained at the decoding side.
[0062] The above is the basic process of the video codec under the block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process can be optimized. The present application is applicable to the basic process of the video codec under the block-based hybrid coding framework, but is not limited to the framework and process.
[0063] With the development of technology, the application of light field video is more and more widely. Light field video is an image or video captured by a multi-view camera or a multi-view camera array, which is currently studied by the MPEG LVC working group. Light field video can be collected by a light field camera. Unlike the general camera imaging model, the light field camera adds a set of microlens arrays in front of the imaging plane, so that the light rays of the same point on the object plane can be captured by multiple microlenses at the same time. Using this set of microlens arrays to collect the light rays of the same point is equivalent to taking the same point from multiple angles at the same time. Because of the unique microlens array design inside the light field camera, the image frames of the light field video will present a series of closely arranged macro-pixels (or micro-images), as shown in FIG. 3. The circular pattern in FIG. 3 is a macro-pixel, and in the example of FIG. 3, the macro-pixels are closely arranged in a hexagonal template.
[0064] FIG. 4 shows an example of a complete processing flow of light field video. The codec in FIG. 4 can use traditional video coding tools (such as advanced video coding (AVC), high efficiency video coding (HEVC), or versatile video coding (VVC), etc.). The architecture of the traditional video coding tool can refer to the description of FIG. 1 and FIG. 2.
[0065] Light field video often has a huge amount of data, which brings challenges to its transmission and compression. Therefore, video coding technology for light field video is particularly important. The coding technology adopted by the traditional video codec may not take into account the characteristics of light field video, so the coding performance needs to be improved. For example, according to the foregoing description, the video codec usually has a motion compensation prediction link to reduce the temporal redundancy of the video sequence. The general process of motion compensation prediction is as follows. At the encoding end, for a to-be-encoded block of a current image, a reference block is obtained on another reference image by using motion estimation and other technologies (the relative position of the reference block and the current block is represented by a motion vector); then, the prediction value of the current block is obtained by using the reference block; finally, the residual value between the prediction value and the actual value of the current block and related information such as the motion vector are written into a code stream. At the decoding end, the motion vector and the residual value corresponding to the current block and other information are first extracted from the code stream; then, the reference block is obtained in the corresponding reference image according to the motion vector, and the prediction value of the current block is predicted by using the reference block; finally, the residual value and the prediction value are added to obtain the reconstruction value of the current block. Since the motion compensation technology can effectively reduce the temporal redundancy and reduce the transmission bits, it is widely used in many video coding standard technologies. The above motion compensation prediction technology can have good performance gain for general video, but the motion compensation prediction technology does not take into account the characteristics of light field video, so there is still room for improvement for light field video compression.
[0066] To solve the above problems, the decoding method provided by the embodiments of the present application is described in detail below.
[0067] FIG. 5 is a flowchart of a decoding method provided by an embodiment of the present application. The method shown in FIG. 5 can be referred to as an image decoding method, or a video decoding method. The decoding method shown in FIG. 5 can be applied to a decoder.
[0068] Referring to FIG. 5, in step S510, a code stream is parsed to determine a reference image of a current block and a first reference block in the reference image.
[0069] In some implementations, the code stream can be parsed to determine the index (refIdx) and the motion vector of the reference image of the current block. Then, the reference image of the current block can be determined according to the index of the reference image. Next, the first reference block can be determined according to the reference image and the motion vector.
[0070] For example, referring to FIG. 6, the reference block pointed to by the motion vector can be determined as the first reference block according to the position of the current block and the motion vector in the manner shown in FIG. 6.
[0071] At step S520, one or more second reference block sets are determined according to the first reference block (e.g., the position of the first reference block). The second reference block set mentioned herein can include one or more second reference blocks. The second reference block can be located in the same reference image as the first reference block, or can be located in different reference images.
[0072] Further, in some implementations, the second reference block and the first reference block are located in different macro-pixels (or micro-pixels) in the same reference image. For example, the first reference block is located in a first macro-pixel in the reference image, and the second reference block is located in a second macro-pixel in the reference image. Since the image content in different macro-pixels has correlation, using the reference block in the second macro-pixel to calculate the prediction value is equivalent to considering the spatial correlation of the light field image, which is beneficial to improve the prediction accuracy of the light field video.
[0073] As an example, the first macro-pixel and the second macro-pixel are spatially adjacent macro-pixels in the same reference image. Taking FIG. 7 as an example, the macro-pixel where the reference block 0 is located is the first macro-pixel, and the macro-pixel where any one of the reference blocks 1 to 6 is located is spatially adjacent to the first macro-pixel, and can be used as the second macro-pixel. Spatially adjacent macro-pixels have high similarity, and therefore, prediction based on the reference block in the spatially adjacent macro-pixel can make the prediction value more approximate to the actual value, thereby improving the prediction accuracy.
[0074] The determination manner of the second macro-pixel is not specifically limited in the embodiments of the present application. For example, the second macro-pixel can be determined based on the first macro-pixel and the first parameter. Alternatively, the second macro-pixel having a spatially adjacent relationship with the first macro-pixel can be determined based on the first macro-pixel and the first parameter. The first parameter mentioned herein can be used to indicate the diameter of the micro-lens (used to collect the macro-pixel) or the diameter of the macro-pixel (for example, the diameter of the macro-pixel can refer to the distance between the opposite sides of the hexagon in FIG. 7). For example, the distance between the other macro-pixel and the center of the first macro-pixel can be determined, and the macro-pixel having a distance less than 2 times the diameter of the micro-lens from the center of the first macro-pixel can be determined as the second macro-pixel.
[0075] As another example, the first macro-pixel and the second macro-pixel can be spatially non-adjacent macro-pixels. For example, the second macro-pixel can be a macro-pixel in the near-neighbor region (which can include spatially adjacent positions and / or spatially non-adjacent positions) of the first macro-pixel, and the size of the near-neighbor region can be determined in a pre-defined manner or by a parameter in the code stream.
[0076] As mentioned above, the first reference block is located in the first macro-pixel, and the second reference block is located in the second macro-pixel. The position of the second reference block in the second macro-pixel can be determined based on the position of the first reference block in the first macro-pixel. For example, the relative position of the second reference block in the second macro-pixel is the same as the relative position of the first reference block in the first macro-pixel. In other words, the second reference block is the macro-pixel co-located block of the first reference block. For example, the macro-pixel co-located block of the first reference block can be found from the neighboring macro-pixels of the first macro-pixel (the macro-pixel in which the first reference block is located) based on the geometric arrangement relationship between the macro-pixels. Referring to FIG. 8, the positions of the macro-pixel co-located blocks 1 to 6 in the neighboring macro-pixels can be determined based on the diameter d of the macro-pixel or the micro-lens. Taking the coordinates of the top-left corner of the first reference block as (x0, y0) for example, the coordinates of the top-left corner of the macro-pixel co-located block 2 can be determined based on the following formula:
[0077] Of course, in some other implementations, the relative position of the second reference block in the second macro-pixel can also be different from the relative position of the first reference block in the first macro-pixel. For example, the second reference block can be determined based on the position of a third reference block. The relative position of the third reference block in the second macro-pixel is the same as the relative position of the first reference block in the first macro-pixel. In other words, the third reference block is the macro-pixel co-located block of the first reference block in the second macro-pixel.
[0078] As an example, the second reference block can be determined based on the correlation between the reference blocks in a first region and the first reference block. The first region is a region in the second macro-pixel, and the first region is determined based on the position of the macro-pixel co-located block in the second macro-pixel (for example, the first region can be a region with a preset step length as the radius and with the macro-pixel co-located block as the center). As a more specific example, the macro-pixel co-located block of the first reference block can be found in the second macro-pixel first, and then a search is performed in a certain region around the macro-pixel co-located block to find the co-located block with the highest correlation with the first reference block as the second reference block.
[0079] Alternatively, the position of the second reference block in the second macro-pixel can also be determined based on the position of the third reference block and a first offset value. As an example, the macro-pixel co-located block of the first reference block (i.e., the third reference block) can be found in the second macro-pixel first, and then an image block with a certain offset relationship (the offset relationship is defined by the first offset value) with the macro-pixel co-located block is found as the second reference block.
[0080] Referring back to FIG. 5, in step S530, the prediction value of the current block is determined based on the first reference block and the second reference block set. For example, the first reference block and the second reference block in the second reference block set can be weighted. In other words, the prediction value of the current block can be the weighted sum of the first reference block and the second reference block.
[0081] The embodiments of the present application do not limit the determination manner of the weight coefficients of the first reference block and the second reference block, and the following gives several possible determination manners.
[0082] In some implementations, the weight coefficients of the first reference block and / or the second reference block are obtained from the code stream. Carrying the weight coefficients in the code stream can simplify the implementation of the decoding end.
[0083] In some other implementations, the residual values of the weight coefficients of the first reference block and / or the second reference block are obtained from the code stream. Carrying the residual values of the weight coefficients in the code stream can save the code stream. The residual values of the weight coefficients can be determined based on the weight coefficients of the first reference block and / or the second reference block and the reference values of the weight coefficients. The reference values of the weight coefficients can be predefined values or the weight coefficients used in the inter prediction of a decoded block.
[0084] In some other implementations, the weight coefficients of the first reference block and / or the second reference block are determined based on the weight coefficients used in the inter prediction of a decoded block. For example, the weight coefficients of the first reference block and / or the second reference block are the same as the weight coefficients used in the inter prediction of the decoded block. That is, the current block and the decoded block can share the weight coefficients, thereby saving the code stream. The decoded block mentioned here can be a decoded block decoded before the current block, such as the previous block of the current block.
[0085] In some other implementations, the values of the weight coefficients of the first reference block and / or the second reference block can be predefined values. For example, the values of the weight coefficients of the first reference block and the second reference block are both 0.5. In the implementation of using the predefined values, the weight coefficients do not need to be transmitted through the code stream, thereby reducing the size of the code stream.
[0086] In some other implementations, the weight coefficients of the second reference block can be determined based on the similarity (which can be based on the sum of absolute differences (SAD), the sum of absolute transform difference (SATD), etc.) between the second reference block and the first reference block. For example, the reference block with high similarity to the first reference block can be set with a larger weight, and the reference block with low similarity to the first reference block can be set with a smaller weight. For another example, the reference block with close distance to the first reference block can be set with a larger weight, and the reference block with far distance to the first reference block can be set with a smaller weight. Determining the weight coefficients of the second reference block based on the similarity or distance between the second reference block and the first reference block can on one hand reduce the size of the code stream by not transmitting the weight coefficients through the code stream, and on the other hand improve the reliability of the weight coefficients, thereby improving the prediction accuracy.
[0087] At step S540, a reconstructed block of the current block is determined according to the prediction value of the current block. For example, the bitstream can be parsed to determine the residual value of the current block. Then, the reconstructed block of the current block can be determined according to the residual value of the current block and the prediction value of the current block.
[0088] In some implementations, the method of FIG. 5 further includes parsing the bitstream to determine first identification information. The first identification information is used to indicate or determine whether the prediction value of the current block is determined based on the second reference block set. For example, if the first identification information has a first value (e.g., 1 or true), it indicates that the prediction value of the current block is determined based on the second reference block set (and the first reference block); if the first identification information has a second value (e.g., 0 or false), it indicates that the prediction value of the current block is determined based on only the first reference block.
[0089] The foregoing describes the decoding method provided by the embodiments of the present application in detail in combination with FIG. 5. The following describes the encoding method provided by the embodiments of the present application in detail in combination with FIG. 9.
[0090] FIG. 9 is a flowchart of an encoding method provided by the embodiments of the present application. The method shown in FIG. 9 can be referred to as an image encoding method, or a video encoding method. The encoding method shown in FIG. 9 can be applied to an encoder.
[0091] Referring to FIG. 9, at step S910, a reference image of a current block and a first reference block in the reference image are determined. For example, the first reference block can be determined in the reference image according to a conventional motion estimation manner, or can be determined according to a motion estimation manner proposed for light field video by related technologies.
[0092] In some implementations, an index (refldx) and a motion vector of the reference image of the current block can be written into the bitstream. The positional relationship between the current block and the first reference block can be represented based on the motion vector, as shown in FIG. 6.
[0093] At step S920, one or more second reference block sets are determined according to the first reference block (e.g., the position of the first reference block). The second reference block set mentioned herein can include one or more second reference blocks. The second reference block and the first reference block can be located in the same reference image, or can be located in different reference images.
[0094] Further, in some implementations, the second reference block and the first reference block are located in different macro-pixels (or micro-images) in the same reference picture. For example, the first reference block is located in a first macro-pixel in a reference picture, and the second reference block is located in a second macro-pixel in the reference picture. Since the image content in different macro-pixels has correlation, using the reference block in the second macro-pixel to calculate the prediction value is equivalent to considering the spatial correlation of the light field image, which is beneficial to improve the prediction accuracy of the light field video.
[0095] As an example, the first macro-pixel and the second macro-pixel are spatially adjacent macro-pixels in the same reference picture. Taking FIG. 7 as an example, the macro-pixel where the reference block 0 is located is the first macro-pixel, and the macro-pixel where any one of the reference blocks 1 to 6 is located is spatially adjacent to the first macro-pixel, and can be used as the second macro-pixel. The spatially adjacent macro-pixels have high similarity, and therefore, the prediction based on the reference block in the spatially adjacent macro-pixel can make the prediction value more approximate to the actual value, thereby improving the prediction accuracy.
[0096] The determination manner of the second macro-pixel is not specifically limited in the embodiments of the present application. For example, the second macro-pixel can be determined based on the first macro-pixel and the first parameter. Or, the second macro-pixel having a spatially adjacent relationship with the first macro-pixel can be determined based on the first macro-pixel and the first parameter. The first parameter mentioned here can be used to indicate the diameter of the micro-lens (used to collect the macro-pixel) or the diameter of the macro-pixel (for example, the diameter of the macro-pixel can be represented by the distance between the mutually parallel opposite sides of the hexagon shown in FIG. 7). For example, the distance between the other macro-pixel and the center of the first macro-pixel can be determined, and the macro-pixel having a distance less than 2 times the diameter of the micro-lens from the center of the first macro-pixel can be determined as the second macro-pixel.
[0097] As another example, the first macro-pixel and the second macro-pixel can be spatially non-adjacent macro-pixels. For example, the second macro-pixel can be a macro-pixel in the near-neighbor region (which can include spatially adjacent positions and / or spatially non-adjacent positions) of the first macro-pixel, and the size of the near-neighbor region can be determined in a pre-defined manner or by a parameter in the code stream.
[0098] As mentioned above, the first reference block is located in the first macro-pixel, and the second reference block is located in the second macro-pixel. The position of the second reference block in the second macro-pixel can be determined based on the position of the first reference block in the first macro-pixel. For example, the relative position of the second reference block in the second macro-pixel is the same as the relative position of the first reference block in the first macro-pixel. In other words, the second reference block is the macro-pixel co-located block of the first reference block. For example, the macro-pixel co-located block of the first reference block can be found from the neighboring macro-pixels of the first macro-pixel (the macro-pixel in which the first reference block is located) based on the geometric arrangement relationship between the macro-pixels. Referring to FIG. 8, the positions of the macro-pixel co-located blocks 1 to 6 in the neighboring macro-pixels can be determined based on the diameter d of the macro-pixel or the micro-lens. Taking the coordinates of the top-left corner of the first reference block as (x0, y0) for example, the coordinates of the top-left corner of the macro-pixel co-located block 2 can be determined based on the following formula:
[0099] Of course, in some other implementations, the relative position of the second reference block in the second macro-pixel can also be different from the relative position of the first reference block in the first macro-pixel. For example, the second reference block can be determined based on the position of a third reference block. The relative position of the third reference block in the second macro-pixel is the same as the relative position of the first reference block in the first macro-pixel. In other words, the third reference block is the macro-pixel co-located block of the first reference block in the second macro-pixel.
[0100] As an example, the second reference block can be determined based on the correlation between the reference blocks in a first region and the first reference block. The first region is a region in the second macro-pixel, and the first region is determined based on the position of the macro-pixel co-located block in the second macro-pixel (for example, the first region can be a region with a preset step length as the radius and with the macro-pixel co-located block as the center). As a more specific example, the macro-pixel co-located block of the first reference block can be found in the second macro-pixel first, and then a search is performed in a certain region around the macro-pixel co-located block to find the co-located block with the highest correlation with the first reference block as the second reference block.
[0101] Alternatively, the position of the second reference block in the second macro-pixel can also be determined based on the position of the third reference block and a first offset value. For example, the macro-pixel co-located block of the first reference block (i.e., the third reference block) can be found in the second macro-pixel first, and then an image block with a certain offset relationship (the offset relationship is defined by the first offset value) with the macro-pixel co-located block is found as the second reference block.
[0102] Referring back to FIG. 9, in step S930, the prediction value of the current block is determined according to the first reference block and the second reference block set. For example, the first reference block and the second reference block in the second reference block set can be weighted. In other words, the prediction value of the current block can be the weighted sum of the first reference block and the second reference block.
[0103] The embodiment of the present application does not specifically limit the method for determining the weight coefficients of the first reference block and the second reference block. Several possible determination methods are given below.
[0104] In some implementations, a weight coefficient of at least one of the first reference block and the second reference block is determined based on a correlation between the at least one reference block and an original block of the current block. The correlation mentioned herein can be determined based on SAD, SATD, or other correlation metrics.
[0105] For example, if the number of second reference blocks is 6, the prediction value of the current block can be determined based on the following formula: Where, is the predicted value of the current block, X0 is the value of the first reference block, X i is the value of each second reference block, w0 is the weight coefficient corresponding to the first reference block, w i is the weight coefficient corresponding to each second reference block, i is a positive integer, and the value of i satisfies 1≤i≤6. Calculate each block X i The correlation with the original value Y of the current block is calculated. Then, the weight coefficient corresponding to each block is obtained based on the correlation, so that the block with greater correlation has a greater weight. As an example, if the correlation metric value is SAD, the calculated weight coefficient is
[0106] Furthermore, the weight coefficients of the first reference block and / or the second reference block may be written into the bitstream.
[0107] In other implementations, the residual values of the weight coefficients of the first reference block and / or the second reference block can be written into the bitstream. Carrying the residual values of the weight coefficients in the bitstream can save bitstream. The residual values of the weight coefficients can be determined based on the weight coefficients of the first reference block and / or the second reference block and a reference value of the weight coefficients. The reference value of the weight coefficients can be a predefined value or a weight coefficient used by a coded block in inter-frame prediction.
[0108] In other implementations, the weight coefficients of the first reference block and / or the second reference block are determined based on the weight coefficients used by the coded block during inter-frame prediction. For example, the weight coefficients of the first reference block and / or the second reference block are the same as the weight coefficients used by the coded block during inter-frame prediction. In other words, the current block and the coded block can share the weight coefficients, thereby saving bitrate. The coded block mentioned here can be a coded block that was coded before the current block, such as the previous block of the current block.
[0109] In some implementations, the weight coefficient of the first reference block and / or the weight coefficient of the second reference block can be a predefined value. For example, the weight coefficient of the first reference block and the weight coefficient of the second reference block can both be 0.5. In the implementation of the predefined value, the weight coefficient does not need to be transmitted through the bitstream, thereby reducing the size of the bitstream.
[0110] In some implementations, the weight coefficient of the second reference block can be determined based on the similarity (which can be based on SAD, SATD, etc.) between the second reference block and the first reference block. For example, a reference block with a higher similarity to the first reference block can be assigned a larger weight, and a reference block with a lower similarity to the first reference block can be assigned a smaller weight. For another example, a reference block closer to the first reference block can be assigned a larger weight, and a reference block farther from the first reference block can be assigned a smaller weight. Determining the weight coefficient of the second reference block based on the similarity or distance between the second reference block and the first reference block, on the one hand, does not need to transmit the weight coefficient through the bitstream, thereby reducing the size of the bitstream, and on the other hand, can improve the reliability of the weight coefficient, thereby improving the prediction accuracy.
[0111] At step S940, a residual value of the current block is determined according to the prediction value of the current block. For example, the residual value of the current block can be determined according to the original value of the current block and the prediction value of the current block.
[0112] In some implementations, the method of FIG. 9 further includes writing first identification information into the bitstream. The first identification information is used to indicate or determine whether the prediction value of the current block is determined based on the second reference block set. For example, if the value of the first identification information is a first value (such as 1 or true), it indicates that the prediction value of the current block is determined based on the second reference block set (and the first reference block); if the value of the first identification information is a second value (such as 0 or false), it indicates that the prediction value of the current block is determined based only on the first reference block.
[0113] It should be noted that the macro-pixel mentioned in various embodiments of the present application can also be referred to as a micro-image. The macro-pixel can be a hexagon shape as mentioned above, or other shapes such as a square.
[0114] The embodiments of the present application will be described in more detail below with specific examples. It should be noted that the following examples are only to help those skilled in the art to understand the embodiments of the present application, and are not intended to limit the embodiments of the present application to the specific values or specific scenarios exemplified. Those skilled in the art can obviously make various equivalent modifications or changes based on the examples given, and such modifications or changes also fall within the scope of the embodiments of the present application.
[0115] The present example proposes an improved scheme for motion-compensated prediction for light field video coding, and the specific process is as follows.
[0116] At the encoder end, based on the initial current block position, a corresponding first reference block is obtained from the reference image, as shown in Figure 6. There are various ways to search for and obtain the first reference block. For example, the first reference block can be determined using a traditional motion estimation algorithm or an improved motion estimation algorithm proposed for light field video.
[0117] Then, based on the position of the first reference block, the position of the corresponding macropixel co-located block can be obtained, as shown in Figure 7 (the current block position and the macropixel co-located block position are in the same relative position within the macropixel to which they belong). Figure 7 shows a hexagonal macropixel; this example is also applicable to other shapes (such as a square). In addition, Figure 7 selects the six macropixels closest to the macropixel where the first reference block is located. In other examples, other numbers or positions of macropixels may also be selected.
[0118] Combined with the reference block and macro pixel co-located block shown in Figure 7, it can be based on the formula The weighted prediction is used to get the predicted value of the current block. is the predicted value of the current block, X0 is the value of the first reference block, X i is the value of each macro pixel co-located block of the first reference block, w0 is the weight coefficient corresponding to the first reference block, w i is the weight coefficient corresponding to each macro pixel co-located block, 1≤i≤6. This example does not restrict the value range of the weight coefficient and the method of obtaining the weight coefficient;
[0119] A possible way to obtain the weight coefficient is to calculate each block X i Correlation with the actual value Y of the current block. For example, the sum of absolute differences (SAD) can be used i Calculate each block X i The correlation with the actual value Y of the current block can then be used to determine the weights of each block, so that blocks with greater correlations have greater weights. For example, if the correlation metric used is SAD, the weight coefficient calculated can be
[0120] Then, calculate the current block prediction value The prediction residual ΔY between the actual value Y and the prediction residual ΔY, motion vector and related flag bits are written into the bitstream. The related flag bits mentioned here can be used to mark whether the motion compensation prediction method proposed in this example is used. i , can be written directly into the code stream or not. i If the bitstream is written, the corresponding weight coefficient can be obtained by some method at the decoding end. For example, the weight coefficient can be set to a fixed value.
[0121] At the decoding end, the motion vector is read from the bitstream, and the corresponding first reference block is obtained. Then, the related flag bit is read. If the flag bit indicates that the motion compensation method proposed in this example is used, the corresponding macro-pixel homonymic block is obtained from the surrounding macro-pixels of the reference block. If the weight coefficient is included in the bitstream, the weight coefficient is read out, and the predicted value of the current block is obtained according to the formula mentioned above. If the weight coefficient is not included in the bitstream, the corresponding weight coefficient can be obtained in some way. For example, the weight coefficient can be set to a preset fixed value. Finally, the prediction residual is read from the bitstream, and the reconstruction value of the current block is obtained by adding the prediction residual to the predicted value.
[0122] The space correlation of the light field image is considered in this example, which is beneficial to improve the prediction accuracy of the light field video coding, thereby reducing the bit consumption of the transmission of the prediction residual.
[0123] The method embodiments of the present application are described in detail above in combination with FIG. 1 to FIG. 9, and the device embodiments of the present application are described in detail below in combination with FIG. 10 to FIG. 13. It should be understood that the description of the method embodiments and the description of the device embodiments correspond to each other, and therefore, the parts not described in detail can be referred to the method embodiments.
[0124] FIG. 10 is a structural schematic diagram of a decoder provided by an embodiment of the present application. The decoder 1000 in FIG. 10 includes a decoding module 1010, a first determining module 1020, a second determining module 1030, and a third determining module 1040. The decoding module 1010 is configured to parse a bitstream, determine a reference image of a current block and a first reference block in the reference image. The first determining module 1020 is configured to determine one or more second reference block sets according to the first reference block. The second determining module 1030 is configured to determine a predicted value of the current block according to the first reference block and the second reference block set. The third determining module 1040 is configured to determine a reconstructed block of the current block according to the predicted value of the current block.
[0125] In some implementations, the decoding module 1010 is configured to parse a bitstream, determine an index and a motion vector of the reference image, determine the reference image according to the index of the reference image, and determine the first reference block according to the reference image and the motion vector.
[0126] In some implementations, the second reference block set includes one or more second reference blocks, and the second reference blocks are located in the reference image.
[0127] In some implementations, the first reference block and the second reference block belong to a first macro-pixel and a second macro-pixel in the reference image, respectively.
[0128] In some embodiments, the first macro-pixel and the second macro-pixel are spatially adjacent macro-pixels.
[0129] In some embodiments, the second macro-pixel is determined based on the first macro-pixel and a first parameter, the first parameter being used to indicate a diameter of a macro-pixel or a diameter of a micro-lens.
[0130] In some embodiments, a position of the second reference block in the second macro-pixel is determined based on a position of the first reference block in the first macro-pixel.
[0131] In some embodiments, a relative position of the second reference block in the second macro-pixel is the same as a relative position of the first reference block in the first macro-pixel.
[0132] In some embodiments, the second reference block is determined based on a third reference block, a relative position of the third reference block in the second macro-pixel being the same as a relative position of the first reference block in the first macro-pixel.
[0133] In some embodiments, the second reference block is determined based on a correlation of a reference block in a first region with the first reference block, the first region being a region within the second macro-pixel, and the first region being determined based on a position of the third reference block.
[0134] In some embodiments, a position of the second reference block in the second macro-pixel is determined based on a position of the third reference block and a first offset value.
[0135] In some embodiments, the second reference block set includes one or more second reference blocks, and the prediction value is a weighted sum of the first reference block and the second reference blocks.
[0136] In some embodiments, a weight coefficient of the first reference block and / or the second reference block is obtained from a bitstream.
[0137] In some embodiments, a residual value of a weight coefficient of the first reference block and / or the second reference block is obtained from a bitstream.
[0138] In some embodiments, a weight coefficient of the first reference block and / or the second reference block is determined based on a weight coefficient used by a decoded block in inter prediction.
[0139] In some embodiments, a weight coefficient of the first reference block and / or the second reference block is a predefined value.
[0140] In some embodiments, a weight coefficient of the second reference block is determined based on a similarity and / or a distance between the second reference block and the first reference block.
[0141] In some embodiments, the decoding module 1010 is further configured to parse the bitstream to determine first identification information, the first identification information being used to determine whether the prediction value of the current block is determined based on the second reference block set.
[0142] In some embodiments, the decoding module 1010 is further configured to parse the bitstream to determine a residual value of the current block; and the third determining module 1040 is configured to determine the reconstructed block of the current block according to the residual value of the current block and the prediction value of the current block.
[0143] It can be understood that, in the embodiments of the present application, the "unit" can be a part of circuit, a part of processor, a part of program or software, etc., and of course can be a module, and can also be non-modular. Moreover, the components in the embodiments can be integrated in one processing unit, or can be physically present individually, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function module.
[0144] The integrated unit, if realized in the form of a software function module and not sold or used as an independent product, can be stored in a computer readable storage medium, based on such understanding, the technical solutions of the embodiments can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the methods described in the embodiments. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program codes that can be stored in the medium.
[0145] Therefore, the embodiments of the present application provide a computer readable storage medium applied to the decoder 1000, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the decoding method in any of the foregoing embodiments.
[0146] Based on the above components of the decoder 1000 and the computer readable storage medium, referring to FIG. 11, a specific hardware structure diagram of the decoder 1000 is shown. As shown in FIG. 11, the decoder 1100 can include a communication interface 1110, a memory 1120 and a processor 1130; each component is coupled together through a bus system 1140. It can be understood that the bus system 1140 is used to realize the connection communication between the components. The bus system 1140 includes not only a data bus, but also a power bus, a control bus and a status signal bus. However, in order to clearly illustrate, various buses are marked as the bus system 1140 in FIG. 11. Among them,
[0147] The communication interface 1110 is configured to receive and send signals in the process of transmitting information with other external network elements;
[0148] The memory 1120 is configured to store a computer program;
[0149] The processor 1130 is configured to execute the following steps when running the computer program:
[0150] Parsing a code stream to determine a reference image of a current block and a first reference block in the reference image;
[0151] Determining one or more second reference block sets according to the first reference block;
[0152] Determining a prediction value of the current block according to the first reference block and the second reference block set;
[0153] Determining a reconstructed block of the current block according to the prediction value of the current block.
[0154] It is to be appreciated that the memory 1120 in the embodiments of this application can be volatile, nonvolatile, or a combination of both. By way of example, the nonvolatile memory can include read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), a flash memory, or a combination of these. The volatile memory can include random access memory (RAM), which acts as external cache. By way of example, the RAM is accessible in a number of forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The memory 1120 of the system and method described herein is intended to include, without being limited to, these and any other suitable types of memory.
[0155] The processor 1130 can be an integrated circuit chip powerfully processing signals. In implementation process, each step of the above method can be completed by the integrated logic circuit or the instruction in the form of software in the processor 1130. The processor 1130 described above can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Each method, step and logic block in the embodiments of the present application can be realized or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor to execute, or be executed by a combination of hardware and software modules in the code processor. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register, or other mature storage medium in the art. The storage medium is located in the storage 1120, and the processor 1130 reads the information in the storage 1120 and combines the hardware to complete the steps of the above method.
[0156] It can be understood that the embodiments described in the present application can be realized by hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general purpose processors, controllers, microcontrollers, microprocessors, other electronic units for executing the functions described in the present application or a combination thereof. For software implementation, the technologies described in the present application can be realized by modules (such as processes, functions, etc.) for executing the functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0157] Optionally, as another embodiment, the processor 1130 is further configured to execute the decoding method described in the foregoing embodiments when running the computer program.
[0158] FIG. 12 is a schematic diagram of an encoder according to an embodiment of the present application. As shown in FIG. 12, the encoder 1200 includes a first determining module 1210, a second determining module 1220, a third determining module 1230, and a fourth determining module 1240. The first determining module 1210 is configured to determine a reference picture of a current block and a first reference block in the reference picture. The second determining module 1220 is configured to determine one or more second reference block sets according to the first reference block. The third determining module 1230 is configured to determine a prediction value of the current block according to the first reference block and the second reference block sets. The fourth determining module 1240 is configured to determine a residual value of the current block according to the prediction value of the current block.
[0159] In some embodiments, the encoder 1200 further includes an encoding module configured to write an index of the reference picture and a motion vector used to determine the first reference block from the reference picture into a bitstream.
[0160] In some embodiments, the second reference block set includes one or more second reference blocks, and the second reference blocks are located in the reference picture.
[0161] In some embodiments, the first reference block and the second reference block belong to a first macro-pixel and a second macro-pixel in the reference picture, respectively.
[0162] In some embodiments, the first macro-pixel and the second macro-pixel are spatially adjacent macro-pixels.
[0163] In some embodiments, the second macro-pixel is determined based on the first macro-pixel and a first parameter, the first parameter being used to indicate a diameter of a macro-pixel or a diameter of a micro-lens.
[0164] In some embodiments, a position of the second reference block in the second macro-pixel is determined based on a position of the first reference block in the first macro-pixel.
[0165] In some embodiments, a relative position of the second reference block in the second macro-pixel is the same as a relative position of the first reference block in the first macro-pixel.
[0166] In some embodiments, the second reference block is determined based on a position of a third reference block, the relative position of the third reference block in the second macro-pixel being the same as the relative position of the first reference block in the first macro-pixel.
[0167] In some embodiments, the second reference block is determined based on a correlation of a reference block within a first region to the first reference block, the first region being a region within the second macro-pixel, and the first region being determined based on a position of the third reference block.
[0168] In some embodiments, a position of the second reference block in the second macro-pixel is determined based on the position of the third reference block and a first offset value.
[0169] In some embodiments, the second reference block set includes one or more second reference blocks, and the prediction value is a weighted sum of the first reference block and the second reference blocks.
[0170] In some embodiments, a weight coefficient of the first reference block and / or the second reference block is written into a bitstream.
[0171] In some embodiments, a weight coefficient of at least one of the first reference block and the second reference block is determined based on a correlation of the at least one reference block to an original block of the current block.
[0172] In some embodiments, the correlation of the at least one reference block to the original block of the current block is determined based on a SAD of the at least one reference block to the original block of the current block.
[0173] In some embodiments, the method further includes writing a residual value of the weight coefficient of the first reference block and / or the second reference block into a bitstream.
[0174] In some embodiments, the weight coefficient of the first reference block and / or the second reference block is determined based on a weight coefficient used by a coded block in inter prediction.
[0175] In some embodiments, the weight coefficient of the first reference block and / or the second reference block is a predefined value.
[0176] In some embodiments, the weight coefficient of the second reference block is determined based on a similarity and / or distance between the second reference block and the first reference block.
[0177] In some embodiments, the encoder further includes an encoding module configured to write first identification information into a bitstream, the first identification information being used to determine whether the prediction value of the current block is determined based on the second reference block set.
[0178] It can be understood that, in the embodiments of the present application, the "unit" can be part of a circuit, part of a processor, part of a program or software, etc., and of course can also be a module, and can also be non-modular. Moreover, the components in the embodiments can be integrated in a processing unit, or can be physically present as individual units, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function module.
[0179] The integrated unit, if realized in the form of a software function module and not sold or used as an independent product, can be stored in a computer-readable storage medium, based on this understanding, the technical solutions of the embodiments essentially or the parts that contribute to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product, the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (processor) execute all or part of the steps of the method described in the embodiments. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0180] Therefore, the embodiments of the present application provide a computer-readable storage medium applied to the encoder 1200, the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the encoding method described in any one of the foregoing embodiments.
[0181] Based on the components of the foregoing encoder 1200 and the computer-readable storage medium, referring to FIG. 13, a specific hardware structure schematic diagram of the encoder 1200 is shown. As shown in FIG. 13, the encoder 1300 can include a communication interface 1310, a memory 1320, and a processor 1330; the components are coupled together through a bus system 1340. It can be understood that the bus system 1340 is used to realize the connection communication between the components. The bus system 1340 includes a data bus, a power supply bus, a control bus, and a state signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the bus system 1340 in FIG. 13. Among them,
[0182] The communication interface 1310 is used for receiving and sending signals in the process of transceiving information with other external network elements;
[0183] The memory 1320 is used for storing a computer program;
[0184] a processor 1330, configured to, when the computer program is run:
[0185] determine a reference picture for a current block and a first reference block in the reference picture;
[0186] determine one or more second reference block sets according to the first reference block;
[0187] determine a prediction value of the current block according to the first reference block and the second reference block sets;
[0188] determine a residual value of the current block according to the prediction value of the current block.
[0189] It can be understood that the memory 1320 in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (Read-Only Memory, ROM), a programmable read-only memory (Programmable ROM, PROM), an erasable programmable read-only memory (Erasable PROM, EPROM), an electrically erasable programmable read-only memory (Electrically EPROM, EEPROM) or a flash memory. The volatile memory can be a random access memory (Random Access Memory, RAM) used as an external cache. By way of example, but not by way of limitation, many forms of RAM are available, such as static random access memory (Static RAM, SRAM), dynamic random access memory (Dynamic RAM, DRAM), synchronous dynamic random access memory (Synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (Double Data Rate SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (Enhanced SDRAM, ESDRAM), synchronous link dynamic random access memory (Synchlink DRAM, SLDRAM) and direct memory bus random access memory (Direct Rambus RAM, DRRAM). The memory 1320 of the system and method described in the present application is intended to include, but not limited to, these and any other suitable types of memory.
[0190] The processor 1330 can be an integrated circuit chip including a processing unit that is configured to process signals. In implementation, the steps of the above-described method can be completed by the integrated logic circuit of the processor 1330 or by an instruction in a form of software. The processor 1330 described above can be a general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The methods, steps and logical block diagrams disclosed in the embodiments of the present application can be implemented or executed by the processor 1330. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code executed by the processor, or a combination of hardware and software modules in the processor. The software module can be located in a storage medium such as random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, or other mature storage media in the art. The storage medium is located in the memory 1320, and the processor 1330 reads information in the memory 1320 and combines the hardware to complete the steps of the above-described method.
[0191] It can be understood that the embodiments described in the present application can be realized by hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field-Programmable Gate Arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for executing the functions described in the present application or a combination thereof. For software implementation, the technologies described in the present application can be implemented by modules (for example, procedures, functions, and so on) for performing the functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0192] Alternatively, as another embodiment, the processor 1330 is further configured to execute the encoding method in the foregoing embodiments when running the computer program.
[0193] The embodiment of the present application further provides a computer readable storage medium, which is a non-volatile computer readable storage medium for storing a bitstream, the bitstream can be generated by using an encoding method of an encoder, or the bitstream is decoded by using a decoding method of a decoder, wherein the decoding method can be the decoding method in any one of the foregoing embodiments, and the encoding method can be the encoding method in any one of the foregoing embodiments.
[0194] It should be noted that, in the present application, the terms "comprising", "containing" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0195] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0196] The methods disclosed in the several method embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments.
[0197] The features disclosed in the several product embodiments provided by the present application can be combined arbitrarily without conflict to obtain new product embodiments.
[0198] The features disclosed in the several method or device embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method or device embodiments.
[0199] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A decoding method applied to a decoder, comprising: parsing a bitstream to determine a reference picture of a current block and a first reference block in the reference picture; determining one or more second reference block sets according to the first reference block; determining a prediction value of the current block according to the first reference block and the second reference block sets; determining a reconstructed block of the current block according to the prediction value of the current block.
2. The method of claim 1, wherein, The parsing of the bitstream to determine the reference picture of the current block and the first reference block in the reference picture comprises: parsing the bitstream to determine an index and a motion vector of the reference picture; determining the reference picture according to the index of the reference picture; determining the first reference block according to the reference picture and the motion vector.
3. The method of claim 1, wherein, The second reference block set comprises one or more second reference blocks, and the second reference blocks are located in the reference picture.
4. The method of claim 3, wherein, The first reference block and the second reference block belong to a first macro-pixel and a second macro-pixel in the reference picture respectively.
5. The method of claim 4, wherein, The first macro-pixel and the second macro-pixel are spatially adjacent macro-pixels.
6. The method of claim 5, wherein, The second macro-pixel is determined based on the first macro-pixel and a first parameter, the first parameter being used to indicate a diameter of a macro-pixel or a diameter of a micro-lens.
7. The method of any one of claims 4 to 6, wherein, A position of the second reference block in the second macro-pixel is determined based on a position of the first reference block in the first macro-pixel.
8. The method of any one of claims 4 to 7, wherein, A relative position of the second reference block in the second macro-pixel is the same as a relative position of the first reference block in the first macro-pixel.
9. The method of any one of claims 4 to 7, wherein, The second reference block is determined based on a third reference block, a relative position of the third reference block in the second macro-pixel being the same as the relative position of the first reference block in the first macro-pixel.
10. The method of claim 9, wherein, The second reference block is determined based on a correlation between a reference block in a first region and the first reference block, the first region being a region within the second macro-pixel, and the first region being determined based on a position of the third reference block.
11. The method of claim 9, wherein, A position of the second reference block in the second macro-pixel is determined based on the position of the third reference block and a first offset value.
12. The method of any one of claims 1 to 11, wherein, The second reference block set comprises one or more second reference blocks, and the prediction value is a weighted sum of the first reference block and the second reference blocks.
13. The method of claim 12, wherein, A weight coefficient of the first reference block and / or the second reference block is obtained from the bitstream.
14. The method of claim 12, wherein, A residual value of the weight coefficient of the first reference block and / or the second reference block is obtained from the bitstream.
15. The method of claim 12, wherein, The weight coefficient of the first reference block and / or the second reference block is determined based on a weight coefficient used by a decoded block in inter prediction.
16. The method of claim 12, wherein, The weight coefficient of the first reference block and / or the second reference block takes a predefined value.
17. The method of claim 12, wherein, The weight coefficient of the second reference block is determined based on a similarity and / or a distance between the second reference block and the first reference block.
18. The method of any one of claims 1 to 17, wherein, The method further comprises: parsing the bitstream to determine first identification information, the first identification information being used to determine whether to determine the prediction value of the current block based on the second reference block sets.
19. The method of any one of claims 1 to 18, wherein, The method further comprises: parsing the bitstream to determine a residual value of the current block; The determination of the reconstructed block of the current block according to the prediction value of the current block comprises: determine a reconstructed block of the current block according to the residual value of the current block and the prediction value of the current block.
20. A method of encoding applied to an encoder, comprising: determining a reference picture of a current block and a first reference block in the reference picture; determining one or more second reference block sets according to the first reference block; determining a prediction value of the current block according to the first reference block and the second reference block sets; determining a residual value of the current block according to the prediction value of the current block.
21. The method of claim 20, wherein, The method further comprises: writing an index of the reference picture and a motion vector used to determine the first reference block from the reference picture into a bitstream.
22. The method of claim 20, wherein, The second reference block set comprises one or more second reference blocks, and the second reference blocks are located in the reference picture.
23. The method of claim 22, wherein, The first reference block and the second reference block belong to a first macro-pixel and a second macro-pixel in the reference picture, respectively.
24. The method of claim 23, wherein, The first macro-pixel and the second macro-pixel are spatially adjacent macro-pixels.
25. The method of claim 24, wherein, The second macro-pixel is determined based on the first macro-pixel and a first parameter, the first parameter being used to indicate a diameter of a macro-pixel or a diameter of a micro-lens.
26. The method of any one of claims 23-25, wherein, A position of the second reference block in the second macro-pixel is determined based on a position of the first reference block in the first macro-pixel.
27. The method of any one of claims 23-26, wherein, A relative position of the second reference block in the second macro-pixel is the same as a relative position of the first reference block in the first macro-pixel.
28. The method of any one of claims 23-26, wherein, The second reference block is determined based on a third reference block, a relative position of the third reference block in the second macro-pixel being the same as the relative position of the first reference block in the first macro-pixel.
29. The method of claim 28, wherein, The second reference block is determined based on a correlation between a reference block in a first region and the first reference block, the first region being a region within the second macro-pixel, and the first region being determined based on a position of the third reference block.
30. The method of claim 28, wherein, A position of the second reference block in the second macro-pixel is determined based on a position of the third reference block and a first offset value.
31. The method of any one of claims 20-30, wherein, The second reference block set comprises one or more second reference blocks, and the prediction value is a weighted sum of the first reference block and the second reference blocks.
32. The method of claim 31, wherein, The method further comprises writing a weight coefficient of the first reference block and / or the second reference block into a bitstream.
33. The method of claim 31 or 32, wherein, The weight coefficient of at least one of the first reference block and the second reference block is determined based on a correlation between the at least one reference block and an original block of the current block.
34. The method of claim 33, wherein, The correlation between the at least one reference block and the original block of the current block is determined based on a SAD between the at least one reference block and the original block of the current block.
35. The method of claim 31, wherein, The method further comprises writing a residual value of the weight coefficient of the first reference block and / or the second reference block into a bitstream.
36. The method of claim 31, wherein, The weight coefficient of the first reference block and / or the second reference block is determined based on a weight coefficient used by a coded block in inter prediction.
37. The method of claim 31, wherein, The weight coefficient of the first reference block and / or the second reference block is a predefined value.
38. The method of claim 31, wherein, The weight coefficient of the second reference block is determined based on a similarity and / or a distance between the second reference block and the first reference block.
39. The method of any one of claims 20 to 38, wherein, The method further comprises: write, into a bitstream, first identification information used to determine whether a prediction value of the current block is determined based on the second reference block set.
40. A decoder comprising: a decoding module configured to parse a bitstream, determine a reference picture of a current block and a first reference block in the reference picture; a first determining module configured to determine one or more second reference block sets according to the first reference block; a second determining module configured to determine a prediction value of the current block according to the first reference block and the second reference block sets; a third determining module configured to determine a reconstructed block of the current block according to the prediction value of the current block.
41. An encoder comprising: a first determining module configured to determine a reference picture of a current block and a first reference block in the reference picture; a second determining module configured to determine one or more second reference block sets according to the first reference block; a third determining module configured to determine a prediction value of the current block according to the first reference block and the second reference block sets; a fourth determining module configured to determine a residual value of the current block according to the prediction value of the current block.
42. A decoder, the decoder comprising: a memory configured to store a computer program; a processor configured to execute a method recited in any of claims 1 to 19 when running the computer program.
43. An encoder, the encoder comprising: a memory configured to store a computer program; a processor configured to execute a method recited in any of claims 20 to 39 when running the computer program.
44. A computer readable storage medium, wherein, The computer readable storage medium stores a computer program which, when executed, implements a method recited in any of claims 1 to 39.
45. A non-volatile computer readable storage medium storing a bitstream, the bitstream being generated by using an encoding method of an encoder, or the bitstream being decoded by using a decoding method of a decoder, the decoding method being a method recited in any of claims 1 to 19, the encoding method being a method recited in any of claims 20 to 39.
Citation Information
Patent Citations
Method, device and equipment for acquiring motion information of video image and template constructing method
CN101931803A
Video inter-frame motion estimation method, device and equipment and readable storage medium
CN110267047A
Image component prediction method, encoder, decoder, and storage medium
CN113992916A
Method for modifying a reference block of a reference image, method for encoding or decoding a block of an image by help of a reference block and device therefore and storage medium or signal carrying a block encoded by help of a modified reference block
US20110135211A1
Methods and devices for encoding and decoding a data stream representative of an image sequence
US20200128239A1