Coding method, coder, and electronic device

By correcting motion parameters in motion estimation, the bandwidth and traffic pressure issues of digital video compression under high video definition are resolved, improving coding performance and accuracy.

WO2025010572A9PCT designated stage expired Publication Date: 2025-11-13GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/106447
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-07-07
Publication Date
2025-11-13

AI Technical Summary

Technical Problem

Existing digital video compression technologies still face bandwidth and traffic pressures under the demand for high video resolution, requiring more efficient encoding methods to reduce transmission requirements.

Method used

By determining the first motion parameters of the initial position in the motion estimation of the current block and finding at least one first position within its search area, the motion parameters are corrected to improve accuracy, thereby enhancing coding performance.

Benefits of technology

It improves the accuracy of motion estimation and coding performance, while reducing the bandwidth and traffic pressure on digital video transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023106447_13112025_PF_FP_ABST
    Figure CN2023106447_13112025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a coding method, a coder, and an electronic device. The coding method comprises: performing motion estimation on the current block in the current image, and determining a first motion parameter of the current block, wherein the first motion parameter points to an initial position in a reference image of the current image; determining at least one first position, other than the initial position, in a first search area in an initial macro pixel where the initial position is located; and determining a second motion parameter of a reference block on the basis of the initial position and the at least one first position. The coding method can improve a motion estimation effect and the coding performance of the coder.
Need to check novelty before this filing date? Find Prior Art

Description

Encoding methods, encoders, and electronic devices Technical Field

[0001] This application relates to the field of encoding and decoding technology, and more specifically, to encoding methods, encoders, and electronic devices. Background Technology

[0002] Digital video compression technology mainly compresses massive amounts of digital video data to facilitate transmission and storage.

[0003] With the surge in internet videos and people's increasing demands for video clarity, although existing digital video compression standards can save a lot of video data, there is still a need to pursue better digital video compression technologies to reduce the bandwidth and traffic pressure of digital video transmission.

[0004] Summary of the Invention

[0005] This application provides an encoding method, an encoder, and an electronic device that can improve the motion estimation effect and encoding performance of the encoder.

[0006] Firstly, this application provides an encoding method, including:

[0007] Motion estimation is performed on the current block in the current image to determine the first motion parameter of the current block; the first motion parameter points to the initial position in the reference image of the current image;

[0008] Within the first search area where the initial position is located, at least one first position other than the initial position is determined; the first search area is less than or equal to the initial macro pixel where the initial position is located;

[0009] Based on the initial position and the at least one first position, the second motion parameters of the reference block are determined.

[0010] Secondly, this application provides an encoder, comprising:

[0011] An estimation unit is used to perform motion estimation on a current block in the current image and determine a first motion parameter of the current block; the first motion parameter points to an initial position in a reference image of the current image;

[0012] The first determining unit is configured to determine at least one first position other than the initial position within the first search area where the initial position is located; the first search area is less than or equal to the initial macro pixel where the initial position is located;

[0013] The second determining unit is used to determine the second motion parameters of the reference block based on the initial position and the at least one first position.

[0014] Thirdly, this application provides an encoder, comprising:

[0015] Processor, adapted to implement computer instructions; and,

[0016] A computer-readable storage medium storing computer instructions adapted for loading by a processor and executing the encoded methods of the second aspect or its various implementations mentioned above.

[0017] In one implementation, there are one or more processors and one or more memories.

[0018] In one implementation, the computer-readable storage medium may be integrated with the processor, or the computer-readable storage medium may be disposed separately from the processor.

[0019] Fourthly, this application provides a computer-readable storage medium storing computer instructions that, when read and executed by a processor of a computer device, cause the computer device to perform the encoding method described in the first aspect above.

[0020] Fifthly, this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the encoding method described in the first aspect above.

[0021] Sixthly, this application provides a bitstream generated by the method described in the first aspect above.

[0022] Based on the above technical solutions, the encoding method provided in this application takes the position pointed to by the first motion parameter determined by motion estimation as the initial position, and then determines at least one first position other than the initial position within the first search area within the initial macro pixel where the initial position is located; based on the initial position and the at least one first position, the second motion parameter of the reference block is determined; equivalently, by introducing the at least one first position to correct the first motion parameter and obtaining the corrected second motion parameter, the accuracy of the second motion parameter can be improved, thereby improving the motion estimation effect and encoding performance of the encoder. Attached Figure Description

[0023] Figure 1 is a schematic block diagram of the encoding and decoding system provided in an embodiment of this application.

[0024] Figure 2 is a schematic block diagram of the encoder provided in an embodiment of this application.

[0025] Figure 3 is a schematic structural diagram of the relationship between the coding tree unit and the coding unit provided in the embodiments of this application.

[0026] Figure 4 is a schematic block diagram of the decoder provided in an embodiment of this application.

[0027] Figure 5 is an example of the transmission process of light field video provided in an embodiment of this application.

[0028] Figure 6 is an example of the initial position and co-location provided in the embodiments of this application.

[0029] Figure 7 is an example of a light field image provided in an embodiment of this application.

[0030] Figure 8 is a schematic flowchart of the encoding method provided in the embodiments of this application.

[0031] Figure 9 is an example of the search principle for at least one first position provided by an embodiment of this application.

[0032] Figure 10 is an example of the same site location provided in the embodiments of this application.

[0033] Figure 11 is an example of at least one layer provided in an embodiment of this application.

[0034] Figure 12 is another example of at least one layer provided in the embodiments of this application.

[0035] Figure 13 is an example of an embodiment of this application where the search spacing of the first search region is equal to the search spacing used by the first layer.

[0036] Figure 14 is an example of a first search region having a search spacing smaller than that used by the first layer, provided in an embodiment of this application.

[0037] Figure 15 is a schematic block diagram of the encoder provided in an embodiment of this application.

[0038] Figure 16 is a schematic structural diagram of the electronic device provided in an embodiment of this application. Detailed Implementation

[0039] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0040] It should be noted that the terminology used in the implementation section of this application is only used to explain the specific embodiments of this application and is not intended to limit this application.

[0041] For example, the term "and / or" in this article simply describes the relationship between related objects, indicating that three relationships can exist. For instance, A and / or B can represent: A alone, A and B simultaneously, and B alone. The term "at least one" simply describes the combination relationship of listed objects, indicating that one or more can exist. For instance, at least one of the following: A, B, C can represent the following combinations: A alone, B alone, C alone, A and B simultaneously, A and C simultaneously, B and C simultaneously, and A, B, and C simultaneously. The term "multiple" refers to two or more. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0042] For example, the term "correspondence" can indicate a direct or indirect correspondence between two things, or an association between them, or a relationship of instruction and being instructed, configuration and being configured, etc. The term "instruction" can be direct, indirect, or indicate an association. For example, A instructing B can mean A directly instructs B, for example, B can be obtained through A; it can also mean A indirectly instructs B, for example, A instructs C, B can be obtained through C; or it can mean an association between A and B. The terms "predefined" or "preconfigured" can refer to pre-stored corresponding codes, tables, or other relevant information that can be used for instruction in the device (e.g., including encoders or decoders), or it can refer to something agreed upon by a protocol. "Protocol" can refer to any standard protocol in the field of encoding and decoding, and this application does not limit it. The term "when..." can be interpreted as "if," "when," or "in response to," etc. Similarly, depending on the context, the phrases "if determined" or "if detected (the stated condition or event)" can be interpreted as "when determined," "in response to determined," "when detected (the stated condition or event)," or "in response to detected (the stated condition or event)," and similar descriptions. The terms "first," "second," "third," "fourth," "A," "B," etc., are used to distinguish different objects, not to describe a specific order. The terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. Digital video compression technology primarily compresses massive amounts of digital video data to facilitate transmission and storage.

[0043] The solutions provided in this application can be applied to the field of digital compression technology.

[0044] For example, the solutions provided in the embodiments of this application can be applied to the field of video coding technology.

[0045] The video coding technology field includes, but is not limited to, at least one of the following: image codec, video codec, hardware video codec, dedicated circuit video codec, and real-time video codec. Furthermore, the solutions provided in this application can be incorporated into the following standards: Audio Video Coding Standard (AVS), AVS2, or AVS3. For example, including but not limited to: H.264 / Audio Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), and H.266 / Versatile Video Coding (VVC). Additionally, the solutions provided in this application can be used for lossy compression or lossless compression of images. This lossless compression can be visually lossless compression or mathematically lossless compression.

[0046] Video codec standards can adopt a block-based hybrid coding framework.

[0047] The hybrid coding framework includes modules such as prediction, transform, quantization, entropy coding, and in-loop filtering. The prediction module includes intra-frame prediction and / or inter-frame prediction. Because there is a strong correlation between adjacent pixels within a video frame, intra-frame prediction is used in video encoding and decoding to eliminate spatial redundancy between adjacent pixels. Intra-frame prediction only references information from the same frame to predict pixel information within the current block. Because there is a strong similarity between adjacent frames in a video, inter-frame prediction is used in video encoding and decoding to eliminate temporal redundancy between adjacent frames, thereby improving coding efficiency. Inter-frame prediction includes motion estimation and motion compensation. Inter-frame prediction can reference image information from different frames and use motion estimation to search for the motion vector information that best matches the current block. Transform converts the predicted image block to the frequency domain, redistributing energy. Combined with quantization, information that is insensitive to the human eye can be removed to eliminate visual redundancy. Entropy coding can eliminate character redundancy based on the current context model and the probability information of the binary bitstream.

[0048] The basic process of a video encoder is as follows:

[0049] The encoder first divides a frame of image into blocks; then it predicts the current block in the current image to obtain the predicted block of the current block; next, it subtracts the predicted block from the original block of the current block to obtain the residual block; the residual block is transformed and quantized to obtain the quantization coefficient matrix; then the quantization coefficient matrix is ​​entropy encoded to obtain the output bitstream.

[0050] The basic process of a video decoder is as follows:

[0051] The decoder performs two operations: firstly, it predicts the current block to obtain the prediction block; secondly, it parses the bitstream to obtain the quantization coefficient matrix, and then performs inverse quantization and inverse transform on the quantization coefficient matrix to obtain the residual block. Finally, it adds the prediction block and the residual block to obtain the reconstructed block. The reconstructed blocks form the reconstructed image, and the decoded image is obtained by performing loop filtering on the reconstructed image or on the blocks.

[0052] It is worth noting that the current block can be the current codec unit (CU) or the current prediction unit (PU), etc.

[0053] Furthermore, the encoder, like the decoder, requires similar operations to obtain the decoded image. The decoded image can serve as a reference frame for inter-frame prediction in subsequent frames. The block partitioning information, prediction, transform, quantization, entropy coding, loop filtering, and other mode or parameter information determined by the encoder need to be written into the bitstream if necessary. The decoder determines the same block partitioning information, prediction, transform, quantization, entropy coding, loop filtering, and other mode or parameter information as the encoder by parsing and analyzing existing information, thus ensuring that the decoded image obtained by the encoder and the decoder are identical. The decoded image obtained by the encoder is often called the reconstructed image. During prediction, the codec can divide the current block into prediction units; during transform, it can divide the current block into transform units. The division of prediction units and transform units can be different. The above describes the basic flow of a video codec under a block-based hybrid coding framework. With technological advancements, some modules or steps of this framework or process may be optimized, and this application does not specifically limit this.

[0054] For ease of understanding, the video encoding and decoding system involved in the embodiments of this application will be introduced first with reference to Figure 1.

[0055] Figure 1 is a schematic block diagram of an encoding / decoding system according to an embodiment of this application.

[0056] As shown in Figure 1, the encoding and decoding system 100 includes an encoding device 110 and a decoding device 120.

[0057] The encoding device 110 encodes (can be understood as compressing) video or image data to generate a bitstream, and transmits the bitstream to the decoding device 120. The decoding device 120 decodes the bitstream generated by the encoding device 110 to obtain the decoded video or image data.

[0058] Encoding device 110 can be understood as a device capable of encoding video or images, and decoding device 120 can be understood as a device capable of decoding video or images. Encoding device 110 can modulate the encoded data according to a communication standard and transmit the modulated data to decoding device 120. The encoding device 110 or decoding device 120 includes a wider range of devices, such as smartphones, desktop computers, mobile computing devices, laptops (e.g., tablet computers), tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, etc.

[0059] Encoding device 110 can transmit encoded data (e.g., bit stream) to decoding device 120 via channel 130.

[0060] Channel 130 may include one or more media and / or means capable of transmitting encoded data from encoding device 110 to decoding device 120. Channel 130 may include one or more communication media enabling encoding device 110 to directly transmit encoded data to decoding device 120 in real time. The communication media may include wireless communication media, such as radio frequency spectrum. The communication media may also include wired communication media, such as one or more physical transmission lines. Channel 130 may include a storage medium that can store the encoded data from encoding device 110. The storage medium includes various local access data storage media, such as optical discs, DVDs, flash memory, etc. Decoding device 120 can retrieve the encoded data from the storage medium. Channel 130 may include a storage server that can store the encoded data from encoding device 110. Decoding device 120 can download the stored encoded data from the storage server. Optionally, the storage server can store the encoded data and transmit it to decoding device 120, such as a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.

[0061] The encoding device 110 includes an encoder 112 and an output interface 113.

[0062] The output interface 113 may include a modulator / demodulator (modem) and / or a transmitter. The encoder 112 transmits the encoded data directly to the decoding device 120 via the output interface 113. The encoded data may also be stored on a storage medium or a storage server for later retrieval by the decoding device 120.

[0063] In addition to the encoder 112 and the input interface 113, the encoding device 110 may also include a video source 111 or an image source.

[0064] Video source 111 may include at least one of a video capture device (e.g., a video camera), a video archive, a video input interface, and a computer graphics system, wherein the video input interface is used to receive video data from a video content provider, and the computer graphics system is used to generate the video data. Encoder 112 encodes the video data from video source 111 to generate a bitstream. The video data may include one or more pictures or a sequence of pictures. The bitstream contains the encoding information of the pictures or the sequence of pictures in the form of a bitstream. The encoding information may include encoded image data and associated data. The associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax structures. The SPS may contain parameters applied to one or more sequences. The PPS may contain parameters applied to one or more pictures. A syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.

[0065] Decoding device 120 includes input interface 121 and decoder 122. Input interface 121 may include receiver and / or modem.

[0066] In addition to the input interface 121 and the decoder 122, the decoding device 120 may also include a display device 123.

[0067] Input interface 121 can receive encoded data via channel 130. Decoder 122 decodes the encoded data to obtain decoded data and transmits the decoded data to display device 123. Display device 123 displays the decoded data. Display device 123 can be integrated with decoding device 120 or external to decoding device 120. Display device 123 can include various display devices, such as liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or other types of display devices.

[0068] It should be understood that Figure 1 is merely an example of this application and should not be construed as a display of this application. That is to say, the technical solutions of the embodiments of this application are not limited to the system framework shown in Figure 1. For example, the technology of this application can also be applied to one-sided video encoding or one-sided video decoding.

[0069] The video coding framework involved in the embodiments of this application is described below.

[0070] Figure 2 is a schematic block diagram of the video encoder 200 involved in an embodiment of this application.

[0071] It should be understood that the video encoder 200 can be applied to image data in luminance-chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4, where Y represents luminance (Luma), Cb(U) represents blue chrominance, Cr(V) represents red chrominance, and U and V represent chrominance (Chroma) used to describe color and saturation. For example, in color format, 4:2:0 means that there are 4 luminance components and 2 chrominance components (YYYYCbCr) per 4 pixels; 4:2:2 means that there are 4 luminance components and 4 chrominance components (YYYYCbCrCbCr) per 4 pixels; and 4:4:4 means full pixel display (YYYYCbCrCbCrCbCrCbCr). Of course, it can also be applied to image data in red-green-blue (RGB) format, and this application does not specifically limit this application.

[0072] After reading the video stream, the video encoder 200 divides each frame of the video stream into several coding tree units (CTUs). In some examples, a CTU may be called a "tree block," "largest coding unit" (LCU), or "coding tree block" (CTB). Each CTU can be associated with a pixel block of equal size within the image. Each pixel can correspond to one luminance (luma) sample and two chrominance (chroma) samples. Therefore, each CTU can be associated with one luminance sampling block and two chrominance sampling blocks. The size of a CTU can be, for example, 128×128, 64×64, 32×32, etc. Figure 3 is a schematic structural diagram of the relationship between coding tree units and coding units provided in the embodiments of this application. As shown in Figure 3, a CTU can be further divided into several coding units (CUs) for encoding. CUs can be rectangular blocks or square blocks. The CU can be further divided into prediction units (PU) and transform units (TU), thus separating encoding, prediction, and transformation for more flexible processing. In one example, the CTU is divided into CUs in a tree (e.g., a quadtree), and the CUs are divided into TUs and PUs in a tree (e.g., a quadtree).

[0073] The video encoder and video decoder support various PU sizes.

[0074] Assuming a specific CU size of 2N×2N, the video encoder and decoder can support PU sizes of 2N×2N or N×N for intra-frame prediction, and support symmetric PUs of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter-frame prediction. The video encoder and decoder can also support asymmetric PUs of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter-frame prediction.

[0075] As shown in Figure 2, the video encoder 200 may include: a prediction unit 210, a residual unit 220, a transform / quantization unit 230, an inverse transform / quantization unit 240, a reconstruction unit 250, a loop filtering unit 260, a decoded image buffer 270, and an entropy coding unit 280. It should be noted that the video encoder 200 may contain more, fewer, or different functional components. In this application, the current block may be referred to as the current coding unit (CU) or the current prediction unit (PU), etc. The prediction block may also be referred to as the predicted image block or the image prediction block, and the reconstructed image block may also be referred to as the reconstruction block or the image reconstruction block.

[0076] Prediction unit 210 includes an inter-prediction unit 211 and an intra-prediction unit 212. Because there is a strong correlation between adjacent pixels in an image within a video, intra-prediction is used in video encoding and decoding to eliminate spatial redundancy between adjacent pixels. Because there is a strong similarity between adjacent images in a video, inter-prediction is used to eliminate temporal redundancy between adjacent images, thereby improving coding efficiency.

[0077] The inter-frame prediction unit 211 can be used for inter-frame prediction, which can include motion estimation and motion compensation. It can reference image information from different frames. Inter-frame prediction uses motion information to find a reference block in the reference frame and generates a prediction block based on the reference block to eliminate temporal redundancy. The reference frame can be a P-frame and / or a B-frame, where P-frame refers to a forward prediction frame and B-frame refers to a bidirectional prediction frame. After finding the reference block using motion information, inter-frame prediction generates a prediction block based on the reference block. Motion information includes the frame list to which the reference frame belongs, the frame index, and motion vectors. Motion vectors can be integer-pixel or fractional-pixel. If the motion vector is fractional-pixel, then interpolation filtering needs to be used in the reference frame to create the required fractional-pixel blocks. The reference block is the integer-pixel or fractional-pixel block found based on the motion vector. Some techniques directly use the reference block as the prediction block, while others process the reference block further to generate the prediction block. Processing the reference block further to generate the prediction block can also be understood as using the reference block as the prediction block and then processing it to generate a new prediction block.

[0078] Intra-prediction unit 212 refers only to information from the same frame image to predict pixel information within the current code image block, thereby eliminating spatial redundancy. The reference frame used for intra-prediction can be an I-frame.

[0079] Intra-frame prediction employs various prediction modes. Angular and non-angle prediction modes can be used to predict the image block to be encoded, resulting in a prediction block. Based on the prediction block and the image block to be encoded, rate-distortion information is calculated, and the optimal prediction mode for the image block to be encoded is selected. This prediction mode is then written into the bitstream for transmission to the decoder. The decoder parses the prediction mode, predicts the target decoded block, and superimposes it with the temporal residual block obtained from the bitstream to obtain the reconstructed block.

[0080] Taking the H-series international digital video coding standards as an example, the H.264 / AVC standard has 8 angular prediction modes and 1 non-angle prediction mode, while H.265 / HEVC extends this to 33 angular prediction modes and 2 non-angle prediction modes. HEVC uses 35 intra-frame prediction modes: Planar, DC, and 33 angular modes. VVC uses 67 intra-frame prediction modes: Planar, DC, and 65 angular modes, including traditional and non-traditional prediction modes. Non-traditional prediction modes can include Matrix-weighted intra-frame prediction (MIP) modes. Traditional prediction modes include: Planar mode (mode number 0), DC mode (mode number 1), and angular prediction modes (mode numbers 2 to 66). It should be noted that as the angle mode increases, the prediction results of intra-frame prediction will be more accurate and better meet the needs of the development of high-definition and ultra-high-definition digital video. The above intra-frame prediction modes are only examples of this application and should not be used to limit this application.

[0081] The residual unit 220 can generate a residual block of the CU based on the pixel block of the CU and the prediction block of the PU of the CU. For example, the residual unit 220 can generate a residual block of the CU such that each sample in the residual block has a value equal to the difference between the sample in the pixel block of the CU and the corresponding sample in the prediction block of the PU of the CU.

[0082] Transform / quantization unit 230 can quantize transform coefficients. Transform / quantization unit 230 can quantize transform coefficients associated with the TU of the CU based on the quantization parameter (QP) value associated with the CU. Video encoder 200 can adjust the degree of quantization applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU.

[0083] The inverse transform / quantization unit 240 can apply inverse quantization and inverse transform to the quantized transform coefficients to reconstruct the residual block from the quantized transform coefficients.

[0084] The reconstruction unit 250 can add samples of the reconstructed residual block to corresponding samples of one or more prediction blocks generated by the prediction unit 210 to produce a reconstructed image block associated with the TU. By reconstructing the sampled blocks of each TU of the CU in this way, the video encoder 200 can reconstruct the pixel blocks of the CU.

[0085] The loop filtering unit 260 processes the pixels after inverse transform and inverse quantization to compensate for distortion information and provide a better reference for subsequent encoded pixels. For example, it can perform deblocking filtering to reduce block artifacts in pixel blocks associated with the CU. In some embodiments, the loop filtering unit 260 includes a deblocking filter (DBF) unit and a sample adaptive compensation / adaptive loop filtering (SAO / ALF) unit, wherein the DBF unit is used to remove block artifacts and the SAO / ALF unit is used to remove ringing artifacts.

[0086] The decoded image buffer 270 can store the reconstructed pixel blocks.

[0087] Specifically, the inter-frame prediction unit 211 can use a reference image containing reconstructed pixel blocks in the decoded image buffer 270 to perform inter-frame prediction on PUs of other images. Additionally, the intra-frame prediction unit 212 can use reconstructed pixel blocks in the decoded image buffer 270 to perform intra-frame prediction on other PUs in the same image as the CU.

[0088] Entropy coding unit 280 can receive quantized transform coefficients from transform / quantization unit 230. Entropy coding unit 280 can perform one or more entropy coding operations on the quantized transform coefficients to produce entropy-coded data.

[0089] Figure 4 is a schematic block diagram of the video decoder involved in the embodiments of this application.

[0090] As shown in Figure 4, the video decoder 300 includes: an entropy decoding unit 310, a prediction unit 320, an inverse quantization / transform unit 330, a reconstruction unit 340, a loop filtering unit 350, and a decoded image buffer 360. It should be noted that the video decoder 300 may contain more, fewer, or different functional components.

[0091] The video decoder 300 can receive a bitstream. The entropy decoding unit 310 can parse the bitstream to extract syntax elements. As part of parsing the bitstream, the entropy decoding unit 310 can parse the entropy-encoded syntax elements in the bitstream. The prediction unit 320, the inverse quantization / transform unit 330, the reconstruction unit 340, and the loop filtering unit 350 can decode the video data based on the syntax elements extracted from the bitstream, i.e., generate decoded video data.

[0092] The prediction unit 320 includes: an intra prediction unit 322 and an inter prediction unit 321.

[0093] Intra-prediction unit 322 can perform intra-prediction to generate prediction blocks for the PU. Intra-prediction unit 322 can use an intra-prediction mode to generate prediction blocks for the PU based on pixel blocks of spatially adjacent PUs. Intra-prediction unit 322 can also determine the intra-prediction mode of the PU based on one or more syntax elements parsed from the bitstream.

[0094] Inter-frame prediction unit 321 can construct a first reference image list (list 0) and a second reference image list (list 1) based on the syntax elements parsed from the bitstream. Furthermore, if the PU uses inter-frame prediction coding, the entropy decoding unit 310 can parse the motion information of the PU. Inter-frame prediction unit 321 can determine one or more reference blocks of the PU based on the motion information of the PU. Inter-frame prediction unit 321 can generate prediction blocks for the PU based on one or more reference blocks of the PU.

[0095] The dequantization / transform unit 330 reversibly quantizes (i.e., dequantizes) the transform coefficients associated with the TU. The dequantization / transform unit 330 can use the QP value associated with the CU of the TU to determine the degree of quantization. After dequantizing the transform coefficients, the dequantization / transform unit 330 can apply one or more inverse transforms to the dequantized transform coefficients to produce a residual block associated with the TU.

[0096] The reconstruction unit 340 uses the residual block associated with the TU of the CU and the prediction block of the PU of the CU to reconstruct the pixel block of the CU. For example, the reconstruction unit 340 can add the sample of the residual block to the corresponding sample of the prediction block to reconstruct the pixel block of the CU, thereby obtaining the reconstructed image block.

[0097] The loop filter unit 350 can perform deblocking filtering operations to reduce the block effect of pixel blocks associated with the CU.

[0098] The video decoder 300 can store the reconstructed image of the CU in the decoded image buffer 360. The video decoder 300 can use the reconstructed image in the decoded image buffer 360 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.

[0099] Combining Figures 2 and 4, the basic process of video encoding and decoding is as follows:

[0100] At the encoding end, a frame of image is divided into image blocks. For the current block, prediction unit 210 uses intra-frame prediction or inter-frame prediction to predict the prediction block of the current block (i.e., the block to be encoded). Residual unit 220 can calculate the residual block, i.e., the difference between the prediction block and the original block, based on the prediction block and the original block of the current block (i.e., the block to be encoded). This residual block can also be called residual information. This residual block can be transformed and quantized by transform / quantization unit 230 to remove information that is not sensitive to the human eye, thereby eliminating visual redundancy. Optionally, the residual block before transformation and quantization by transform / quantization unit 230 can be called a temporal residual block, and the temporal residual block after transformation and quantization by transform / quantization unit 230 can be called a frequency residual block or a frequency domain residual block. Entropy coding unit 280 receives the quantized change coefficients output by change quantization unit 230 and can perform entropy coding on the quantized change coefficients to output a bitstream. For example, entropy coding unit 280 can eliminate character redundancy based on the target context model and the probability information of the binary bitstream.

[0101] At the decoding end, the entropy decoding unit 310 can parse the bitstream to obtain the prediction information and quantization coefficient matrix of the current block (i.e., the block to be decoded). Based on the prediction information, the prediction unit 320 uses intra-frame prediction or inter-frame prediction to predict the prediction block of the current block (i.e., the block to be decoded). The dequantization / transform unit 330 uses the quantization coefficient matrix obtained from the bitstream to perform dequantization and inverse transform on the quantization coefficient matrix to obtain the residual block. The reconstruction unit 340 adds the prediction block and the residual block to obtain the reconstructed block. The reconstructed blocks form the reconstructed image. The loop filtering unit 350 performs loop filtering on the reconstructed image based on the image or based on the blocks to obtain the decoded image. It is worth noting that the encoding end also needs to use similar operations as the decoder to obtain the decoded image. This decoded image can also be called the reconstructed image, and the reconstructed image can be used as a reference frame for inter-frame prediction in subsequent frames.

[0102] Furthermore, the block partitioning information determined by the encoder, as well as mode information or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering, are carried in the bitstream when necessary. The decoder determines the same block partitioning information, prediction, transform, quantization, entropy coding, and loop filtering mode information or parameter information as the encoder by parsing the bitstream and analyzing existing information, thereby ensuring that the decoded image obtained by the encoder is the same as the decoded image obtained by the decoder.

[0103] It should be noted that, due to the need for parallel processing, images can be divided into slices, etc. Slices within the same image can be processed in parallel, meaning there is no data dependency between them. The term "frame" can be understood as an image or slice, etc. Furthermore, Figures 1 to 4 are merely examples of this application and should not be construed as limiting this application. In other alternative embodiments, the decoding and encoding methods provided in this application can also be applied to any other type of encoding / decoding system, encoding framework, or decoding framework that meets its application conditions. For example, with technological advancements, some modules in the systems or frameworks mentioned above, or some steps in the processes described above, may be optimized. In such cases, the decoding and encoding methods provided in this application can also be applied to systems, frameworks, and processes optimized based on them.

[0104] This application applies to the encoding and decoding of light field video.

[0105] Light field video is video captured by multiple cameras or arrays of multiple cameras, and it is a research topic of the MPEG Lenslet Video Coding (LVC) working group. Unlike typical camera imaging models, light field cameras add a microlens array in front of the imaging plane, allowing light rays from the same point on the object's plane to be captured simultaneously by multiple microlenses, essentially capturing the same point from multiple angles at the same time.

[0106] Figure 5 is an example of the transmission process of light field video provided in an embodiment of this application.

[0107] As shown in Figure 5, light field video can be acquired by a capture / shooting device (such as a light field camera), processed into a light field video of a specific data format, and input into a codec. The codec encodes and decodes the video and outputs it as a light field video of a specific data format. The display device receives the light field video of the specific data format and displays it.

[0108] Due to its unique imaging model, the visual effect of light field images differs significantly from traditional images, leading to poor performance of compression methods typically used for general images or videos when processing light field images or videos. The LVC working group was established to address this issue by researching compression methods more suitable for light field videos. However, it is worth noting that LVC can still utilize conventional video codec tools (AVC, HEVC, or VVC, etc.) for its encoding and / or decoding parts. For example, LVC can use the system shown in Figure 1 or the framework shown in Figures 2 and 3.

[0109] The following describes the technical solutions provided in the embodiments of this application.

[0110] (1) Improvement of motion estimation search algorithm.

[0111] To improve the compression performance of light field video, the motion estimation search algorithm in the inter-frame prediction stage can be improved by utilizing the arrangement rules of macro pixels in the light field image. Motion estimation refers to finding an optimal matching block on the reference image that minimizes the rate-distortion value of the current coding block in the current image. The optimization goal of the search algorithm is to find the optimal matching block faster and more accurately, and then point the final motion vector to this optimal matching block. A specific improvement to the motion estimation search algorithm can be achieved as follows: In the motion estimation stage, the position pointed to by the initially predicted motion vector is used as the initial position (also called the initial point, initial search position, or initial search point). Within a certain search range, the positions of corresponding points within each macro pixel are searched sequentially according to the macro pixel arrangement template; then, the optimal position is selected from the initial position and the corresponding point positions.

[0112] In this context, the position of a corresponding location within its own macropixel is the same as the position of the initial location within its own macropixel. The spacing between corresponding locations is the spacing between macropixels. For example, the image block corresponding to a corresponding location within its own macropixel is the same as the image block corresponding to the initial location within its own macropixel. The image block corresponding to the corresponding location can be an image block whose corresponding location is located at the top left corner (or bottom left corner, top right corner, bottom right corner, center, or other specific location); similarly, the image block corresponding to the initial location can be an image block whose initial location is located at the top left corner (or bottom left corner, top right corner, bottom right corner, center, or other specific location).

[0113] Figure 6 is an example of the initial position and co-location provided in the embodiments of this application.

[0114] As shown in Figure 6, the search range for the same location includes multiple macro pixels, including an initial macro pixel, which in turn includes an image block corresponding to the initial position. Furthermore, for any macro pixel other than the initial macro pixel, this individual macro pixel includes an image block corresponding to the same location. The position of this image block within its macro pixel is the same as the position of the initial location within the initial macro pixel. Specifically, as shown in Figure 6(a), both the image block corresponding to the initial position and the image block corresponding to the same location are located in the middle of their respective macro pixels; as shown in Figure 6(b), both the image block corresponding to the initial position and the image block corresponding to the same location are located at the lower right corner of their respective macro pixels.

[0115] Figure 7 is an example of a light field image provided in an embodiment of this application.

[0116] As shown in Figure 7, since a light field image is composed of a series of regularly arranged macropixels, there is a strong correlation between adjacent macropixels according to the imaging principle of a light field camera. Therefore, in the motion estimation search process, searching for the same-site locations of each macropixel can make fuller use of the correlation of the light field image, thereby bringing more efficient compression performance.

[0117] It should be understood that the term "macropixel" can also be referred to as "microimage," "microlens image," or other terms with similar meanings.

[0118] However, while this improved method considers the macro-pixel arrangement pattern of the light field image, the optimal candidates for matching blocks are not necessarily arranged strictly according to the macro-pixel spacing. Therefore, searching only for co-location positions within macro-pixels may miss some local optima. In view of this, this application provides an encoding method that can make fuller use of the correlation of the light field image, thereby further optimizing the motion estimation effect of light field video compression. Specifically, the encoder can search for at least one first position within a first search region within the initial macro-pixel where the initial position is located, and optimize the initial position based on the at least one first position; optionally, the encoder can also search for co-location positions of the initial position based on the search for at least one first position, and then optimize the initial position based on the co-location positions; optionally, the encoder can also search for at least one second position within a second search region within the macro-pixel where the co-location position is located, and then optimize the initial position based on the at least one second position. For example, the encoder can appropriately add local grid search on the basis of searching for the same position within the macro pixel, that is, to offset and correct the position of the candidate reference block obtained based on the same position within the macro pixel, so as to maximize the correlation with the current block, thereby further optimizing the motion estimation effect of light field video compression.

[0119] Figure 8 is a schematic flowchart of the encoding method 400 provided in an embodiment of this application. It should be understood that the encoding method 400 can be executed by an encoder. For example, the encoding method 400 can be executed by the encoding device 110 or encoder 112 shown in Figure 1. As another example, the encoding method 400 can be executed by the video encoder 200 shown in Figure 2. For ease of description, the encoder will be used as the executing entity of the encoding method 400 in the following exemplary description.

[0120] As shown in Figure 8, the encoding method 400 may include:

[0121] S410, the encoder performs motion estimation on the current block in the current image and determines the first motion parameter of the current block; the first motion parameter points to the initial position in the reference image of the current image.

[0122] For example, the current image and the reference image of the current image are light field images.

[0123] For example, the encoder may use an Advanced Motion Vector Prediction Model (AMVP) mode or other types of prediction modes to perform motion estimation on the current block to obtain the first motion parameters.

[0124] For example, the initial position may be a location of a reference block in the reference image to which the first motion parameter points. Alternatively, the first motion parameter points to a reference block in the reference image, and the initial position is a location of that reference block. For instance, the initial position may be the upper left corner, lower left corner, upper right corner, lower right corner, center, or other specific location of the reference block.

[0125] For example, the initial position may be the position of the pixel that the first motion parameter points to in the reference image.

[0126] It is worth noting that the term "position" can be equivalently replaced by "pixel," "the location of the pixel," or other descriptions with similar meanings, and this application does not specifically limit it in this regard. For example, the initial position can be equivalently replaced by the initial pixel, the location of the initial pixel, or other terms with similar meanings.

[0127] S420, the encoder determines at least one first position other than the initial position within the first search area where the initial position is located; the first search area is less than or equal to the initial macro pixel where the initial position is located.

[0128] For example, the first search area includes a region centered on the initial position.

[0129] For example, the center position of the first search area is the center, upper left corner, lower left corner, upper right corner, lower right corner, or other positions of the initial macro pixel.

[0130] For example, the range of the first search region can be any numerical value. For instance, the range of the first search region can be any positive integer.

[0131] For example, the first search area may be a predefined search area. For instance, the first search area may be implemented by pre-saving corresponding codes, tables, or other means in the encoder that can be used to indicate relevant information, or the first search area may be agreed upon or defined by a standard protocol.

[0132] For example, the first search region may be equal to the initial macro pixel, in which case the concept of the first search region may not be introduced. That is, the encoder determines the at least one first position within the initial macro pixel.

[0133] For example, the first search area may be smaller than the initial macro pixel; in this case, the first search area may also be referred to as a local search area. That is, the encoder determines the at least one first position within a local area of ​​the initial macro pixel.

[0134] For example, the encoder can search for the at least one first position within the first search area using any search method. For instance, the encoder can search for the at least one first position within the first search area using a grid search method (also known as a mesh search method). Alternatively, the encoder can search for the at least one first position within the first search area using a random sampling method.

[0135] Figure 9 is an example of the search principle for at least one first position provided by an embodiment of this application.

[0136] As shown in Figure 9, the encoder can search for the at least one first position within the first search area using a grid search method. Specifically, assuming the first search area includes 7×7 pixels and the grid search interval is 1 pixel, the encoder determines the pixels at the intersection of rows 1, 3, 5, and 7 and columns 1, 3, 5, and 7 as the searched 4×4 pixels, i.e., the at least one first position.

[0137] S430, the encoder determines the second motion parameters of the reference block based on the initial position and the at least one first position.

[0138] For example, the encoder may determine an optimal position among the initial position and the at least one first position, and determine the motion parameter pointing to the optimal position as the second motion parameter.

[0139] For example, the encoder can consider other positions based on the initial position and the at least one first position. That is, the encoder can determine an optimal position among the initial position, the at least one first position, and other positions to be considered, and determine the motion parameter pointing to the optimal position as the second motion parameter. For example, the other positions may include one or more co-location positions.

[0140] For example, the second motion parameter may be the final motion parameter.

[0141] For example, the second motion parameter may be a motion parameter that needs further optimization, that is, the encoder may further process the second motion parameter to obtain the final motion parameter.

[0142] It is worth noting that the term "motion parameter" is intended to include motion vectors or other motion parameters that can represent the motion of an image patch in a reference image relative to a specific position. This specific position can be a position in the reference image that is the same as the position of the current patch in the current image.

[0143] For example, after the encoder obtains the second motion parameter, it can use the reference block pointed to by the second motion parameter to predict the current block and obtain the predicted block of the current block. Then, based on the original block of the current block and the predicted block of the current block, the encoder determines the residual block of the current block. Further, the encoder can quantize and entropy encode the residual block to obtain a bitstream, such as a bitstream of a light field image or a light field video.

[0144] In this embodiment, the position pointed to by the first motion parameter determined by motion estimation is taken as the initial position. Then, within the first search area within the initial macro pixel where the initial position is located, at least one first position other than the initial position is determined. Based on the initial position and the at least one first position, the second motion parameter of the reference block is determined. In other words, by introducing the at least one first position to correct the first motion parameter and obtaining the corrected second motion parameter, the accuracy of the second motion parameter can be improved, thereby improving the motion estimation effect and coding performance of the encoder.

[0145] In some embodiments, S430 may include:

[0146] The encoder determines the motion parameters pointing to the position with the minimum rate-distortion cost among the initial position and the at least one first position as the second motion parameters.

[0147] For example, the encoder first determines the rate-distortion cost corresponding to the initial position and the rate-distortion cost corresponding to any one of the at least one first position. Then, based on the rate-distortion cost corresponding to the initial position and the rate-distortion cost corresponding to any one of the first positions, the position with the minimum rate-distortion cost is determined as the optimal position. Then, the encoder determines the motion parameters of the optimal position as the second motion parameters.

[0148] It is worth noting that the term "rate-distortion cost" is intended to be used to characterize parameters of distortion and bit rate. For example, "rate-distortion cost" may include Peak Signal-to-Noise Ratio (PSNR), Mean Structural Similarity Index Measure (MSSIM), or other parameters with similar functions. This application does not specifically limit the calculation method of "rate-distortion cost". Of course, in other alternative embodiments, "rate-distortion cost" can also be replaced with parameters used only to characterize distortion or bit rate to reduce complexity, and even "rate-distortion cost" can be replaced with other types of metrics used to characterize encoding or decoding performance; this application does not specifically limit this.

[0149] For example, when there are multiple positions with the minimum rate-distortion cost among the initial position and the at least one first position, the encoder can determine the motion parameter pointing to any one of the multiple positions as the second motion parameter. For instance, the encoder can determine the motion parameter pointing to the position closest to the initial position among the multiple positions as the second motion parameter. Alternatively, the encoder can randomly select a position from the multiple positions and determine the motion parameter pointing to that position as the second motion parameter.

[0150] Of course, in other alternative embodiments, the encoder may determine the second motion parameter as the motion parameter that satisfies a first condition among the initial position and the at least one first position. For example, the first condition may include: the rate-distortion cost is less than or equal to a preset threshold, or the first condition may include: the rate-distortion cost is equal to a minimum value (e.g., the minimum of the rate-distortion cost corresponding to the initial position and the rate-distortion cost corresponding to the at least one first position). The first condition may be a predefined condition. For example, the first condition may be implemented by pre-saving corresponding codes, tables, or other means in the encoder that can be used to indicate relevant information, or the first condition may be agreed upon or defined by a standard protocol.

[0151] In some embodiments, S430 may include:

[0152] Based on the initial position within the initial macropixel, at least one co-location of the initial position is determined; based on the initial position, the at least one first position, and the at least one co-location, the second motion parameter is determined.

[0153] For example, for any one of the at least one corresponding position, the position of the corresponding position within the macro pixel where the corresponding position is located is the same as the position of the initial position within the initial macro pixel.

[0154] For example, the position of the image block corresponding to any given location within the macropixel containing that location is the same as the position of the image block corresponding to the initial location within the initial macropixel. The image block corresponding to any given location can be an image block whose given location is the top-left corner (or bottom-left, top-right, bottom-right, center, or other specific location); similarly, the image block corresponding to the initial location can be an image block whose initial location is the top-left corner (or bottom-left, top-right, bottom-right, center, or other specific location).

[0155] It should be understood that the image block corresponding to any one of the same position and the image block corresponding to the initial position can be specifically referred to in the example in Figure 6. To avoid repetition, it will not be described again here.

[0156] In this embodiment, the co-location is considered based on the initial position and the at least one first position. That is, the first motion parameter is corrected by introducing the at least one first position and the at least one co-location, and the corrected second motion parameter is obtained. Compared with the scheme of correcting the first motion parameter without considering the co-location, the correlation between the reference block pointed to by the second motion parameter and the current block is increased, thereby improving the motion estimation effect and coding performance of the encoder.

[0157] In some embodiments, at least one macro pixel is determined within the search range of the at least one co-location; within any one of the at least one macro pixels, a position that is the same as the position of the initial position within the initial macro pixel is determined as the co-location position among the at least one co-location positions.

[0158] For example, the search range of the at least one co-location covers multiple macro pixels.

[0159] For example, when the search range for the at least one co-location is a rectangular range, the rectangular range can be any value larger than the distance between co-locations. The distance between co-locations is equal to the distance between macropixels.

[0160] For example, the search range for the at least one co-location can be a predefined search area. For instance, the search range can be implemented by pre-saving corresponding codes, tables, or other means of indicating relevant information in the encoder, or the search range for the at least one co-location can be agreed upon or defined by a standard protocol.

[0161] For example, the encoder determines the at least one macropixel based on the arrangement pattern of macropixels, and then, within any macropixel within the at least one macropixel, determines the corresponding position of the at least one corresponding position located within the corresponding macropixel based on the position of the initial position within the initial macropixel. The arrangement pattern includes, but is not limited to, at least one of the following: the shape of the macropixel, the spacing between macropixels, and the arrangement direction of the macropixels.

[0162] For example, within each of the at least one macropixels, the encoder determines the position that is the same as the position of the initial position within the initial macropixel as the corresponding position among the at least one corresponding position. In other words, the corresponding position within each macropixel is the same as the position of the initial position within the initial macropixel.

[0163] Figure 10 is an example of the same site location provided in the embodiments of this application.

[0164] As shown in Figure 10, assuming the macropixels are hexagonal pixels arranged closely together, the search range for the same location includes multiple macropixels, including the initial macropixel containing the initial position. For any macropixel other than the initial macropixel among these multiple macropixels, this arbitrary macropixel includes a same location; wherein the position of this arbitrary location within its macropixel is the same as the position of the initial position within the initial macropixel. Specifically, as shown in Figure 10, both the initial position and this arbitrary location are the center positions of their respective macropixels.

[0165] In some embodiments, the motion parameter pointing to the position with the minimum rate-distortion cost among the initial position, the at least one first position, and the at least one co-located position is determined as the second motion parameter.

[0166] For example, the encoder first determines the rate-distortion cost corresponding to the initial position, the rate-distortion cost corresponding to any one of the at least one first position, and the rate-distortion cost corresponding to any one of the at least one same-site position. Then, based on the rate-distortion cost corresponding to the initial position, the rate-distortion cost corresponding to any one of the first positions, and the rate-distortion cost corresponding to any one of the same-site positions, the position with the minimum rate-distortion cost is determined as the optimal position. The encoder then executes the motion parameters of the optimal position and determines them as the second motion parameters. The term "rate-distortion cost" can be referred to in the above description; to avoid repetition, it will not be repeated here.

[0167] For example, when there are multiple positions among the initial position, the at least one first position, and the at least one corresponding position that have the lowest rate-distortion cost, the encoder can determine the motion parameter pointing to any one of these multiple positions as the second motion parameter. For instance, the encoder can determine the motion parameter pointing to the position among the multiple positions that is closest to the initial position as the second motion parameter. Alternatively, the encoder can randomly select one of the multiple positions and determine the motion parameter pointing to that position as the second motion parameter.

[0168] Of course, in other alternative embodiments, the encoder may determine the second motion parameter as a motion parameter that satisfies a second condition among the initial position, the at least one first position, and the at least one corresponding position. For example, the second condition may include: rate-distortion cost is less than or equal to a preset threshold, or the second condition may include: rate-distortion cost equal to a minimum value (e.g., the minimum of the rate-distortion cost corresponding to the initial position, the rate-distortion cost corresponding to the at least one first position, and the rate-distortion cost corresponding to the at least one corresponding position). The second condition may be a predefined condition. For example, the second condition may be implemented by pre-saving corresponding codes, tables, or other methods that can be used to indicate relevant information in the encoder, or the second condition may be agreed upon or defined by a standard protocol.

[0169] In some embodiments, within a second search region where the first co-location is located among the at least one co-location locations, at least one second location other than the first co-location is determined; the second search region is less than or equal to the macro pixel where the first co-location is located; and the motion parameter pointing to the location with the minimum rate-distortion cost among the initial location, the at least one first location, the at least one co-location location, and the at least one second location is determined as the second motion parameter.

[0170] For example, the second search area includes a region centered on the first co-location.

[0171] For example, the center position of the second search area is the center, upper left corner, lower left corner, upper right corner, lower right corner, or other position of the macro pixel where the first co-location is located.

[0172] For example, the range of the second search region can be any numerical value. For instance, the range of the second search region can be any positive integer.

[0173] For example, the second search area can be a predefined search area. For instance, the second search area can be implemented by pre-saving corresponding codes, tables, or other means in the encoder that can be used to indicate relevant information, or the second search area can be agreed upon or defined by a standard protocol.

[0174] For example, the second search region may be equal to the macropixel where the first co-location is located. In this case, the concept of a second search region may not be introduced. That is, the encoder determines the at least one second location within the macropixel where the first co-location is located.

[0175] For example, the second search region may be smaller than the macropixel where the first co-location is located. In this case, the second search region may also be referred to as a local search region. That is, the encoder determines the at least one second position within a local region of the macropixel where the first co-location is located.

[0176] For example, the encoder can search for the at least one second position within the second search area using any search method. For instance, the encoder can search for the at least one second position within the second search area using a grid search method (also known as a mesh search method). Alternatively, the encoder can search for the at least one second position within the second search area using a random sampling method. It is worth noting that the method by which the encoder searches for the at least one second position within the second search area may be the same as or different from the method by which the encoder searches for the at least one first position within the first search area. For details, please refer to the scheme illustrated in Figure 9 above; to avoid repetition, it will not be elaborated further here.

[0177] For example, the encoder first determines the rate-distortion cost corresponding to the initial position, the rate-distortion cost corresponding to any one of the at least one first position, the rate-distortion cost corresponding to any one of the at least one same-site position, and the rate-distortion cost corresponding to any one of the at least one second position. Then, based on the rate-distortion costs corresponding to the initial position, the rate-distortion costs corresponding to any one of the first positions, the rate-distortion costs corresponding to any one of the same-site positions, and the rate-distortion costs corresponding to any one of the second positions, the position with the minimum rate-distortion cost is determined as the optimal position. The encoder then executes the motion parameters of the optimal position and determines them as the second motion parameters. The term "rate-distortion cost" can be referred to in the above explanation; to avoid repetition, it will not be repeated here.

[0178] For example, when there are multiple positions among the initial position, the at least one first position, the at least one corresponding position, and the at least one second position that have the lowest rate-distortion cost, the encoder can determine the motion parameter pointing to any one of the multiple positions as the second motion parameter. For instance, the encoder can determine the motion parameter pointing to the position among the multiple positions that is closest to the initial position as the second motion parameter. Alternatively, the encoder can randomly select one position from the multiple positions and determine the motion parameter pointing to that position as the second motion parameter.

[0179] Of course, in other alternative embodiments, the encoder may determine the second motion parameter as a motion parameter that satisfies a third condition among the initial position, the at least one first position, the at least one corresponding position, and the at least one second position. For example, the third condition may include: rate-distortion cost is less than or equal to a preset threshold, or the third condition may include: rate-distortion cost equal to a minimum value (e.g., the minimum of the rate-distortion cost corresponding to the initial position, the rate-distortion cost corresponding to the at least one first position, the rate-distortion cost corresponding to the at least one corresponding position, and the rate-distortion cost corresponding to the at least one second position). The third condition may be a predefined condition. For example, the third condition may be implemented by pre-saving corresponding codes, tables, or other methods that can be used to indicate relevant information in the encoder, or the third condition may be agreed upon or defined by a standard protocol.

[0180] In this embodiment, at least one second position is considered based on the initial position, the at least one first position, and the at least one co-location position. That is, the first motion parameter is corrected by introducing the at least one first position, the at least one co-location position, and the at least one second position. Compared with the scheme of correcting the first motion parameter without considering the second position, this increases the correlation between the reference block pointed to by the second motion parameter and the current block, thereby improving the motion estimation effect and coding performance of the encoder.

[0181] In some embodiments, the range of the second search region is the same as the range of the first search region.

[0182] Of course, in other alternative embodiments, the range of the second search area and the range of the first search area may also be different, and this application does not specifically limit this.

[0183] For example, the range of the second search area can be smaller than the range of the first search area.

[0184] In some embodiments, a first search interval is determined based on a first distance between the first co-location and the initial location; based on the first search interval, the at least one second location is searched in the second search area in a grid search manner.

[0185] For example, the encoder determines the first search spacing based on a predefined first mapping relationship, specifically the spacing corresponding to the first distance within that mapping relationship. The first mapping relationship may include multiple distances and the spacing corresponding to each distance. The first mapping relationship can be implemented by pre-storing corresponding codes, tables, or other methods that can indicate relevant information in the encoder, or it can be agreed upon or defined by a standard protocol.

[0186] For example, the encoder determines the spacing corresponding to the first interval containing the first distance as the first search spacing. Alternatively, the encoder can determine the spacing corresponding to the first interval in a predefined second mapping relationship as the first search spacing. The second mapping relationship may include multiple distance intervals and the spacing corresponding to each distance interval. The second mapping relationship can be implemented by pre-storing corresponding codes, tables, or other methods that can indicate relevant information in the encoder, or the second mapping relationship can be agreed upon or defined by a standard protocol.

[0187] For example, the first search interval can be any value. For instance, the first search interval can be 0 or a positive integer.

[0188] Of course, in other alternative embodiments, the first search interval can be determined based on the search interval of the first search region.

[0189] For example, the first search interval can be the sum of the search interval of the first search region and a first value, which can be determined based on the distance between the first co-location and the initial position. For instance, the first value is positively correlated with the distance between the first co-location and the initial position; that is, the greater the distance between the first co-location and the initial position, the larger the first value. Since macropixels closer to the initial position have a stronger correlation with the current block, the positive correlation between the first value and the distance between the first co-location and the initial position allows for a denser search within the search region of macropixels close to the initial position, and a sparser search or no search at all within the search region of macropixels farther away. This improves coding efficiency while ensuring the encoder's motion estimation effect and coding performance.

[0190] In this embodiment, since the spacing of macro pixels in the light field image is difficult to guarantee to be strictly consistent, a first search spacing is determined based on the first distance between the first co-location and the initial location; based on the first search spacing, the at least one second location is searched in the second search area in a grid search manner, which can improve the robustness of the search.

[0191] In some embodiments, the first distance is positively correlated with the first search interval.

[0192] For example, the larger the first distance, the larger the first search interval.

[0193] Since macropixels closer to the initial position are more correlated with the current block, the first distance is positively correlated with the first search interval. This allows for a denser search in the search region within macropixels closer to the initial position and a sparser search or no search in the search region within macropixels farther away. This can improve coding efficiency while ensuring the motion estimation effect and coding performance of the encoder.

[0194] In some embodiments, the search interval of the first search region is less than or equal to the search interval corresponding to the distance between the closest corresponding position to the initial position and the initial position.

[0195] For example, assuming that the closest co-location to the initial position among the at least one co-location is recorded as the closest co-location, the search spacing of the first search region is less than the search spacing corresponding to the distance between the closest co-location and the initial position. This can ensure that the search region in the initial macro-pixel is the most densely searched, while the search region in the macro-pixel where the closest co-location is located is sparser. This can improve coding efficiency while ensuring the motion estimation effect and coding performance of the encoder.

[0196] For example, assuming that the closest position to the initial position among the at least one corresponding position is recorded as the closest corresponding position, the search spacing of the first search region is equal to the search spacing corresponding to the distance between the closest corresponding position and the initial position. This can ensure that the search is most densely applied to the search region in the initial macro-pixel and the search region in the macro-pixel where the closest corresponding position is located, thus ensuring the motion estimation effect and coding performance of the encoder.

[0197] In some embodiments, the at least one co-location is divided into at least one layer; based on the spacing corresponding to the first layer where the first co-location is located, the at least one second location is searched in the second search area in a grid search manner.

[0198] For example, any one of the at least one layers includes one or more co-locations among the at least one co-location.

[0199] For example, the encoder divides the at least one co-location into at least one layer; then determines the spacing corresponding to the first layer where the first co-location is located; then, based on the spacing corresponding to the first layer, the encoder performs a grid search in a grid search manner within the search area where some co-locations (e.g., filtered co-locations, which include the first co-location) or all co-locations are located on the first layer.

[0200] In other words, starting from the initial position and working outwards, the encoder can divide the at least one corresponding position into M layers (numbered 1 to M, with smaller numbers closer to the initial position). Taking a hexagonal macropixel as an example, each layer is a hexagon (or possibly a rhombus or other shape) formed by several corresponding positions. The encoder performs a local grid search within the search area containing the corresponding positions (not all of them, but a few) at the initial position and within each layer. The search interval used for the local grid search within the search area at the initial position is denoted as P0, and the search interval used for the local grid search within the search area containing the corresponding positions at each layer is denoted as P. i(0 < i ≤ M), and satisfying that when x < y, P x ≤ P y . Meanwhile, the range sizes of the local grid searches adopted everywhere are all the same.

[0201] Exemplarily, the encoder determines, based on a predefined third mapping relationship, the spacing corresponding to the first layer in the third mapping relationship as the first search spacing. Wherein, the third mapping relationship may include multiple layers and the spacings corresponding to each layer. The third mapping relationship can be implemented by pre-saving corresponding codes, tables or other means for indicating relevant information in the encoder, or the third mapping relationship can be agreed or defined by a standard protocol.

[0202] It should be noted that the term "layer" can also be equivalently replaced by terms with similar meanings such as "level", "group", "sub-group", "set", etc. in other alternative embodiments, and the present application does not make specific limitations thereto.

[0203] In some embodiments, starting from the initial position, along a plurality of predefined directions, the same-position point positions in the at least one same-position point position that have the same number of same-position point positions spaced from the initial position are divided into the same-position point positions included in one layer of the at least one layer.

[0204] Exemplarily, the plurality of directions can be implemented by pre-saving corresponding codes, tables or other means for indicating relevant information in the encoder, or the plurality of directions can be agreed or defined by a standard protocol.

[0205] Exemplarily, starting from the initial position, along a plurality of predefined directions, the encoder may only divide the same-position point positions in the at least one same-position point position that have the same number of same-position point positions spaced from the initial position into the same-position point positions included in one layer of the at least one layer. [[ID=!21]]<!!

[0206] Exemplarily, starting from the initial position, along a plurality of predefined directions, the encoder only divides the same-position point positions in the at least one same-position point position that have 0 same-position point positions spaced from the initial position into the same-position point positions included in the first layer of the at least one layer; starting from the initial position, along a plurality of predefined directions, the encoder only divides the same-position point positions in the at least one same-position point position that have 1 same-position point position spaced from the initial position into the same-position point positions included in the second layer of the at least one layer; and so on, until when the at least one same-position point position is completely divided, the same-position point positions included in each layer of the at least one layer are obtained.

[0207] For example, the encoder, starting from the initial position, divides only the at least one corresponding position with a distance of 1 from the initial position into the corresponding positions included in the first layer of the at least one layer along multiple predefined directions; the encoder, starting from the initial position, divides only the at least one corresponding position with a distance of 2 from the initial position into the corresponding positions included in the second layer of the at least one layer along multiple predefined directions; and so on, until all at least one corresponding position is divided, thus obtaining the corresponding positions included in each of the at least one layers.

[0208] For example, starting from the initial position, along a number of predefined directions, the encoder can classify all the same-site positions on a line formed by the at least one same-site position and the same number of same-site positions spaced from the initial position into the same-site positions included in one of the at least one layers.

[0209] For example, starting from the initial position, the encoder can divide all the same-site positions on the line formed by the at least one same-site position and the number of same-site positions separated from the initial position being 0, into the same-site positions included in the first layer of the at least one layer, along a predefined plurality of directions; starting from the initial position, the encoder can divide all the same-site positions on the line formed by the line formed by the at least one same-site position and the number of same-site positions separated from the initial position being 1, into the same-site positions included in the second layer of the at least one layer, and so on, until the at least one same-site position is divided, thus obtaining the same-site positions included in each of the at least one layers.

[0210] For example, starting from the initial position, the encoder can divide all the same-site positions on a line formed by the at least one same-site position and the number of same-site positions spaced apart from the initial position into the same-site positions included in the first layer of the at least one layer, along a predefined plurality of directions; starting from the initial position, the encoder can divide all the same-site positions on a line formed by the at least one same-site position and the number of same-site positions spaced apart from the initial position and the number of same-site positions spaced apart from the initial position into the same-site positions included in the second layer of the at least one layer, along a predefined plurality of directions; and so on, until the at least one same-site position is divided, thus obtaining the same-site positions included in each of the at least one layer.

[0211] Of course, the encoder can divide the at least one corresponding position into at least one layer in other ways, and this application does not specifically limit this. For example, in other alternative embodiments, the encoder takes the initial position as the starting point and divides the corresponding position among the at least one corresponding position and located on the circle corresponding to any one of the predefined radii into corresponding position positions in the layer corresponding to any one of the predefined radii, along multiple predefined radii. The multiple predefined radii can be implemented by pre-storing corresponding codes, tables, or other methods that can be used to indicate relevant information in the encoder, or the multiple predefined radii can be agreed upon or defined by a standard protocol.

[0212] In some embodiments, the plurality of directions include directions perpendicular to the edges of the initial macropixel and / or directions parallel to the diagonals of the initial macropixel.

[0213] For example, the plurality of directions includes directions perpendicular to the respective edges of the initial macro pixel.

[0214] Figure 11 is an example of at least one layer provided in an embodiment of this application.

[0215] As shown in Figure 11, assuming the initial macro pixel is a hexagonal macro pixel, the plurality of directions include 6 directions perpendicular to the 6 sides of the initial macro pixel. Starting from the initial position, the encoder divides all corresponding positions (e.g., including 6 corresponding positions) along these 6 directions from the lines formed by corresponding positions of the at least one corresponding position that are 0 times apart from the initial position into the corresponding positions included in the first layer of the at least one layer; along these 6 directions, all corresponding positions (e.g., including 12 corresponding positions) along the lines formed by corresponding positions of the at least one corresponding position that are 1 times apart from the initial position into the corresponding positions included in the second layer of the at least one layer; along these 6 directions, all corresponding positions (e.g., including 18 corresponding positions) along the lines formed by corresponding positions of the at least one corresponding position that are 2 times apart from the initial position into the corresponding positions included in the third layer of the at least one layer; and so on, until all corresponding positions are divided, thus obtaining the corresponding positions included in each of the at least one layer.

[0216] Figure 12 is an example of at least one layer provided in an embodiment of this application.

[0217] As shown in Figure 12, assuming the initial macropixel is a hexagonal macropixel, the multiple directions include 6 directions perpendicular to the six sides of the initial macropixel and 2 directions parallel to the horizontal diagonal of the initial macropixel, totaling 8 directions. Starting from the initial position, the encoder divides all corresponding positions (e.g., 8 corresponding positions) along these 8 directions from the line formed by corresponding positions of at least one corresponding position with a distance of 0 from the initial position into the corresponding positions included in the first layer of the at least one layer; along these 8 directions, it divides all corresponding positions (e.g., 16 corresponding positions) along the line formed by corresponding positions of at least one corresponding position with a distance of 1 from the initial position into the corresponding positions included in the second layer of the at least one layer; and so on, until all corresponding positions are divided, thus obtaining the corresponding positions included in each of the at least one layers.

[0218] It is worth noting that the multiple directions, including directions perpendicular to the edges of the initial macropixel and / or directions parallel to the diagonals of the initial macropixel, are merely examples of this application and should not be construed as limiting this application. For example, in other alternative embodiments, the multiple directions may also be directions determined by the encoder based on the arrangement pattern of macropixels in the image; the arrangement pattern includes, but is not limited to, at least one of the following: the shape of the macropixels, the spacing between macropixels, and the arrangement direction of the macropixels; for example, the multiple directions may include the arrangement direction of the macropixels.

[0219] In some embodiments, the shape formed by connecting the corresponding positions on the at least one layer is the same as or different from the shape of the initial macropixel.

[0220] For example, the positions of the same points on the at least one layer can be connected to form a specific shape, such as a hexagon, a rhombus, or other shapes.

[0221] For example, as shown in FIG11, the shape formed by connecting the corresponding positions on the at least one layer is a hexagon, which is the same as the shape of the initial macropixel. As shown in FIG12, the shape formed by connecting the corresponding positions on the at least one layer is a rhombus, which is different from the shape of the initial macropixel.

[0222] In some embodiments, the distance between the first layer and the initial position is positively correlated with the spacing corresponding to the first layer.

[0223] For example, the greater the distance between the first layer and the initial position, the greater the first search interval.

[0224] Since macropixels closer to the initial position are more correlated with the current block, the distance between the first layer and the initial position is positively correlated with the spacing of the first layer. This allows for a denser search in the search region within macropixels closer to the initial position, and a sparser search or no search in the search region within macropixels farther away. This can improve coding efficiency while ensuring the motion estimation effect and coding performance of the encoder.

[0225] In some embodiments, the search spacing of the first search region is less than or equal to the search spacing of the layer with the smallest distance from the initial position among the at least one layers.

[0226] For example, the search spacing of the first search region is less than or equal to the search spacing of the layer with the smallest minimum distance from the initial position among the at least one layers.

[0227] For example, the search spacing of the first search region is less than or equal to the search spacing of the layer with the smallest maximum minimum distance from the initial position in the at least one layer.

[0228] For example, starting from the initial position and working outwards, the encoder can divide the at least one co-location position into M layers (numbered 1 to M, with the number decreasing as it gets closer to the initial position). Then, the search spacing of the first search area is less than or equal to the search spacing corresponding to the layer numbered 1.

[0229] In this embodiment, the search spacing of the first search region is equal to the search spacing corresponding to the layer with the smallest distance from the initial position among the at least one layer. This ensures the densest search is applied to the search region in the initial macropixel and the nearest layer, guaranteeing the encoder's motion estimation effect and coding performance. Alternatively, the search spacing of the first search region is less than the search spacing corresponding to the layer with the smallest distance from the initial position among the at least one layer. This ensures the densest search is applied to the search region in the initial macropixel, while a sparser search spacing is applied to the nearest layer. This improves coding efficiency while maintaining the encoder's motion estimation effect and coding performance.

[0230] Figure 13 is an example of an embodiment of this application where the search spacing of the first search region is equal to the search spacing used by the first layer.

[0231] As shown in Figure 13, the search spacing of the first search region can be equal to the search spacing of the layer with the smallest distance from the initial position among the at least one layer. For example, the search spacing of the first search region and the search spacing of the search region where the same position is located on the first layer from the inside out among the at least one layer are both 2, and the search spacing of the search region where the same position is located on the second layer is 4.

[0232] Figure 14 is an example of a first search region having a search spacing smaller than that used by the first layer, provided in an embodiment of this application.

[0233] As shown in Figure 14, the search spacing of the first search region can be less than the search spacing of the layer with the smallest distance from the initial position among the at least one layer. For example, the search spacing of the first search region is 2, the search spacing of the search regions where the same position is located on the first layer from the inside out among the at least one layer is 3, and the search spacing of the search regions where the same position is located on the second layer is 4.

[0234] In some embodiments, the first search area includes a region centered at the initial position and with a side length or radius of a preset value.

[0235] For example, the preset value can be implemented by pre-saving the corresponding code, table or other means that can be used to indicate relevant information in the encoder, or the preset value can be agreed or defined by a standard protocol.

[0236] For example, the first search area includes a circular area centered at the initial position and with a radius equal to the preset value.

[0237] For example, the first search area includes a rectangular area centered at the initial position and with a side length (e.g., the maximum or minimum side length) equal to the preset value.

[0238] For example, the first search area includes a diamond-shaped area centered at the initial position and with a diagonal side length (e.g., the maximum or minimum side length) equal to the preset value.

[0239] In some embodiments, the first search region is a rectangular search region, wherein the maximum side length of the rectangular search region is less than or equal to the minimum value of the horizontal and vertical spacing between two adjacent macro pixels in the reference image.

[0240] For example, the maximum side length of the rectangular search region is less than or equal to the minimum value of the horizontal and vertical spacing between two adjacent macro pixels in the reference image, and is circumscribed by the initial macro pixel.

[0241] The preferred embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the specific details of the embodiments described above. Within the scope of the technical concept of this application, various simple modifications can be made to the technical solutions of this application, and these simple modifications all fall within the protection scope of this application. For example, the various specific technical features described in the specific embodiments described above can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this application will not describe the various possible combinations separately. Furthermore, various different embodiments of this application can also be arbitrarily combined, as long as they do not violate the spirit of this application, they should also be considered as the content disclosed in this application.

[0242] It should also be understood that, in the various method embodiments of this application, the order of the processes mentioned above does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0243] The device embodiments of this application are described in detail below with reference to Figures 15 and 16.

[0244] Figure 15 is a schematic block diagram of the encoder 500 provided in an embodiment of this application.

[0245] As shown in Figure 15, the encoder 500 may include:

[0246] The estimation unit 510 is used to perform motion estimation on the current block in the current image and determine the first motion parameter of the current block; the first motion parameter points to the initial position in the reference image of the current image;

[0247] Determine unit 520, used for:

[0248] Within the first search area where the initial position is located, at least one first position other than the initial position is determined; the first search area is less than or equal to the initial macro pixel where the initial position is located;

[0249] A determining unit is configured to determine a second motion parameter of the reference block based on the initial position and the at least one first position.

[0250] In some embodiments, the determining unit 520 is specifically used for:

[0251] The motion parameter pointing to the position with the minimum rate-distortion cost among the initial position and the at least one first position is determined as the second motion parameter.

[0252] In some embodiments, the determining unit 520 is specifically used for:

[0253] Based on the position of the initial position within the initial macropixel, at least one co-location of the initial position is determined;

[0254] The second motion parameters are determined based on the initial position, the at least one first position, and the at least one co-location position.

[0255] In some embodiments, the determining unit 520 is specifically used for:

[0256] Within the search range of the at least one co-location location, at least one macro pixel is determined;

[0257] Within any one of the at least one macropixels, the position that is the same as the position of the initial position within the initial macropixel is determined as the corresponding position among the at least one corresponding position.

[0258] In some embodiments, the determining unit 520 is specifically used for:

[0259] The motion parameter pointing to the position with the minimum rate-distortion cost among the initial position, the at least one first position, and the at least one corresponding position is determined as the second motion parameter.

[0260] In some embodiments, the determining unit 520 is specifically used for:

[0261] Within a second search region where the first corresponding position is located in the at least one corresponding position, at least one second position other than the first corresponding position is determined; the second search region is less than or equal to the macro pixel where the first corresponding position is located;

[0262] The motion parameter pointing to the position with the minimum rate-distortion cost among the initial position, the at least one first position, the at least one co-location position, and the at least one second position is determined as the second motion parameter.

[0263] In some embodiments, the determining unit 520 is specifically used for:

[0264] A first search interval is determined based on a first distance between the first co-location and the initial location;

[0265] Based on the first search interval, the at least one second position is searched within the second search area using a grid search method.

[0266] In some embodiments, the first distance is positively correlated with the first search interval.

[0267] In some embodiments, the search interval of the first search region is less than or equal to the search interval corresponding to the distance between the closest corresponding position to the initial position and the initial position.

[0268] In some embodiments, the determining unit 520 is specifically used for:

[0269] Divide the at least one homologous site into at least one layer;

[0270] Based on the spacing corresponding to the first layer where the first co-location is located, the at least one second location is searched in the second search area using a grid search method.

[0271] In some embodiments, the determining unit 520 is specifically used for:

[0272] Starting from the initial position, along multiple predefined directions, the same number of co-locations among the at least one co-location positions that are spaced apart from the initial position are divided into co-location positions included in one of the at least one layers.

[0273] In some embodiments, the plurality of directions include directions perpendicular to the edges of the initial macropixel and / or directions parallel to the diagonals of the initial macropixel.

[0274] In some embodiments, the shape formed by connecting the corresponding positions on the at least one layer is the same as or different from the shape of the initial macropixel.

[0275] In some embodiments, the distance between the first layer and the initial position is positively correlated with the spacing corresponding to the first layer.

[0276] In some embodiments, the search spacing of the first search region is less than or equal to the search spacing of the layer with the smallest distance from the initial position among the at least one layers.

[0277] In some embodiments, the range of the second search region is the same as the range of the first search region.

[0278] In some embodiments, the first search area includes a region centered at the initial position and with a side length or radius of a preset value.

[0279] In some embodiments, the first search region is a rectangular search region, wherein the maximum side length of the rectangular search region is less than or equal to the minimum value of the horizontal and vertical spacing between two adjacent macro pixels in the reference image.

[0280] It should be understood that the device embodiments of the encoder and the method embodiments of the encoding method can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, further details are omitted here. Specifically, the encoder 500 shown in FIG15 can correspond to the corresponding subject in the encoding method 400 of the present application embodiments, and the foregoing and other operations and / or functions of each unit in the encoder 500 are respectively for implementing the corresponding processes in the encoding method 400 and other methods.

[0281] It should also be understood that the units in the encoder 500 involved in the embodiments of this application are based on logical functional division. In practical applications, the function of one unit can be implemented by multiple units, or the function of multiple units can be implemented by one unit. Furthermore, these functions can be implemented with the assistance of one or more other units. For example, some or all of the units in the encoder 500 can be merged into one or more additional units. As another example, some units(s) in the encoder 500 can be further divided into multiple functionally smaller units, which can achieve the same operation without affecting the technical effects of the embodiments of this application. Furthermore, the encoder 500 can also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0282] According to another embodiment of this application, the encoder 500 involved in the embodiments of this application, and the encoding method of the embodiments of this application, can be constructed and implemented by running a computer program (including program code) capable of executing the steps involved in the corresponding method on a general-purpose computing device including processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). The computer program can be recorded on, for example, a computer-readable storage medium, loaded into an electronic device through the computer-readable storage medium, and run therein to implement the corresponding method of the embodiments of this application. In other words, the units mentioned above can be implemented in hardware, in software instructions, or in a combination of hardware and software. Specifically, the steps of the method embodiments in the embodiments of this application can be completed by the integrated logic circuits of the hardware in the processor and / or by the instructions in software. The steps of the method disclosed in the embodiments of this application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software in the decoding processor. Optionally, the software can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps in the method embodiments described above.

[0283] Figure 16 is a schematic structural diagram of the electronic device 600 provided in an embodiment of this application.

[0284] As shown in Figure 16, the electronic device 600 includes at least a processor 610 and a computer-readable storage medium 620. The processor 610 and the computer-readable storage medium 620 can be connected via a bus or other means. The computer-readable storage medium 620 stores a computer program 621, which includes computer instructions. The processor 610 executes the computer instructions stored in the computer-readable storage medium 620. The processor 610 is the computing and control core of the electronic device 600, and is suitable for implementing one or more computer instructions, specifically for loading and executing one or more computer instructions to achieve corresponding method flows or corresponding functions.

[0285] For example, processor 610 may also be referred to as a central processing unit (CPU). Processor 610 may include, but is not limited to: general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, discrete hardware components, etc.

[0286] Exemplarily, the computer-readable storage medium 620 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk storage device; optionally, it may also be at least one computer-readable storage medium located remotely from the aforementioned processor 610. Specifically, the computer-readable storage medium 620 includes, but is not limited to, volatile memory and / or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0287] For example, the electronic device 600 may be a decoder or decoding framework involved in the embodiments of this application; the computer-readable storage medium 620 stores computer instructions; the processor 610 loads and executes the computer instructions stored in the computer-readable storage medium 620 to implement the corresponding steps in the decoding method provided in the embodiments of this application; in other words, the computer instructions in the computer-readable storage medium 620 are loaded and executed by the processor 610 to implement the corresponding steps, which will not be described again here to avoid repetition.

[0288] According to another aspect of this application, this application also provides an encoding and decoding system, including the encoder and decoder mentioned above.

[0289] According to another aspect of this application, a computer-readable storage medium (Memory) is also provided. This computer-readable storage medium is a memory device in the electronic device 600 for storing programs and data. For example, a computer-readable storage medium 620. It is understood that the computer-readable storage medium 620 here may include both the built-in storage medium in the electronic device 600 and extended storage media supported by the electronic device 600. The computer-readable storage medium provides storage space that stores the operating system of the electronic device 600. Furthermore, this storage space also stores one or more computer instructions suitable for loading and execution by the processor 610. These computer instructions may be one or more computer programs 621 (including program code).

[0290] According to another aspect of this application, this application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. For example, computer program 621. In this case, the data processing device 600 may be a computer, and the processor 610 reads the computer instructions from the computer-readable storage medium 620. The processor 610 executes the computer instructions, causing the computer to perform the encoding methods provided in the various alternative methods described above. In other words, when implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes of the embodiments of this application are run or the functions of the embodiments of this application are implemented. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0291] According to another aspect of this application, this application also provides a bitstream, which may be a bitstream generated using the encoding method provided in the embodiments of this application.

[0292] Those skilled in the art will recognize that the units and process steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0293] Finally, it should be noted that the above content is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An encoding method, characterized in that, include: Perform motion estimation on the current block in the current image to determine the first motion parameters of the current block; The first motion parameter points to the initial position in the reference image of the current image; Within the first search area where the initial position is located, at least one first position other than the initial position is determined; the first search area is less than or equal to the initial macro pixel where the initial position is located; Based on the initial position and the at least one first position, the second motion parameters of the reference block are determined.

2. The method according to claim 1, characterized in that, Determining the second motion parameters of the reference block based on the initial position and the at least one first position includes: The motion parameter pointing to the position with the minimum rate-distortion cost among the initial position and the at least one first position is determined as the second motion parameter.

3. The method according to claim 1, characterized in that, Determining the second motion parameters of the reference block based on the initial position and the at least one first position includes: Based on the position of the initial position within the initial macropixel, at least one co-location of the initial position is determined; The second motion parameters are determined based on the initial position, the at least one first position, and the at least one co-location position.

4. The method according to claim 3, characterized in that, Determining at least one co-location of the initial position based on its position within the initial macropixel includes: Within the search range of the at least one co-location location, at least one macro pixel is determined; Within any one of the at least one macropixels, the position that is the same as the position of the initial position within the initial macropixel is determined as the corresponding position among the at least one corresponding position.

5. The method according to claim 3 or 4, characterized in that, Determining the second motion parameter based on the initial position, the at least one first position, and the at least one co-location position includes: The motion parameter pointing to the position with the minimum rate-distortion cost among the initial position, the at least one first position, and the at least one corresponding position is determined as the second motion parameter.

6. The method according to claim 3 or 4, characterized in that, Determining the second motion parameter based on the initial position, the at least one first position, and the at least one co-location position includes: Within a second search region where the first corresponding position is located in the at least one corresponding position, at least one second position other than the first corresponding position is determined; the second search region is less than or equal to the macro pixel where the first corresponding position is located; The motion parameter pointing to the position with the minimum rate-distortion cost among the initial position, the at least one first position, the at least one co-location position, and the at least one second position is determined as the second motion parameter.

7. The method according to claim 6, characterized in that, Determining at least one second location other than the first co-location within the second search area where the first co-location is located among the at least one co-location locations includes: A first search interval is determined based on a first distance between the first co-location and the initial location; Based on the first search interval, the at least one second position is searched within the second search area using a grid search method.

8. The method according to claim 7, characterized in that, The first distance is positively correlated with the first search interval.

9. The method according to claim 7, characterized in that, The search interval of the first search area is less than or equal to the search interval corresponding to the distance between the closest corresponding position to the initial position and the initial position.

10. The method according to claim 6, characterized in that, Determining at least one second location other than the first co-location within the second search area where the first co-location is located among the at least one co-location locations includes: Divide the at least one homologous site into at least one layer; Based on the spacing corresponding to the first layer where the first co-location is located, the at least one second location is searched in the second search area using a grid search method.

11. The method according to claim 10, characterized in that, The step of dividing the at least one homologous site into at least one layer includes: Starting from the initial position, along multiple predefined directions, the same number of co-locations among the at least one co-location positions that are spaced apart from the initial position are divided into co-location positions included in one of the at least one layers.

12. The method according to claim 11, characterized in that, The plurality of directions include directions perpendicular to the edges of the initial macropixel and / or directions parallel to the diagonals of the initial macropixel.

13. The method according to claim 11, characterized in that, The shape formed by connecting the corresponding positions on the at least one layer is the same as or different from the shape of the initial macropixel.

14. The method according to any one of claims 10 to 13, characterized in that, The distance between the first layer and the initial position is positively correlated with the spacing corresponding to the first layer.

15. The method according to any one of claims 10 to 14, characterized in that, The search spacing of the first search region is less than or equal to the search spacing of the layer with the smallest distance from the initial position among the at least one layers.

16. The method according to any one of claims 6 to 15, characterized in that, The range of the second search area is the same as the range of the first search area.

17. The method according to any one of claims 1 to 16, characterized in that, The first search area includes a region centered at the initial position with a side length or radius of a preset value.

18. The method according to claim 17, characterized in that, The first search area is a rectangular search area, and the maximum side length of the rectangular search area is less than or equal to the minimum value of the horizontal and vertical spacing between two adjacent macro pixels in the reference image.

19. The method according to any one of claims 1 to 18, characterized in that, The current image and the reference image of the current image are light field images.

20. An encoder, characterized in that, include: The estimation unit is used to perform motion estimation on the current block in the current image and determine the first motion parameters of the current block; The first motion parameter points to the initial position in the reference image of the current image; Determine the unit, used for: Within the first search area where the initial position is located, at least one first position other than the initial position is determined; the first search area is less than or equal to the initial macro pixel where the initial position is located; A determining unit is configured to determine a second motion parameter of the reference block based on the initial position and the at least one first position.

21. An electronic device, characterized in that, include: A processor, adapted to execute computer programs; A computer-readable storage medium storing a computer program that, when executed by the processor, implements the method according to any one of claims 1 to 19.

22. A computer-readable storage medium, characterized in that, Used to store a computer program that, when run on a computer, causes the computer to perform the method according to any one of claims 1 to 19.

23. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method according to any one of claims 1 to 19.

24. A bitstream, characterized in that, The bitstream is a bitstream generated by the method according to any one of claims 1 to 19.