Coding method and device, equipment, code stream and storage medium

CN120958816APending Publication Date: 2025-11-14GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380096022.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-04-16
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In the motion estimation process of existing video coding technology, the search method based on macropixels may lose local optimal points, resulting in a decrease in video image compression performance.

Method used

By searching in the reference image based on multiple adjacent blocks of the current block, the second reference block is determined, thereby improving the motion estimation accuracy and generating more accurate motion information to save code stream bit overhead.

Benefits of technology

It improves the compression performance of video images, reduces the bit overhead of the code stream, and enhances the compression effect of light field videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120958816A_ABST
    Figure CN120958816A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a coding method and device, equipment, a code stream and a storage medium, and the method comprises the steps: searching a second reference block of a current block in a first reference image according to a plurality of adjacent blocks of the first reference block of the current block in the first reference image; determining first motion information of the current block according to the second reference block; and generating a code stream according to the first motion information.
Need to check novelty before this filing date? Find Prior Art

Description

Coding method and device, equipment, code stream, storage medium Technical Field

[0001] The embodiments of the present application relate to video coding technology, including but not limited to coding methods and devices, equipment, code streams, and storage media. Background Art

[0002] Light field images are composed of a series of regularly arranged macropixels. Due to the imaging principles of light field cameras, adjacent macropixels are highly correlated. Therefore, during the motion estimation search process, a macropixel-based search method can fully exploit this correlation in light field images, resulting in more efficient compression performance. However, while this search method takes into account the regular arrangement of macropixels in light field images, searching in macropixel units can miss some local optimal points, thereby reducing the compression performance of the video image.

[0003] Summary of the Invention

[0004] In view of this, the encoding method, apparatus, device, bitstream, and storage medium provided in the embodiments of the present application can improve the motion estimation accuracy of the current block, thereby saving bit overhead of the bitstream and improving the compression performance of the video image. The encoding method, apparatus, device, bitstream, and storage medium provided in the embodiments of the present application are implemented as follows:

[0005] According to one aspect of an embodiment of the present application, a coding method is provided, which is applied to an encoder. The method includes: searching for a second reference block of the current block in a first reference image based on multiple adjacent blocks of the first reference block of the current block in the first reference image; determining first motion information of the current block based on the second reference block; and generating a code stream based on the first motion information.

[0006] According to one aspect of an embodiment of the present application, there is provided an encoding device, applied to an encoder, the device comprising: a first search module, configured to search for a second reference block of the current block in a first reference image based on multiple adjacent blocks of the first reference block of the current block in the first reference image; a second search module, configured to determine first motion information of the current block based on the second reference block; and an encoding module, configured to generate a code stream based on the first motion information.

[0007] According to one aspect of an embodiment of the present application, an encoder is provided, comprising a first memory and a first processor; wherein the first memory is used to store a computer program that can be run on the first processor; and the first processor is used to execute the encoding method as described in the embodiment of the present application when running the computer program.

[0008] According to one aspect of an embodiment of the present application, a code stream is provided. The code stream is generated by the encoding method described in the embodiment of the present application.

[0009] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor adapted to execute a computer program; and a computer-readable storage medium storing a computer program, wherein when the computer program is executed by the processor, the encoding method described in the embodiment of the present application is implemented.

[0010] In an embodiment of the present application, a second reference block of the current block in the first reference image is searched for based on multiple adjacent blocks of the first reference block of the current block in the first reference image, rather than directly determining the first motion information of the current block based on the first reference block; in this way, the motion estimation accuracy of the current block is improved, and the first motion information is made closer to the true value, which is beneficial to saving the bit overhead of the code stream and improving the compression performance of the video image. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings herein are incorporated into and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, serve to illustrate the technical solutions of the present application. Obviously, the drawings described below are merely some embodiments of the present application. Those skilled in the art can, without inventive effort, derive other drawings from these drawings.

[0012] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0013] FIG1 is a schematic diagram of a processing flow of an encoding end of a video encoding and decoding framework provided in an embodiment of the present application;

[0014] FIG2 is a schematic diagram of a processing flow of a decoding end of a video encoding and decoding framework provided in an embodiment of the present application;

[0015] FIG3 is a schematic diagram of a network architecture of a coding and decoding system provided in an embodiment of the present application;

[0016] FIG4 is a schematic diagram of a coding and decoding system provided in an embodiment of the present application;

[0017] FIG5 is a schematic diagram of an implementation flow of the encoding method provided in an embodiment of the present application;

[0018] FIG6 is a schematic diagram of an implementation flow of the encoding method provided in an embodiment of the present application;

[0019] FIG7 is a schematic diagram of pixel-level fine correction of the initial search point of a light field video according to an embodiment of the present application;

[0020] FIG8 is a schematic diagram of a preliminary search at the macro-pixel level for a light field video according to an embodiment of the present application;

[0021] FIG9 is a schematic diagram of a macro-pixel-level fine search for a light field video with a step size of 1 according to an embodiment of the present application;

[0022] FIG10 is a schematic diagram of a macro-pixel-level fine search for a light field video when a step size is greater than a certain threshold according to an embodiment of the present application;

[0023] FIG11A is a schematic diagram of a partial implementation flow of the encoding method provided in an embodiment of the present application;

[0024] FIG11B is a schematic diagram of another implementation flow of the encoding method provided in an embodiment of the present application;

[0025] FIG12 is a schematic diagram of a light field image;

[0026] FIG13 is a schematic diagram of the structure of an encoding device provided in an embodiment of the present application;

[0027] FIG14 is a schematic diagram of the structure of the encoder provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the specific technical solutions of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0030] In the following description, references to “some embodiments,” “this embodiment,” “embodiments of the present application,” and examples, etc., describe a subset of all possible embodiments. However, it can be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.

[0031] The descriptions such as "first, second, third" appearing in the embodiments of this application are only for illustration and distinction of the described objects. There is no order, nor does it indicate any special limitation on the number of devices in the embodiments of this application, and cannot constitute any limitation on the embodiments of this application.

[0032] Most video coding standards use a block-based hybrid coding framework. Each image, sub-image, or frame in a video is divided into square Largest Coding Units (LCUs) or Coding Tree Units (CTUs) of the same size (e.g., 128x128 or 64x64). Each LCU or CTU can be divided into rectangular Coding Units (CUs) according to a rule. Coding Units may also be divided into Prediction Units (PUs) and / or Transform Units (TUs). The hybrid coding framework includes modules such as prediction, transform, quantization, entropy coding, and in-loop filtering. The prediction module includes intra-frame prediction and inter-frame prediction. Inter-frame prediction includes motion estimation and motion compensation. Due to the strong correlation between adjacent pixels in a video image, intra-frame prediction is used in video coding technology to eliminate spatial redundancy between adjacent pixels. Since there is a strong similarity between adjacent images in a video, the inter-frame prediction method is used in video coding and decoding technology to eliminate the temporal redundancy between adjacent images, thereby improving coding and decoding efficiency.

[0033] FIG1 is a schematic diagram of the processing flow of the encoding end of the video codec framework provided by an embodiment of the present application. As shown in FIG1 , a frame of image 101 is divided into blocks. Intra-frame prediction or inter-frame prediction is used for the current block to generate a prediction block for the current block. The prediction block is subtracted from the original block of the current block to obtain a residual block. The residual block is transformed and quantized to obtain a quantization coefficient matrix. The quantization coefficient matrix is ​​entropy encoded and output to the bitstream. As shown in FIG2 , at the decoding end, intra-frame prediction or inter-frame prediction is used for the current block to generate a prediction block / prediction value for the current block. On the other hand, the bitstream is parsed, the residual is transformed and quantized to obtain a quantization coefficient matrix. The quantization coefficient matrix is ​​inversely quantized and inversely transformed to obtain a residual block. The prediction block and the residual block are added to obtain a reconstructed block. The reconstructed blocks form a reconstructed image, and the reconstructed image is loop-filtered based on the image or block to obtain a decoded image. The encoding end also requires similar operations as the decoding end to obtain a decoded image. At the encoding end, the decoded image obtained can be used as a reference image for inter-frame prediction of subsequent images. The block division information, prediction, transformation, quantization, entropy coding, loop filtering, and other mode information or parameter information determined by the encoding end need to be included in the output bitstream if necessary. It can be understood that at the decoding end, the block division information, prediction, transformation, quantization, entropy coding, loop filtering, and other mode information or parameter information determined by the encoding end are the same through parsing and analysis based on existing information, thereby ensuring that the decoded image obtained by the encoding end is the same as the decoded image obtained by the decoding end. The decoded image obtained by the encoding end is also usually called a reconstructed image. During prediction, the current block can be divided into prediction units, and during transformation, the current block can be divided into transformation units. The division of prediction units and transformation units can be different.

[0034] The above is the basic process of the video codec under the block-based hybrid coding framework. As technology develops, some modules or steps of the framework or process may be optimized. The encoding and decoding method provided in the embodiment of the present application is applicable to the basic process of the video codec under the block-based hybrid coding framework, but is not limited to this framework and process. It is known to those skilled in the art that with the evolution of encoders and decoders and the emergence of new business scenarios, the method provided in the embodiment of the present application is also applicable to similar technical problems.

[0035] The current block may be a current coding unit (CU) or a current prediction unit (PU), etc.

[0036] The embodiment of the present application also provides a network architecture of a coding and decoding system including an encoder and a decoder, wherein FIG3 shows a schematic diagram of a network architecture of a coding and decoding system provided by an embodiment of the present application. As shown in FIG3 , the network architecture includes one or more electronic devices 13 to 1N and a communication network 01, wherein the electronic devices 13 to 1N can perform video interaction through the communication network 01. During implementation, the electronic device can be various types of devices with video coding and decoding functions. For example, the electronic device can include a smart phone, a tablet computer, a personal computer, a personal digital assistant, a navigator, a digital phone, a video phone, a television, a sensing device, a server, etc., and the embodiment of the present application is not specifically limited. Here, the decoder or encoder described in the embodiment of the present application can be the above-mentioned electronic device.

[0037] It should be noted that the method of the embodiment of the present application is mainly applied to the inter-frame prediction module shown in Figure 1. When applied to the inter-frame prediction module at the encoding end, the "current block" specifically refers to the coding block currently to be inter-frame predicted.

[0038] The encoding and decoding method provided in the embodiments of the present application is applicable to the encoding and decoding of light field video images / light field images. It can be understood that light field video is a light field video captured by multiple cameras or a multi-camera array, and is currently being studied by the MPEG LVC (Lenslet video coding) working group. Unlike the general camera imaging model, the light field camera adds a set of microlens arrays in front of the imaging plane, so that the light at the same point on the object plane can be captured by multiple microlenses at the same time, which is equivalent to shooting the same point from multiple angles at the same time.

[0039] Due to their unique imaging model, the visual quality of light field images differs significantly from that of traditional images. This results in poor performance for compression methods used for conventional images or videos. The LVC group was established to address this issue, researching compression methods more suitable for light field videos.

[0040] For the encoding end, as shown in Figure 4, the captured / shot light field video can be converted into a data format that meets the codec input requirements, and then encoded into a bit stream through the codec and transmitted to the decoding end. The decoding end parses the bit stream, obtains the light field video in this data format, and then converts it into a light field video that meets the display requirements.

[0041] The embodiment of the present application provides an encoding method, which is applied to an encoder. FIG5 is a schematic diagram of an implementation flow of the encoding method provided in the embodiment of the present application. As shown in FIG5 , the method includes the following steps 501 to 503:

[0042] Step 501: Search and obtain a second reference block of a current block in a first reference image based on a plurality of neighboring blocks of the first reference block of the current block in the first reference image.

[0043] In some embodiments, the encoder may implement step 501 according to steps 603 to 607 described in the following embodiments.

[0044] Step 502: Determine first motion information of the current block according to the second reference block.

[0045] In some embodiments, the encoder may implement step 502 according to steps 608 to 609 as described in the following embodiments. Further, the encoder may implement step 502 according to steps 708 to 724 as described in the following embodiments.

[0046] Step 503: Generate a code stream according to the first motion information.

[0047] In an embodiment of the present application, a second reference block of the current block in the first reference image is searched for based on multiple adjacent blocks of the first reference block of the current block in the first reference image, rather than directly determining the first motion information of the current block based on the first reference block; in this way, the motion estimation accuracy of the current block is improved, and the first motion information is made closer to the true value, which is beneficial to saving the bit overhead of the code stream and improving the compression performance of the video image.

[0048] It should be noted that the encoding method provided in the embodiments of this application is not limited to any specific application scenario and can be used for light field video compression or other types of video compression. In some embodiments, the first reference image can be a light field image, and the current block is a light field image block. In other words, the encoding method provided in the embodiments of this application can be applied to light field video compression.

[0049] As you can understand, light field images are composed of a series of regularly arranged macropixels. Due to the imaging principles of light field cameras, there is a strong correlation between adjacent macropixels. Therefore, during the motion estimation search process, compared to conventional pixel-based methods, the macropixel-based search method can fully utilize the correlation of light field images, resulting in more efficient compression performance.

[0050] However, searching for motion estimation in macropixels can miss some local optima. Considering that the matching block may move across macropixels between different images, and that adjacent macropixels may have disparity and size differences, the optimal reference blocks obtained by the macropixel-based motion estimation search method may not be strictly aligned with the macropixel spacing.

[0051] In view of this, in the application scenario of light field video compression, for the motion estimation search of the current block, a pixel-level motion estimation search is performed in an embodiment of the present application, that is, based on multiple adjacent blocks of the first reference block of the current block in the first reference image, a second reference block of the current block in the first reference image is searched for, and then the first motion information of the current block is determined based on the second reference block; in this way, the first motion information obtained by the search is closer to the true value, which is beneficial to saving the bit overhead of the code stream and improving the compression performance of the video image.

[0052] The embodiment of the present application provides an encoding method. FIG6 is a schematic diagram of an implementation flow of the encoding method provided in the embodiment of the present application. As shown in FIG6 , the method includes the following steps 601 to 610:

[0053] Step 601: Perform motion estimation on the current block according to the motion information candidates of the current block to obtain second motion information of the current block; wherein the second motion information includes a first motion vector and an index value of the first reference image.

[0054] In some embodiments, the encoder may use the Merge mode to construct a Merge list for the current block, which records the motion information candidates for the current block. In other embodiments, the encoder may also use the AMVP mode to construct an AMVP list for the current block, which records the motion information candidates for the current block.

[0055] Step 602: Determine a first reference block of the current block in the first reference image according to the second motion information.

[0056] It can be understood that the second motion information includes the first motion vector (MV) of the current block and the index value of the first reference image. Therefore, the encoder can easily find the block pointed to by the first motion vector in the first reference image (i.e., the first reference block) based on the second motion information.

[0057] Step 603: Execute a first search process with the first reference block as the central block. The first search process includes: searching for a neighboring block with the smallest rate-distortion cost from multiple neighboring blocks of the central block as a first candidate reference block; wherein the size of the neighboring block is equal to the size of the first reference block.

[0058] It can be understood that in the embodiment of the present application, the rate-distortion cost of the adjacent block refers to the rate-distortion cost caused by assuming that the current block is inter-frame predicted and encoded based on the adjacent block.

[0059] Step 604: Using the first candidate reference block currently searched as the center block, perform the first search process to obtain another first candidate reference block.

[0060] Step 605: Determine whether two central blocks in the selected central blocks are the same region block; if so, use the same region block as the second reference block and proceed to step 608; otherwise, proceed to step 606;

[0061] Step 606: Determine whether the number of executions of the first search process is equal to the first number threshold; if so, execute step 607; otherwise, use the first candidate reference block currently searched as the center block and return to execute the first search process;

[0062] In the embodiment of the present application, there is no limitation on the value of the first number threshold, which may be 2, 3, or 4, etc. In short, the first number threshold is a value greater than or equal to 2.

[0063] Step 607 , using the first candidate reference block with the minimum rate-distortion cost among the searched first candidate reference blocks as the second reference block, and proceeding to step 608 ;

[0064] It can be understood that the rate-distortion cost of the first candidate reference block mentioned here refers to the rate-distortion cost caused by assuming that inter-frame prediction and encoding are performed on the current block based on the first candidate reference block.

[0065] In the embodiment of the present application, steps 603 to 607 can be understood as operations of performing pixel-level fine correction on the first reference block. To facilitate understanding of the scheme described in steps 603 to 607, an example is given here, as shown in Figure 7, assuming that all blocks are calibrated using the pixel / sample in the upper left corner of their own blocks (hereinafter referred to as "pixel point"), pixel point "0" is the initial pixel point (which calibrates the first reference block), and the eight surrounding pixel points are searched with it as the center point. These eight pixel points calibrate the eight adjacent blocks of the first reference block, and the adjacent block with the smallest rate-distortion cost is selected from these eight adjacent blocks as the first candidate reference block. Assume that the pixel point calibrated for the first candidate reference block is pixel point "1"; then, the same eight-point search is performed with pixel point "1" as the center point to obtain pixel point "2" (which calibrates another first candidate reference block); and so on to obtain pixel point "3"; finally, the same eight-point search is performed with pixel point "3" as the center point to obtain pixel point "2", which is consistent with the center point selected previously, so the block calibrated by pixel point "2" is the second reference block.

[0066] It can be understood that steps 603 to 607 are operations for performing pixel-level fine correction on the first reference block, that is, the second reference block is the result of the pixel-level fine correction on the first reference block; based on this, the second reference block is further corrected at the macro-pixel level through the following step 608; thus, compared with the motion estimation search based only on the macro-pixel level, the combination of pixel-level fine correction and macro-pixel-level correction is beneficial to reducing the impact of inconsistent macro-pixel spacing in the same light field image on the motion estimation accuracy, thereby improving the motion estimation accuracy, and further saving the bit overhead of the code stream and enhancing the light field video compression performance.

[0067] Step 608 : Using the macropixel where the second reference block is located as a starting macropixel, a second search process is performed to obtain a third reference block with the minimum rate-distortion cost.

[0068] Furthermore, in some embodiments, the second search process is performed with the macropixel where the specific sample of the second reference block is located as a starting macropixel.

[0069] Exemplarily, in some embodiments, the specific sample includes a sample at the upper left corner of the second reference block, for example, the specific sample is located at the upper left corner vertex of the second reference block.

[0070] It can be understood that the rate-distortion cost of the third reference block mentioned here refers to the rate-distortion cost caused by assuming that inter-frame prediction and encoding are performed on the current block based on the third reference block.

[0071] It is understood that the second search process is a macro-pixel-level search operation. In the embodiment of the present application, step 608 may include a preliminary search at the macro-pixel level, or step 608 may include a preliminary search at the macro-pixel level and a fine search at the macro-pixel level based on the preliminary search results.

[0072] Specifically, in some embodiments, the encoder may implement step 608 as follows: using at least one macropixel as a search step, searching for a fourth reference block with the lowest macropixel rate-distortion cost in the first reference image; and determining the third reference block based on the fourth reference block.

[0073] It can be understood that the rate-distortion cost of the fourth reference block mentioned here refers to the rate-distortion cost caused by assuming that inter-frame prediction and encoding are performed on the current block based on the fourth reference block.

[0074] In an embodiment of the present application, when searching for the fourth reference block, there is no limitation on which macropixels to search. In short, the macropixel where the second reference block is located is used as the starting macropixel, and the search is performed on the first reference image according to a search step equal to at least one macropixel spacing.

[0075] Exemplarily, in some embodiments, the fourth reference block can be obtained by the following macropixel-level search: searching for a second candidate reference block having the same size as the second reference block on macropixels on one or more search templates according to one or more search steps; selecting a block with the smallest rate-distortion cost from the second reference block and the second candidate reference block as the fifth reference block; wherein the one or more search steps are multiples of macropixels; and determining the fourth reference block based on the fifth reference block.

[0076] It can be understood that the one or more search steps are multiples of macro pixels, which can be understood as one or more search steps are multiples of the macro pixel spacing.

[0077] The operation / process of searching for the fifth reference block can be understood as a preliminary search at the macropixel level. The one or more search templates described in this search process are not limited. In some embodiments, the one or more search templates can be diamond-shaped templates, square templates, or other search templates. Blocks (the same size as the second reference block) marked by specific points (such as vertices or midpoints) of these templates can be searched to find the fifth reference block.

[0078] Taking the search template as a diamond template as an example, as shown in Figure 8, the second reference block is used as the initial search point, the search step starts from 1 macropixel unit, and increases in the form of an integer power of 2. The search is performed within the specified search range according to the diamond template to obtain the second candidate reference block (i.e., the initial macropixel search), and the block with the smallest rate-distortion cost is selected from the second reference block and the second candidate reference block as the fifth reference block.

[0079] Regarding the determination of the fourth reference block based on the fifth reference block, further, in some embodiments, the fifth reference block can be directly used as the fourth reference block.

[0080] In some other embodiments, a macro-pixel-level fine search may be performed based on the fifth reference block to find the fourth reference block. Specifically, determining the fourth reference block based on the fifth reference block includes:

[0081] When the search step between the fifth reference block and the second reference block is equal to one macropixel, a third candidate reference block with the same size as the fifth reference block is searched from N adjacent macropixels of the macropixel where the fifth reference block is located, with one macropixel as the search step; wherein N is greater than or equal to 1; from the fifth reference block and the third candidate reference block, the block with the smallest rate-distortion cost is selected as the fourth reference block; in this embodiment of the present application, the value of N is not limited and can be any value, for example, N=2.

[0082] When a search step between the fifth reference block and the second reference block is greater than one macropixel and less than or equal to a step threshold, using the fifth reference block as the fourth reference block; wherein the step threshold is greater than one macropixel;

[0083] When the search step between the fifth reference block and the second reference block is greater than a step threshold, a fourth candidate reference block with the same size as the fifth reference block is searched on the remaining macropixels different from the macropixels where the fifth reference block is located within a specific range; wherein the macropixels where the fifth reference block is located are within the specific range; and from the fifth reference block and the fourth candidate reference block, a block with the smallest rate-distortion cost is selected as the fourth reference block.

[0084] In the embodiment of the present application, there is no limitation on the specific range used when searching for the fourth reference block. It can be within the range of 1 macropixel between the macropixel where the fifth reference block is located, or it can be within the range of multiple macropixels between the macropixel where the fifth reference block is located.

[0085] For example, as shown in Figure 9, when the search step between the fifth reference block and the second reference block is equal to one macropixel, the block with the smallest rate-distortion cost is selected from blocks 902 and 903 (i.e., the third candidate reference block) on the two adjacent macropixels of the macropixel where the fifth reference block 901 is located and the fifth reference block 901 as the fourth reference block; and as shown in Figure 10, when the search step between the fifth reference block and the second reference block is greater than the step threshold, with the fifth reference block as the center and one macropixel as the search step, the fourth candidate reference block is searched for on all adjacent pixels of the macropixel where the fifth reference block is located, and the block with the smallest rate-distortion cost is selected from the fifth reference block and the fourth candidate reference block as the fourth reference block.

[0086] In an embodiment of the present application, the encoder may directly use the fourth reference block as the third reference block, or may perform further pixel-level search based on the fourth reference block, etc. For details, see steps 717 to 723 of the following embodiment.

[0087] Step 609: Determine first motion information of the current block according to the third reference block.

[0088] In some embodiments, the encoder may determine the first motion information of the current block based on the motion vector of the current block relative to the third reference block; in other embodiments, the encoder may also determine the first motion information of the current block based on the motion vector of the third reference block relative to the first reference block (that is, the difference between the motion vector of the third reference block relative to the current block and the motion vector of the first reference block relative to the current block).

[0089] Exemplarily, in some embodiments, the first motion information includes: a motion vector difference of the third reference block relative to the first reference block, an index value of the first motion information, and an index value of the first reference image.

[0090] Alternatively, in some other embodiments, the first motion information includes a motion vector difference of the third reference block relative to the first reference block and an index value of the second motion information.

[0091] Step 610: Generate a bitstream according to the first motion information.

[0092] In some embodiments, the encoder further determines an inter-frame prediction value of the current block based on the first motion information, determines a residual value of the current block based on the inter-frame prediction value of the current block and a sample value of the current block; and generates a code stream based on the residual value of the current block.

[0093] Correspondingly, at the decoding end, the decoder parses the bitstream and obtains the residual value of the current block based on the parsing result; the decoder also parses the bitstream to obtain the index value of the second motion information and the motion vector difference, based on which the second motion information (which includes the first motion vector and the index value of the first reference image) is obtained; the decoder determines the third reference block of the current block based on the first motion vector and the motion vector difference (based on which the predicted value of the current block is obtained); the decoder obtains the reconstructed value of the current block based on the residual value of the current block and the predicted value of the current block.

[0094] An embodiment of the present application provides an encoding method. FIG. 11A and FIG. 11B are schematic diagrams of an implementation flow of the encoding method provided in the embodiment of the present application. As shown in FIG. 11A and FIG. 11B , the method includes the following steps 701 to 725, wherein FIG. 11A shows steps 701 to 716, and FIG. 11B shows steps 717 to 725:

[0095] Step 701: performing motion estimation on the current block based on the motion information candidate of the current block to obtain second motion information of the current block; wherein the second motion information includes a first motion vector and an index value of the first reference image;

[0096] Step 702: Determine the first reference block of the current block in the first reference image according to the second motion information;

[0097] Step 703: Using the first reference block as a central block, perform a first search process, wherein the first search process includes: searching for a neighboring block with the minimum rate-distortion cost from a plurality of neighboring blocks of the central block as a first candidate reference block;

[0098] In some embodiments, the size of the neighboring block is equal to the size of the first reference block.

[0099] Step 704: Execute the first search process using the first candidate reference block currently searched as the center block;

[0100] Step 705: determine whether two central blocks in the selected central blocks are the same region block; if so, use the same region block as the second reference block and proceed to step 708; otherwise, execute step 706;

[0101] Step 706: Determine whether the number of executions of the first search process is equal to the first number threshold; if so, execute step 707; otherwise, use the first candidate reference block currently searched as the center block and return to execute the first search process;

[0102] Step 707 , using the first candidate reference block with the minimum rate-distortion cost among the searched first candidate reference blocks as the second reference block, and proceeding to step 708 ;

[0103] Step 708: Using the macropixel where the second reference block is located as a starting macropixel, search for a second candidate reference block having the same macropixel size as the second reference block on one or more search templates according to one or more search steps.

[0104] Step 709: Select a block with the smallest rate-distortion cost from the second reference block and the second candidate reference block as a fifth reference block; wherein the one or more search steps are multiples of macropixels;

[0105] Step 710, determining whether the search step length between the fifth reference block and the second reference block is equal to one macro pixel; if so, executing step 711; otherwise, executing step 713;

[0106] Step 711, using one macropixel as a search step, searching for a third candidate reference block having the same size as the fifth reference block from N adjacent macropixels of the macropixel where the fifth reference block is located, and proceeding to step 712; wherein N is greater than or equal to 1;

[0107] Step 712: Select the block with the smallest rate-distortion cost from the fifth reference block and the third candidate reference block as the fourth reference block, and proceed to step 717;

[0108] Step 713: Determine whether the search step between the fifth reference block and the second reference block is greater than a step threshold; if so, proceed to step 714; otherwise, proceed to step 716; wherein the step threshold is greater than 1 macropixel;

[0109] Step 714: Search for a fourth candidate reference block of the same size as the fifth reference block on the remaining macropixels within the specific range that are different from the macropixel where the fifth reference block is located; and proceed to step 715; wherein the macropixel where the fifth reference block is located is within the specific range;

[0110] Step 715 : Select the block with the smallest rate-distortion cost from the fifth reference block and the fourth candidate reference block as the fourth reference block, and proceed to step 717 .

[0111] Step 716: Use the fifth reference block as the fourth reference block and proceed to step 717; wherein the step threshold is greater than 1.

[0112] Step 717: Using the fourth reference block as the center block, perform the first search process to obtain a new second reference block.

[0113] Step 718: Determine the third reference block based on the new second reference block;

[0114] In some embodiments, as shown in FIG7B , step 718 may be implemented through steps 719 to 723 as follows:

[0115] Step 719: Using the new second reference block as the center block, perform a third search process to obtain a new second reference block; wherein the third search process includes the first search process and the second search process;

[0116] 720, determining whether the new second reference blocks obtained by two adjacent searches in the third search process are the same region block; if so, executing step 721; otherwise, executing step 722;

[0117] Step 721: Use the new second reference block of the same region block as the third reference block, and proceed to step 724;

[0118] Step 722: Determine whether the number of executions of the third search process is equal to the second number threshold; if so, execute step 723; otherwise, use the current new second reference block as the center block and return to step 719;

[0119] In the embodiment of the present application, the value of the second number threshold is not limited and can be 2, 3, or 4. In short, the first number threshold is a value greater than or equal to 2. The first number threshold and the second number threshold can be equal or different.

[0120] Step 723: Use the block with the smallest rate-distortion cost among the new second reference blocks as the third reference block, and proceed to step 724;

[0121] Step 724: determine first motion information of the current block based on the third reference block;

[0122] Step 725: Generate a bitstream according to the first motion information.

[0123] It should be noted that the rate-distortion cost of a block described in the embodiment of the present application refers to the rate-distortion cost caused by assuming that inter-frame prediction and encoding are performed on the current block based on the block.

[0124] In order to improve the compression performance of light field videos, in some embodiments, the arrangement pattern of macro pixels in light field images is used to improve motion estimation. Specifically, based on the TZSearch algorithm of traditional motion estimation, the motion vector search step size is set to a multiple of the light field macro pixel spacing, as shown in the following formula (1):

[0125] In formula (1), d x and d y is the motion vector offset in this iteration (including horizontal and vertical directions), d x , and d y , is the motion vector offset in the previous iteration, ΔstepX and ΔstepY are the number of search steps increased in this iteration, and L x and L y is the step length of each step, for example, the step length is set to the spacing size of each light field macro pixel (divided into spacing in the horizontal direction and the vertical direction).

[0126] This method is based on the fact that light field images are composed of a series of regularly arranged macropixels, as shown in Figure 12. Due to the imaging principle of light field cameras, there is a strong correlation between adjacent macropixels. Therefore, during the motion estimation search process, a macropixel-based search method can more fully utilize the correlation of light field images, resulting in more efficient compression performance.

[0127] Although the technical solution for motion estimation using the arrangement pattern of macropixels in light field images takes into account the macropixel arrangement pattern of light field images, searching only in macropixel units may miss some local optimal points. Considering that matching blocks may move across macropixels between different images, and that there are differences such as parallax and size between adjacent macropixels, the best candidates for matching blocks may not be strictly arranged according to the macropixel spacing. In view of this, in an embodiment of the present application, a pixel-level fine search can be appropriately added to the macropixel search, and the candidate reference block positions obtained based on the macropixel search can be offset and corrected at the pixel level to increase their correlation with the current prediction unit to a greater extent, thereby further optimizing the motion estimation effect of light field video compression.

[0128] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.

[0129] Based on the correlation between macro-pixels in light field images, in order to further refine the results of motion estimation, a fast search method combining macro-pixel level search and pixel level search is proposed in the embodiment of the present application. This method is based on the TZSearch algorithm in HEVC and is improved. The specific process is as follows: Steps (1) to (6):

[0130] Step (1): Adaptive Motion Vector Prediction (AMVP) algorithm based on HEVC

[0131] Method to determine the initial search point.

[0132] Specifically, in some embodiments, the encoder selects the MV with the lowest rate-distortion cost from the candidate predicted motion vectors (MVs) given by AMVP, and uses the position pointed to by the MV as the initial search point;

[0133] Step (2) performs pixel-level fine correction on the initial search point: taking the initial search point obtained in step (1) as the center, perform multiple steps of pixel-level fine search in the neighborhood around the point and perform fine correction on it.

[0134] Specifically, in some embodiments, eight pixels within the neighborhood of the point are searched, and the optimal point with the lowest rate-distortion cost is selected. This optimal point is then used as the new center point, and the same pixel-level fine search within the neighborhood is performed again. If the optimal point selected in a search coincides with the previously selected center point, this optimal point is considered the result of this step. If the number of searches reaches a certain threshold, the search is terminated early, and the optimal point with the lowest rate-distortion cost is used as the result of this step.

[0135] For example, in Figure 7, pixel "0" is the initial pixel, and eight surrounding pixels are searched with it as the center point to obtain the optimal pixel "1". Then, with "1" as the center point, the same eight-point search is performed to obtain "2". And so on, "3" is obtained. Finally, with "3" as the center point, "2" is obtained, which is the same as the center point selected before. Therefore, the motion vector corresponding to pixel "2" is the result of this step.

[0136] Step (3): Based on the search results of step (2), a preliminary search at the macro-pixel level is performed: as shown in FIG8 , the search step starts from 1 unit and increases in the form of an integer power of 2. The search is performed within the specified search range according to the diamond template (or square template), and the search point with the minimum rate-distortion cost is selected as the result of this step. Here, the step unit is set to the macro-pixel spacing of the light field image.

[0137] Step (4): Perform a fine search at the macro-pixel level based on the search results of step (3): If the step size corresponding to the optimal point obtained in step (3) is 1, then as shown in Figure 9, perform two point searches around this point, and select the search point with the minimum rate-distortion cost as the result of this step. The search here is also based on macro-pixels.

[0138] If the step size corresponding to the optimal point obtained in step (3) is greater than a certain threshold, then as shown in Figure 10, a full search is performed within a certain range with this point as the center, and the search point with the minimum rate-distortion cost is selected as the result of this step. The search here is also in macropixels.

[0139] Step (5): perform a conventional pixel-level fine search based on the search results of step (4): take the result point obtained in step (4) as the center point, and perform multiple steps of pixel-level fine search in the neighborhood around the point. The specific method is the same as that described in step (2);

[0140] In step (6), the optimal result point obtained in step (5) is used as the new initial search point, and steps (2) to (5) are repeated. When the result points obtained from two adjacent searches are consistent or the number of searches reaches a certain threshold, the search is stopped, and the motion vector corresponding to the last search result point is the final motion vector.

[0141] In the embodiment of the present application, a conventional pixel-level fine search is added on the basis of the macro-pixel search, which more fully considers the local optimum, further refines the search results, and improves the motion estimation performance;

[0142] In the embodiment of the present application, it is difficult to ensure strict consistency between the spacing between macropixels on the light field image. The introduction of pixel-level fine search improves the robustness of this. That is, even if the spacing between macropixels on the light field image is different, a good motion estimation result can still be obtained, which is beneficial to saving bit overhead of the code stream.

[0143] It should be noted that although the steps of the method of the present application are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all steps must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps; or steps in different embodiments may be combined to form a new technical solution.

[0144] Based on the above embodiments, an embodiment of the present application provides an encoding device, which is applied to an encoder. FIG13 is a schematic structural diagram of the encoding device provided in an embodiment of the present application. As shown in FIG13 , the encoding device 13 includes:

[0145] A first search module 131 is configured to search for a second reference block of the current block in the first reference image based on a plurality of neighboring blocks of the first reference block of the current block in the first reference image;

[0146] A second search module 132 is configured to determine first motion information of the current block based on the second reference block;

[0147] The encoding module 133 is configured to generate a code stream according to the first motion information.

[0148] In some embodiments, the first reference image is a light field image, and the current block is a light field image block.

[0149] In some embodiments, the first search module 131 is configured to: perform a first search process with the first reference block as the central block, the first search process including: searching for an adjacent block with the smallest rate-distortion cost from multiple adjacent blocks of the central block as a first candidate reference block; perform the first search process with the first candidate reference block currently searched as the central block; and when there are two central blocks in the selected central blocks that are the same area block, the same area block is the second reference block.

[0150] In some embodiments, the first search module 131 is further configured to: when there are no two center blocks in the selected center blocks that are the same area block, perform the first search process with the first candidate reference block currently searched as the center block; when two first candidate reference blocks in the searched first candidate reference blocks are the same area block, the same area block is the second reference block.

[0151] In some embodiments, the first search module 131 is further configured to: when there are no two central blocks in the selected central blocks that are the same area blocks and the number of executions of the first search process is equal to the first number threshold, the first candidate reference block with the smallest rate-distortion cost among the first candidate reference blocks searched is the second reference block.

[0152] In some embodiments, the first search module 131 is further configured to: when there are no two central blocks in the selected central blocks that are the same area block and the number of executions of the first search process is less than the first number threshold, use the first candidate reference block currently searched as the central block and iteratively execute the first search process until the second reference block is searched out or the number of executions of the first search process is equal to the first number threshold.

[0153] In some embodiments, the size of the neighboring block is equal to the size of the first reference block.

[0154] In some embodiments, the second search module 132 is configured to perform a second search process with the macropixel where the second reference block is located as the starting macropixel, to search for a third reference block with the minimum rate-distortion cost; and determine the first motion information of the current block based on the third reference block.

[0155] In some embodiments, the second search process includes: using at least one macropixel as a search step, searching for a fourth reference block with the minimum macropixel rate-distortion cost in the first reference image; and determining the third reference block based on the fourth reference block.

[0156] In some embodiments, the method of searching for a fourth reference block with the minimum rate-distortion cost on macropixels in the first reference image using at least one macropixel as a search step size includes: searching for a second candidate reference block with the same size as the second reference block on macropixels on one or more search templates according to one or more search steps; selecting a block with the minimum rate-distortion cost from the second reference block and the second candidate reference blocks as the fifth reference block; wherein the one or more search steps are multiples of macropixels; and determining the fourth reference block based on the fifth reference block.

[0157] In some embodiments, determining the fourth reference block based on the fifth reference block includes: when the search step between the fifth reference block and the second reference block is equal to one macropixel, using one macropixel as the search step, searching for a third candidate reference block with the same size as the fifth reference block from N adjacent macropixels of the macropixel where the fifth reference block is located; wherein N is greater than or equal to 1; and selecting a block with the smallest rate-distortion cost from the fifth reference block and the third candidate reference block as the fourth reference block.

[0158] In some embodiments, the step of searching for a fourth reference block with the lowest macropixel rate-distortion cost in the first reference image with at least one macropixel as the search step includes: when the search step between the fifth reference block and the second reference block is greater than 1 macropixel and less than or equal to a step threshold, using the fifth reference block as the fourth reference block; wherein the step threshold is greater than 1 macropixel.

[0159] In some embodiments, the step of searching for a fourth reference block with the minimum rate-distortion cost on macropixels in the first reference image using at least one macropixel as a search step includes: when the search step between the fifth reference block and the second reference block is greater than a step threshold, searching for a fourth candidate reference block with the same size as the fifth reference block on the remaining macropixels that are different from the macropixels where the fifth reference block is located within a specific range; wherein the macropixels where the fifth reference block is located are within the specific range; and selecting the block with the minimum rate-distortion cost from the fifth reference block and the fourth candidate reference block as the fourth reference block.

[0160] In some embodiments, a fourth candidate reference block having the same size as the fifth reference block is searched for on all neighboring macropixels of the macropixel where the fifth reference block is located.

[0161] In some embodiments, determining the third reference block based on the fourth reference block includes: performing the first search process with the fourth reference block as the center block to obtain a new second reference block; and determining the third reference block based on the new second reference block.

[0162] In some embodiments, determining the third reference block based on the new second reference block includes: iteratively performing a third search process with the new second reference block as the center block, the third search process including the first search process and the second search process, until the new second reference blocks obtained by two adjacent searches of the third search process are the same regional block, and the new second reference block that is the same regional block is used as the third reference block; or, until the number of executions of the third search process is equal to a second number threshold, the block with the smallest rate-distortion cost in the new second reference block is used as the third reference block.

[0163] In some embodiments, the encoding device 13 also includes a motion estimation module and a determination module; wherein, the motion estimation module is configured to perform motion estimation on the current block based on the motion information candidates of the current block before searching for the second reference block based on multiple adjacent blocks of the first reference block of the current block in the first reference image, so as to obtain the second motion information of the current block; wherein, the second motion information includes a first motion vector and an index value of the first reference image; and the determination module is configured to determine the first reference block of the current block in the first reference image based on the second motion information.

[0164] In some embodiments, the first motion information includes a motion vector difference of the third reference block relative to the first reference block, an index value of the first motion vector, and an index value of the first reference image.

[0165] Alternatively, in some other embodiments, the first motion information includes a motion vector difference of the third reference block relative to the first reference block and an index value of the second motion information.

[0166] The description of the above encoding device embodiment is similar to the description of the above encoding method embodiment, and has similar beneficial effects as the encoding method embodiment. For technical details not disclosed in the encoding device embodiment of this application, please refer to the description of the encoding method embodiment of this application for understanding.

[0167] It should be noted that the division of modules in the device shown in Figure 13 in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation. In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or they can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. It can also be implemented in the form of a combination of software and hardware.

[0168] It should be noted that, in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling an electronic device to execute all or part of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0169] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, the encoding method or the decoding method as described in the embodiment of the present application is implemented.

[0170] The embodiment of the present application provides an encoder, as shown in FIG14 , the encoder 14 includes: a communication interface 141, a memory 142, and a processor 143; each component is coupled together via a bus system 144. It is understood that the bus system 144 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 144 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in FIG14 , various buses are labeled as the bus system 144. Among them,

[0171] Communication interface 141, used for receiving and sending signals during the process of sending and receiving information with other external network elements;

[0172] Memory 142, for storing computer programs that can be run on processor 143;

[0173] The processor 143 is configured to execute the encoding method described in the embodiment of the present application when running the computer program.

[0174] It is understood that the memory 142 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 142 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0175] The processor 143 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 143 or by software instructions. The processor 143 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or the like. The storage medium is located in the memory 142 , and the processor 143 reads the information in the memory 142 and completes the steps of the above method in combination with its hardware.

[0176] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0177] Optionally, as another embodiment, the processor 143 is further configured to execute any of the aforementioned encoding method embodiments when running the computer program.

[0178] The embodiment of the present application further provides a code stream, which is obtained by using the aforementioned encoding method.

[0179] An embodiment of the present application provides an electronic device, comprising: a processor adapted to execute a computer program; and a computer-readable storage medium storing the computer program, wherein when the computer program is executed by the processor, the encoding method and / or decoding method described in the embodiment of the present application are implemented. The electronic device can be any type of device capable of video encoding and / or video decoding, such as a mobile phone, tablet computer, laptop computer, personal computer, television, projection device, or monitoring device.

[0180] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0181] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments. The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other. For the sake of brevity, they will not be repeated here.

[0182] The term "and / or" in this article is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, object A and / or object B can mean: object A exists alone, object A and object B exist at the same time, and object B exists alone.

[0183] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0184] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.

[0185] The modules described above as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules; they may be located in one place or distributed across multiple network units; some or all of the modules may be selected according to actual needs to achieve the purpose of this embodiment.

[0186] In addition, all functional modules in the embodiments of the present application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the above-mentioned integrated modules can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0187] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.

[0188] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks or optical disks.

[0189] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0190] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0191] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0192] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A coding method, the method being applied to an encoder, the method comprising: Searching and obtaining a second reference block of the current block in the first reference image according to a plurality of neighboring blocks of the first reference block of the current block in the first reference image; Determining first motion information of the current block according to the second reference block; Generate a code stream according to the first motion information.

2. The method according to claim 1, wherein: The first reference image is a light field image, and the current block is a light field image block.

3. The method according to claim 1 or 2, wherein: The step of searching for a second reference block of the current block in the first reference image according to a plurality of neighboring blocks of the first reference block of the current block in the first reference image comprises: Taking the first reference block as the central block, performing a first search process, the first search process comprising: searching for an adjacent block with the minimum rate distortion cost from a plurality of adjacent blocks of the central block as a first candidate reference block; Taking the first candidate reference block currently searched as the center block, executing the first search process; In the case that two central blocks among the selected central blocks are the same region block, the same region block is the second reference block.

4. The method according to claim 3, wherein: The method further comprises: In the case that no two central blocks in the searched central blocks are the same region block, the first candidate reference block currently searched is used as the central block to perform the first search process; In the case that two central blocks among the selected central blocks are the same region block, the same region block is the second reference block.

5. The method according to claim 3 or 4, wherein: When there are no two central blocks in the selected central blocks that are the same area blocks and the number of executions of the first search process is equal to the first number threshold, the first candidate reference block with the smallest rate-distortion cost among the searched first candidate reference blocks is the second reference block.

6. The method according to claim 5, wherein: The method further comprises: When there are no two central blocks in the selected central blocks that are the same area blocks and the number of executions of the first search process is less than the first number threshold, the first candidate reference block currently searched is used as the central block, and the first search process is iteratively executed until the second reference block is searched out or the number of executions of the first search process is equal to the first number threshold.

7. The method according to any one of claims 1 to 6, wherein: The size of the neighboring block is equal to the size of the first reference block.

8. The method according to any one of claims 3 to 6, wherein: The determining, according to the second reference block, first motion information of the current block includes: Taking the macro pixel where the second reference block is located as the starting macro pixel, performing a second search process to search for a third reference block with the minimum rate-distortion cost; The first motion information of the current block is determined according to the third reference block.

9. The method according to claim 8, wherein: The step of performing a second search process with the macro pixel where the second reference block is located as a starting macro pixel includes: The second search process is performed with the macro pixel where the specific sample of the second reference block is located as the starting macro pixel.

10. The method according to claim 9, wherein: The specific sample is located at the top left vertex in the second reference block.

11. The method according to any one of claims 8 to 10, wherein: The second search process comprises: Using at least one macro pixel as a search step, searching the first reference image for a fourth reference block having a minimum rate-distortion cost on the macro pixel; The third reference block is determined according to the fourth reference block.

12. The method according to claim 11, wherein: The step of searching the first reference image for a fourth reference block having the minimum rate-distortion cost on a macro pixel by using at least one macro pixel as a search step size comprises: Searching for a second candidate reference block having the same size as the second reference block in macropixels on one or more search templates according to one or more search steps; Selecting a block with the smallest rate-distortion cost from the second reference block and the second candidate reference block as a fifth reference block; wherein the one or more search steps are multiples of macro pixels; The fourth reference block is determined based on the fifth reference block.

13. The method according to claim 12, wherein: The step of determining the fourth reference block according to the fifth reference block comprises: When the search step length between the fifth reference block and the second reference block is equal to one macropixel, searching for a third candidate reference block having the same size as the fifth reference block from N adjacent macropixels of the macropixel where the fifth reference block is located, with one macropixel as the search step length; wherein N is greater than or equal to 1; From the fifth reference block and the third candidate reference block, a block with the smallest rate-distortion cost is selected as the fourth reference block.

14. The method according to claim 12, wherein: The method further comprises: When the search step between the fifth reference block and the second reference block is greater than 1 macropixel and less than or equal to a step threshold, the fifth reference block is used as the fourth reference block; wherein the step threshold is greater than 1 macropixel.

15. The method according to claim 12, wherein: The method further comprises: When the search step length between the fifth reference block and the second reference block is greater than a step length threshold, searching for a fourth candidate reference block having the same size as the fifth reference block on remaining macro pixels different from the macro pixels where the fifth reference block is located within a specific range; wherein the macro pixels where the fifth reference block is located are within the specific range; From the fifth reference block and the fourth candidate reference block, a block with the smallest rate-distortion cost is selected as the fourth reference block.

16. The method according to claim 15, wherein: A fourth candidate reference block having the same size as the fifth reference block is searched on all neighboring macro pixels of the macro pixel where the fifth reference block is located.

17. The method according to claim 13, 15 or 16, wherein: The macro-pixel where the fifth reference block is located refers to the macro-pixel where the specific sample of the fifth reference block is located.

18. The method according to claim 17, wherein: The specific sample is located at the top left vertex in the fifth reference block.

19. The method according to claim 11, wherein: The step of determining the third reference block according to the fourth reference block comprises: Taking the fourth reference block as the central block, executing the first search process to obtain a new second reference block; The third reference block is determined according to the new second reference block.

20. The method according to claim 19, wherein: The step of determining the third reference block according to the new second reference block includes: Taking the new second reference block as the central block, iteratively execute a third search process, wherein the third search process includes the first search process and the second search process, until the new second reference blocks obtained by two adjacent searches of the third search process are the same regional block, and the new second reference block that is the same regional block is used as the third reference block; or, until the number of executions of the third search process is equal to the second number threshold, the block with the smallest rate-distortion cost in the new second reference block is used as the third reference block.

21. The method according to claim 8, wherein: Before searching for the second reference block according to a plurality of neighboring blocks of the first reference block of the current block in the first reference image, the method further includes: According to the motion information candidate of the current block, motion estimation is performed on the current block to obtain second motion information of the current block; wherein the second motion information includes a first motion vector and an index value of the first reference image; Determine the first reference block of the current block in the first reference image according to the second motion information.

22. The method according to claim 21, wherein: The first motion information includes a motion vector difference of the third reference block relative to the first reference block and an index value of the second motion information.

23. An encoding device, applied to an encoder, the device comprising: A first search module, configured to search for a second reference block of the current block in the first reference image based on a plurality of neighboring blocks of the first reference block of the current block in the first reference image; A second search module, configured to determine first motion information of the current block according to the second reference block; The encoding module is configured to generate a code stream according to the first motion information.

24. An encoder comprising a first memory and a first processor; wherein: The first memory is used to store a computer program that can be run on the first processor; The first processor is configured to execute the method according to any one of claims 1 to 22 when running the computer program.

25. A code stream, wherein the code stream is generated by the encoding method according to any one of claims 1 to 22.

26. An electronic device comprising: a processor adapted to execute a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the method according to any one of claims 1 to 22 is implemented.

27. A computer-readable storage medium, wherein: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 22 is implemented.