Inter-frame prediction method, apparatus and system

By obtaining the mapping relationship between the current frame and the reference frame in the video encoding and decoding system, and utilizing the initial motion information and image distortion processing, the problem of low efficiency of motion vector search in the existing technology is solved, and more efficient and accurate inter-frame prediction is achieved, reducing the amount of data.

WO2025209084A1PCT designated stage Publication Date: 2025-10-09HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2025/080334
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-30
Filing Date
2025-03-03
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

In the prior art, the processing device is slow and difficult in searching for motion vectors on a reference image, resulting in poor inter-frame prediction accuracy, low coding efficiency, and large amounts of coded and transmitted data.

Method used

By obtaining the mapping relationship between the current frame and the reference frame, the initial motion information is used to find the mapping image block in the reference frame that has a mapping relationship with the current image block, and image warping is performed to improve the accuracy and efficiency of motion information. The homography matrix is ​​used to describe the complex transformation relationship, and the mapping relationship is determined using camera motion parameters and intrinsic parameters.

Benefits of technology

It improves the efficiency and accuracy of motion vector search, reduces the amount of encoding and transmission data, improves the accuracy and coding efficiency of inter-frame prediction, and is suitable for global motion and complex irregular motion scenes in video encoding and decoding systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025080334_09102025_PF_FP_ABST
    Figure CN2025080334_09102025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are an inter-frame prediction method, apparatus and system, relating to the field of video / image encoding and decoding. The method comprises: acquiring a mapping relationship between a current frame and a reference frame of the current frame; determining initial motion information on the basis of a current image block in the current frame and the mapping relationship, the initial motion information being used for indicating motion information between the current image block and a mapped image block, in the reference frame, having a mapping relationship with the current image block; further determining target motion information of the current image block on the basis of the initial motion information; and then determining predicted pixels of the current image block on the basis of the target motion information. The method provides a more efficient and less complex process to determine the target motion information, resulting in higher accuracy of the obtained motion information, and therefore, the method has higher prediction accuracy and higher encoding / decoding efficiency, and achieves smaller data volumes for both encoding and transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Inter-frame prediction method, device and system

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on March 30, 2024, with application number 202410385854.X and application name “Method, device and system for inter-frame prediction”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of video coding and decoding technology, and in particular to a method, device and system for inter-frame prediction. Background Art

[0003] In video coding and decoding technology, inter-frame prediction is generally achieved using a block-based prediction method. In this method, starting with a zero motion vector or a motion vector inherited from surrounding blocks, a search window is used to search the reference image for the current block's motion vector. The resulting motion vector is then used to find a matching reference block in the reference image. This reference block is then used to determine the predicted image block for the current block. However, due to limitations in the processing power of the processing device, the range of the reference image search that the processing device can perform at a time is limited.

[0004] This results in slower motion vector searches, leading to lower coding efficiency. Furthermore, it can make motion vector searches more difficult, making it difficult to find a reference image block that matches the current image block. This leads to lower inter-frame references, poor inter-frame prediction accuracy, and larger amounts of encoded and transmitted data. Summary of the Invention

[0005] The present application provides a method, device and system for inter-frame prediction, which can solve the problems of slow efficiency and high difficulty of searching motion vectors by processing equipment, resulting in poor effect of searching motion vectors and thus low coding efficiency, poor inter-frame prediction accuracy, and large amount of encoded data and transmission number.

[0006] In a first aspect, an inter-frame prediction method is provided, which obtains a mapping relationship between a current frame and a reference frame of the current frame, and determines initial motion information based on a current image block in the current frame and the mapping relationship. The initial motion information is used to indicate motion information between the current image block and a mapping image block in the reference frame that has a mapping relationship with the current image block, and then determines target motion information of the current image block based on the initial motion information, and then determines predicted pixels of the current image block based on the target motion information.

[0007] The image content of the reference frame is of reference value to the current frame. For example, the image content in the current frame and the reference frame may have a linear transformation and / or perspective transformation (projection transformation) relationship, and thus a mapping relationship exists between the current frame and the reference frame. Based on this mapping relationship, for the current image block (i.e., the current block), a mapping image block (i.e., the mapping block) with which it has a mapping relationship can be found in the reference frame. This mapping block has a high reference value to the current block. For example, the image block obtained by transforming the image content of the mapping block (e.g., affine transformation, projection transformation, etc.) is close to / similar to the content of the current block.

[0008] After obtaining the motion information between the current block and the mapped block, i.e., the initial motion information, based on the mapping relationship, the method of the present application directly obtains the accurate target motion information of the current block (such as a motion vector, a motion field, which can be used to indicate the image block in the reference frame that matches the current block, i.e., the motion information between the reference block and the current block) based on the initial motion information, and then determines the predicted pixels of the current block based on the target motion information. Compared with the method of using a zero motion vector or a motion vector inherited from an image block surrounding the current block as a starting point and searching for motion vectors in the reference frame with a smaller search range each time, the method of determining the target motion information based on the initial motion information in the method of the present application is more efficient, less complex, and the obtained motion information is more accurate, thereby making the prediction accuracy of the prediction method of the present application higher, the encoding / decoding efficiency higher, and the amount of encoded data (which can be represented by the encoding bit rate) and the amount of transmitted data smaller.

[0009] In one possible implementation, the method maps the original pixel positions of the pixels in the current block according to the mapping relationship to obtain the mapped pixel positions, and further determines the motion information between the current block and the mapped block, i.e., the initial motion information, based on the position change between the mapped pixel positions and the original pixel positions.

[0010] By directly mapping the pixels in the current frame to the reference frame based on the mapping relationship, pixels similar to / corresponding to the current block pixels (i.e., pixels at the mapped pixel positions) can be found in the reference frame efficiently and accurately, thereby efficiently and accurately determining the initial motion information based on the position change between the mapped pixel positions and the original pixel positions.

[0011] In another possible implementation, in the method, target motion information of the current image block is determined by performing a motion search in a reference frame based on the current image block and using the initial motion information as an initial value for the motion search.

[0012] The mapped block can be relatively close to the image block in the reference frame that matches the current block (i.e., the reference block), so that the initial motion information can be close to the accurate target motion information of the current block. By performing a motion search (i.e., searching for motion vectors) in the reference frame using the initial motion information as the initial value for the motion search, accurate target motion information can be found more quickly. This improves the efficiency and accuracy of the motion vector search compared to searching for motion vectors in the reference frame using a zero motion vector or a motion vector inherited from image blocks surrounding the current block as the initial value (i.e., starting point) for the motion search.

[0013] In another possible implementation, in the method, image warping is performed on a mapping image block in a reference frame according to initial motion information to obtain a warped mapping image block; a first predicted pixel of a current image block is determined based on the warped mapping image block; a first reference image block matching the current image block is determined based on target motion information, wherein the first reference image block is an image block in the reference frame, and a positional relationship between the current image block and the first reference image block is a positional relationship indicated by the target motion information; a second predicted pixel of the current image block is determined based on the first reference image block; and a predicted pixel of the current image block is determined based on the first predicted pixel and the second predicted pixel.

[0014] The image block obtained by warping the image content of the mapped block based on the initial motion information is close to / similar to the content of the current block. Therefore, the warped mapped image block serves as a reference for the current image block. By determining a first reference image block based on the target motion information and deriving predicted pixels for the current block based on the first reference image block, and then fusing the predicted pixels for the current block derived based on the warped mapped image block, more accurate predicted pixels can be obtained, improving prediction performance and, in turn, enhancing video encoding and decoding performance.

[0015] In another possible implementation, in the method, image warping is performed on the mapped image block in the reference frame according to the initial motion information to obtain a warped reference frame; and motion search is performed in the warped reference frame according to the current image block to determine the target motion information.

[0016] The image content of the mapped block is warped based on the initial motion information, resulting in a warped image block that is close to / similar to the content of the current block. Consequently, after warping one or more image blocks of the current frame mapped onto one or more mapping blocks in the reference frame, the resulting warped reference frame can be aligned with the current frame. This means that the difference (e.g., object displacement) between the warped reference frame and the current frame can be minimized. Consequently, performing a motion search, i.e., searching for target motion information of the current image block, on the warped reference frame can be more efficient and improve the accuracy of the target motion information found.

[0017] In another possible implementation, in the method, a second reference image block matching the current image block is determined based on the target motion information, where the second reference image block is an image block in the distorted reference frame, and the positional relationship between the current image block and the second reference image block is the positional relationship indicated by the target motion information; and based on the second reference image block, predicted pixels of the current image block are determined.

[0018] Based on the method of performing motion search on the distorted reference frame to obtain the target motion information of the current block, a second reference image block matching the current block is further found in the distorted reference frame according to the target motion information, so that accurate predicted pixels of the current block can be obtained according to the second reference image block.

[0019] In another possible implementation, in the method, the mapping relationship includes a homography matrix between the current frame and the reference frame.

[0020] The homography matrix can effectively describe the mapping relationship between two images with complex transformation relationships (such as projective transformations). Therefore, implementing inter-frame prediction methods based on the homography matrix can make the target motion information determined in the method more accurate and achieve better prediction results.

[0021] In another possible implementation, in the method, camera motion parameters and camera intrinsic parameters between the current frame captured by the camera and the reference frame captured by the camera are obtained, and then the mapping relationship between the current frame and the reference frame is determined according to the camera motion parameters and the camera intrinsic parameters.

[0022] By determining the mapping relationship between the current frame and the reference frame based on camera motion parameters and camera intrinsic parameters, the mapping relationship between the current frame and the reference frame can be obtained relatively simply and accurately, reducing the computational complexity of the mapping relationship. As a result, the computational efficiency of the encoder / decoder when applying this prediction method can also be improved.

[0023] In another possible implementation, in the method, identification information of the current frame is obtained; the mapping relationship is obtained based on the identification information of the current frame, the identification information of the current frame includes the mapping relationship or camera motion parameters and camera intrinsic parameters between the current frame captured by the camera and the reference frame captured by the camera, and the camera motion parameters and the camera memory are used to determine the mapping relationship.

[0024] The method of this application can include the identification information of the current frame in the video bitstream. This identification information includes the mapping relationship between the current frame and the reference frame, or the camera motion parameters and camera internal parameters used to calculate the mapping relationship. The decoder can then implement the inter-frame prediction method of this application using the obtained identification information. As a result, the motion information of the current frame does not need to be transmitted in the video bitstream, or the transmitted motion information parameters are reduced, thereby reducing the amount of transmitted data.

[0025] In a second aspect, a coding and decoding device is provided, including an acquisition module and a prediction module, wherein the acquisition module is used to obtain a mapping relationship between a current frame and a reference frame of the current frame; the prediction module is used to determine initial motion information based on a current image block in the current frame and the mapping relationship, the initial motion information being used to indicate motion information between the current image block and a mapping image block that has a mapping relationship with the current image block in the reference frame; the prediction module is also used to determine target motion information of the current image block based on the initial motion information; the prediction module is also used to determine predicted pixels of the current image block based on the target motion information.

[0026] In one possible implementation, the prediction module is further used to map the original pixel positions of the pixels in the current block according to the mapping relationship to obtain the mapped pixel positions, and further determine the motion information between the current block and the mapped block, i.e., the initial motion information, based on the position change between the mapped pixel positions and the original pixel positions.

[0027] In another possible implementation, the prediction module is further configured to perform a motion search in a reference frame based on the current image block using the initial motion information as an initial value for the motion search, and determine target motion information of the current image block.

[0028] In another possible implementation, the prediction module is further used to perform image warping on the mapped image block in the reference frame according to the initial motion information to obtain a warped mapped image block; determine a first predicted pixel of the current image block according to the warped mapped image block; determine a first reference image block that matches the current image block according to the target motion information, wherein the first reference image block is an image block in the reference frame, and a positional relationship between the current image block and the first reference image block is a positional relationship indicated by the target motion information; determine a second predicted pixel of the current image block according to the first reference image block; and determine a predicted pixel of the current image block according to the first predicted pixel and the second predicted pixel.

[0029] In another possible implementation, the prediction module is further configured to perform image warping on the mapped image block in the reference frame according to the initial motion information to obtain a warped reference frame; and to perform motion search in the warped reference frame according to the current image block to determine target motion information.

[0030] In another possible implementation, the prediction module is further used to determine a second reference image block that matches the current image block based on the target motion information, where the second reference image block is an image block in the distorted reference frame, and the positional relationship between the current image block and the second reference image block is the positional relationship indicated by the target motion information; based on the second reference image block, the predicted pixels of the current image block are determined.

[0031] In another possible implementation, the mapping relationship includes a homography matrix between the current frame and the reference frame.

[0032] In another possible implementation, the acquisition module is further used to obtain camera motion parameters and camera intrinsic parameters between the current frame captured by the camera and the reference frame captured by the camera, and then determine the mapping relationship between the current frame and the reference frame based on the camera motion parameters and camera intrinsic parameters.

[0033] In another possible implementation, in the method, the acquisition module is further used to obtain identification information of the current frame; the mapping relationship is obtained based on the identification information of the current frame, the identification information of the current frame includes the mapping relationship or the camera motion parameters and camera intrinsic parameters between the current frame captured by the camera and the reference frame captured by the camera, and the camera motion parameters and the camera memory are used to determine the mapping relationship.

[0034] According to a third aspect, a decoder is provided, comprising at least one processor and a memory, wherein the memory is used to store a computer program so that when the computer program is executed by the at least one processor, the inter-frame prediction method as described in the first aspect is implemented.

[0035] In a fourth aspect, an encoder is provided, comprising at least one processor and a memory, wherein the memory is used to store a computer program, so that when the computer program is executed by the at least one processor, the inter-frame prediction method as described in the first aspect is implemented.

[0036] In a fifth aspect, a coding and decoding system is provided, which includes the encoder as described in the fourth aspect and the decoder as described in the third aspect, and the encoder and the decoder are used to perform the operating steps of the inter-frame prediction method as described in the first aspect.

[0037] In a sixth aspect, a computing device is provided, comprising: one or more processors, a memory, and a communication interface. The memory and the communication interface are connected to the one or more processors; the computing device communicates with other devices via the communication interface; the memory is used to store computer program code, which includes instructions. When the one or more processors execute the instructions, the computing device performs the inter-frame prediction method described in the first aspect.

[0038] In a seventh aspect, a computer program product comprising instructions is provided. When the instructions are executed by a computer device, the computer device executes the inter-frame prediction method as described in the first aspect.

[0039] In an eighth aspect, a computer-readable storage medium is provided, comprising instructions, which, when executed on a computer device, enable the computer device to perform the inter-frame prediction method as described in the first aspect.

[0040] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods.

[0041] The following description includes more details about the implementation methods provided by the above aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] FIG1 is a schematic diagram of motion information of a block provided by an embodiment of the present application;

[0043] FIG2A is a first schematic diagram of an affine motion model of a block provided in an embodiment of the present application;

[0044] FIG2B is a second schematic diagram of an affine motion model of a block provided in an embodiment of the present application;

[0045] FIG3 is a schematic diagram of the structure of a video encoding and decoding system provided in an embodiment of the present application;

[0046] FIG4 is a schematic diagram of a flow chart of an inter-frame prediction method provided in an embodiment of the present application;

[0047] FIG5 is a schematic diagram of pixel mapping provided in an embodiment of the present application;

[0048] FIG6 is a schematic diagram of dividing an image block into sub-blocks according to an embodiment of the present application;

[0049] FIG7 is a schematic diagram of searching for a motion vector according to an embodiment of the present application;

[0050] FIG8 is a schematic diagram of the structure of an inter-frame prediction apparatus provided in an embodiment of the present application;

[0051] FIG9 is a schematic structural diagram of an encoding device provided in an embodiment of the present application;

[0052] FIG10 is a schematic diagram of the structure of a decoding device provided in an embodiment of the present application;

[0053] FIG11 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] To facilitate understanding, the terms / concepts involved in this application are first introduced.

[0055] Image encoding: The process of compressing an image sequence into a code stream.

[0056] Image decoding: The process of restoring the code stream into a reconstructed image according to specific syntax rules and processing methods.

[0057] Video sequences have strong temporal correlations, meaning that adjacent frames have little difference. Therefore, by encoding the difference between two frames, redundant information between them can be eliminated, achieving the goal of video compression.

[0058] Generally, the video encoding process is as follows: the encoder first divides a frame of the original image into multiple parts, each of which is called an image block. The encoder then performs prediction, transform, and quantization on each image block to generate the corresponding bitstream. Prediction is performed to obtain a predicted block for the image block, allowing only the difference between the image block and its predicted block (also known as the residual or residual block) to be encoded and transmitted, thereby reducing transmission overhead. Finally, the encoder sends the corresponding bitstream to the decoder.

[0059] Accordingly, after receiving the bitstream, the decoder performs a video decoding process. Generally, the video decoding process involves performing operations such as prediction, inverse quantization, and inverse transformation on the received bitstream to obtain reconstructed image blocks (or reconstructed image blocks). This process is called image reconstruction (or image reconstruction). The decoder then assembles the reconstructed blocks of each image block in the original image to obtain a reconstructed image of the original image, which can be used for playback.

[0060] In most coding frameworks, a video sequence consists of a series of pictures, each of which is divided into at least one slice, each of which is further divided into blocks. Video encoding / decoding uses blocks as units, starting from the upper left corner of the picture and proceeding from left to right, top to bottom, and row by row. In some video coding standards, the concept of blocks is further expanded. For example, the H.264 standard includes macroblocks (MBs), which can be further divided into multiple prediction blocks (partitions) for predictive coding. The High Efficiency Video Coding (HEVC) standard uses basic concepts such as coding units (CUs), prediction units (PUs), and transform units (TUs), functionally classifying various block units and describing them using a new tree-based structure. For example, a CU can be divided into smaller CUs using a quadtree, and smaller CUs can be further divided, forming a quadtree structure. The CU is the basic unit for partitioning and encoding a coded picture. There is a similar tree structure for PU and TU. PU can correspond to the prediction block and is the basic unit of prediction coding. The CU is further divided into multiple PUs according to the partitioning mode. TU can correspond to the transform block and is the basic unit for transforming the prediction residual. However, whether CU, PU or TU, they all essentially belong to the concept of block (or coding unit). The embodiment of the present application does not specifically limit the concept of block.

[0061] In this application, an image block undergoing encoding / decoding processing is referred to as a current image block (current block), and the image in which the current image block is located is referred to as a current frame.

[0062] Prediction methods in video image coding and decoding technology include intra-frame prediction and inter-frame prediction. Among them, inter-frame prediction refers to the prediction of the coded image block / decoded image block using the correlation between the current frame and its reference frame. The current frame can have one or more reference frames, for example, the previous frame, the next frame, the second previous frame, the second next frame, etc. of the current frame can all be used as reference frames for the current frame. Specifically, inter-frame prediction generates predicted pixels (also called predicted image blocks) of the current image block of the current frame based on the pixels in the reference frame of the current frame.

[0063] Motion compensation (MC) is a process of predicting a current image block using a reference image block.

[0064] Generally, a current frame has a reference frame list. The reference frame list contains at least one reconstructed frame used as a reference frame for the current frame. The reference frame is used to provide reference pixels for inter-frame prediction of the current frame.

[0065] In the current frame, image blocks adjacent to the current image block (e.g., to the left, above, or right of the current block) may have already completed encoding / decoding, resulting in reconstructed images. These are called reconstructed image blocks. Information such as the coding mode and reconstructed pixels of the reconstructed image blocks is available.

[0066] A frame that has completed encoding / decoding processing before the current frame is encoded / decoded is called a reconstructed frame.

[0067] Motion Vector (MV): A type of motion information that can refer to the spatial displacement of an object across multiple images (e.g., a video). A motion vector is a key parameter in the inter-frame prediction process, representing the spatial displacement of a reference image block (e.g., an encoded or decoded image block) relative to the current image block. The motion vector helps the encoder / decoder determine the position and direction of motion of an object across multiple images, such as a video, thereby locating the reference image block for the current image block in the reference frame.

[0068] Generally, a motion estimation (ME) method, such as motion search, may be used to obtain a motion vector.

[0069] In the early days of inter-frame prediction, the encoder transmitted the motion vector of the current image block in the bitstream, allowing the decoder to reproduce the predicted pixels of the current image block and thus obtain the reconstructed block. To further improve coding efficiency, it was later proposed to differentially encode the motion vector using a reference motion vector, that is, only encoding the motion vector difference (MVD).

[0070] In order for the decoder and encoder to use the same reference image blocks, the encoder needs to send the motion information of each image block to the decoder in the code stream. If the encoder directly encodes the motion vector of each image block, it will consume a large amount of transmission resources. Because the motion vectors of adjacent image blocks in the spatial domain are highly correlated, the motion vector of the current image block can be predicted based on the motion vectors of adjacent encoded image blocks. The predicted motion vector is called MVP, and the difference between the motion vector of the current image block and the MVP is called MVD.

[0071] Currently, there are multiple video coding and decoding standards, such as the Advanced Video Coding (H.264 / AVC) standard, the High Efficiency Video Coding (H.265 / HEVC) standard, and the next-generation international video coding standard (Versatile Video Coding, H.266 / VVC). Various inter-frame prediction modes are used in these standards.

[0072] As an example: the H.264 / AVC standard uses an inter-frame prediction mode based on block motion estimation and motion compensation; the H.265 / HEVC standard proposes a block-based inter-frame prediction mode that can maintain similar video image quality and peak signal-to-noise ratio (PSNR) but reduce the amount of compressed data compared to the H.264 / AVC standard, such as the Advanced Motion Vector Prediction (AMVP) mode, the Merge mode and the non-translational motion model prediction mode; the inter-frame prediction modes used in the H.266 / VVC standard include a block-based affine transform motion compensation prediction mode, which can more accurately perform motion compensation prediction on complex motions, overcome the limitations of the translational motion model and maintain low computational complexity.

[0073] In the H.264 / AVC standard and the H.265 / HEVC standard, in the inter-frame prediction mode adopted, the motion vector of a block can be represented by the motion vector of a pixel of the block. For example, for the AMVP mode, the encoder constructs a candidate motion vector list through the motion information of the encoded image blocks adjacent to the current image block in the spatial or temporal domain, and determines the optimal motion vector from the candidate motion vector list as the MVP of the current image block based on the rate-distortion cost. In addition, the encoder performs a motion search in the neighborhood centered on the MVP to obtain the motion vector of the current image block (as shown in Figure 1, Figure 1 is a schematic diagram of the motion information of the current block provided by an embodiment of the present application. The motion vector of the block can be the motion vector (vx, vy) of the upper left vertex pixel of the block). Among them, the candidate motion vector list can include the following types of candidates: candidates in the candidate motion vector list of its neighboring CU, candidates constructed by the MV of the translation motion of the neighboring CU, and 0 vector. The encoder transmits the index value of the MVP in the candidate motion vector list (i.e., the above-mentioned MVP flag), the index value of the reference frame, and the MVD to the decoder.

[0074] For block-based affine transform motion compensation prediction mode, the codec uses the same motion model to derive the affine motion information of the current image block or each sub-block within the current image block, and performs motion compensation based on the affine motion information of the current image block or all sub-blocks to obtain the predicted image block. Commonly used motion models for the codec are the 4-parameter affine motion model and the 6-parameter affine motion model. The 4-parameter affine motion model saves two parameter bits compared to the 6-parameter affine motion model, resulting in higher coding efficiency.

[0075] For example, as shown in FIG2A , FIG2A is a schematic diagram of the affine motion model of the block provided in an embodiment of the present application. The 4-parameter affine motion model can be represented by the motion vectors of two pixels of the block (such as the current block or a sub-block of the current block). Here, the pixel points used to represent the motion model parameters are called control points, and the motion vectors of the control points are called CPMV. If the upper left vertex P1 (0, 0) and the upper right vertex P2 (W, 0) pixel points of the block are control points, and the motion vectors of the upper left vertex P1 and the upper right vertex P2 of the block are MV0 (vx0, vy0) and MV1 (vx1, vy1) respectively, then the motion vector of the center pixel of the block is obtained according to the following formula (1), and the motion vector of the center pixel is used as the required affine motion information. In the following formula (1), (x, y) is the coordinate of the center pixel of the block, (vx, vy) is the motion vector of the center pixel of the block, and W is the width of the current image block.

[0076] For example, as shown in FIG2B , FIG2B is a second schematic diagram of the affine motion model of the block provided in an embodiment of the present application. The 6-parameter affine motion model can be represented by the motion vectors of three pixels of the block (such as the current block or a sub-block of the current block). If the upper left vertex P1 (0, 0), the upper right vertex P2 (W, 0) and the lower left vertex P3 (0, H) pixel points of the block are control points, and the motion vectors of the upper left vertex, the upper right vertex and the lower left vertex of the block are MV0 (vx0, vy0), MV1 (vx1, vy1) and MV2 (vx2, vy2) respectively, then the motion vector of the center pixel of the block is obtained according to the following formula (2), and the motion vector of the center pixel is used as the required affine motion information. In the following formula (2), (x, y) is the coordinate of the center pixel of the block, (vx, vy) is the motion vector of the center pixel of the block, and W and H are the width and height of the block respectively.

[0077] The motion vector (CPMV) of the control point can be determined using motion estimation methods such as motion search. For example, in the AMVP mode, the encoder constructs a list of candidate motion vectors based on the motion information of spatially or temporally adjacent coded image blocks of the current image block and determines the optimal motion vector from this list as the MVP of the current image block based on the rate-distortion cost. Furthermore, the encoder performs a motion search within a neighborhood centered on the MVP to obtain the motion vector (CPMV) of the control point of the current image block.

[0078] The block-based affine transformation motion compensation prediction model can achieve better motion estimation and motion compensation for linear transformations such as translation, rotation, shearing and scaling of objects in video images. However, affine motion is a linear two-dimensional transformation. The relative positions and attributes of the coordinate points do not change during the transformation process. It is a transformation between two-dimensional coordinates and does not involve depth. Therefore, the block-based affine transformation motion compensation prediction model is not effective for motion estimation and motion compensation of some global motions (such as the rotation of the image caused by the up, down, left, and right rotation of the camera, and the scaling, perspective transformation / projection transformation of objects in the image caused by forward and backward movement) and complex irregular motions.

[0079] It can be seen that some existing block-based inter-frame prediction modes use a certain motion vector (generally the 0 vector or the motion vector of the adjacent block of the current block) as the starting point, and search on the reference image to obtain the motion vector MV of the current image block or the motion vector CPMV of the control point. However, due to reasons such as the processing power limitations of the processing device, the range that the processing device can search on the reference image each time will be smaller. As a result, the efficiency of searching for the motion vector MV is slow, the search time is long, and the encoding efficiency is low. In addition, it may also make it difficult to search for the motion vector MV, and it is difficult to find a reference image block that matches the current image block, resulting in low inter-frame reference and poor inter-frame prediction accuracy, and thus the amount of encoded data and the amount of transmitted data will be large.

[0080] To address the above-mentioned problems, the present application provides an inter-frame prediction method that obtains a mapping relationship between a current frame and a reference frame. Based on this mapping relationship, for a current image block (i.e., the current block), a mapping image block (i.e., the mapping block) with which it has a mapping relationship can be found in the reference frame. This mapping block has a high reference value for the current block. For example, the image block obtained by transforming the image content of the mapping block (e.g., an affine transformation, a projective transformation, etc.) is close to / similar to the content of the current block. Furthermore, after obtaining motion information, i.e., initial motion information, between the current block and the mapping block based on this mapping relationship, accurate target motion information of the current block (used to indicate the motion information between the image block matching the current block in the reference frame, i.e., the reference block and the current block, such as a motion vector or motion field) can be directly obtained based on the initial motion information. The predicted pixels of the current block are then determined based on the target motion information. This makes determining the motion information of the current block more efficient, less complex, and more accurate, thereby increasing the prediction accuracy, encoding / decoding efficiency, and the amount of encoded data (which can be represented by the encoding bit rate) and the amount of transmitted data using the prediction method of the present application.

[0081] In addition, for scenes where global motion or complex irregular motion occurs in the video image, the existing block-based inter-frame prediction mode has poor motion estimation and motion compensation effects. The mapping relationship between the current frame and the reference frame used in the inter-frame prediction method of the present application can include the entire image motion information between the two frames, such as global motion information and complex irregular motion information. By using this mapping relationship to obtain the motion information of the current block for determining the predicted pixels, it can achieve better motion estimation and motion compensation effects for global motion and complex irregular motion in the video image.

[0082] The inter-frame prediction method provided in this application is applicable to a video codec system. FIG3 shows the structure of a video codec system. As shown in FIG3 , video codec system 300 includes a source device 10 and a destination device 20. Source device 10 generates encoded video data and may also be referred to as a video encoding device or video encoding apparatus. Destination device 20 may decode the encoded video data generated by source device 10 and may also be referred to as a video decoding device or video decoding apparatus. Source device 10 and / or destination device 20 may include at least one processor and a memory coupled to the at least one processor. The memory may include, but is not limited to, read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), flash memory, or any other medium that can be used to store desired program code in the form of computer-accessible instructions or data structures, and this application does not specifically limit this.

[0083] Source device 10 and destination device 20 may comprise a variety of devices, including desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called “smart” phones, televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, or the like.

[0084] Destination device 20 may receive encoded video data from source device 10 via link 30. Link 30 may include one or more media and / or devices capable of moving the encoded video data from source device 10 to destination device 20. In one example, link 30 may include one or more communication media that enable source device 10 to transmit the encoded video data directly to destination device 20 in real time. In this example, source device 10 may modulate the encoded video data according to a communication standard (e.g., a wireless communication protocol) and transmit the modulated video data to destination device 20. The one or more communication media may include wireless and / or wired communication media, such as radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media may form part of a packet-based network, such as a local area network, a wide area network, or a global network (e.g., the Internet). The one or more communication media may include routers, switches, base stations, or other equipment that enables communication from source device 10 to destination device 20.

[0085] In another example, the encoded video data can be output from output interface 140 to storage device 40. Similarly, the encoded video data can be accessed from storage device 40 via input interface 240. Storage device 40 can include a variety of locally accessible data storage media, such as Blu-ray discs, high-density digital video discs (DVDs), compact disc read-only memories (CD-ROMs), flash memory, or other suitable digital storage media for storing encoded video data.

[0086] In another example, storage device 40 may correspond to a file server or another intermediate storage device that stores the encoded video data generated by source device 10. In this example, destination device 20 may obtain the stored video data from storage device 40 via streaming or downloading. The file server may be any type of server capable of storing encoded video data and transmitting the encoded video data to destination device 20. For example, the file server may include a World Wide Web (Web) server (e.g., for a website), a File Transfer Protocol (FTP) server, a Network Attached Storage (NAS) device, and a local disk drive.

[0087] Destination device 20 may access the encoded video data via any standard data connection, such as an Internet connection. Example types of data connections include a wireless channel suitable for accessing encoded video data stored on a file server, a wired connection (e.g., a cable modem), or a combination of both. Transmission of the encoded video data from the file server may be a streaming transmission, a download transmission, or a combination of both.

[0088] The inter-frame prediction method of the present invention is not limited to wireless application scenarios. For example, the inter-frame prediction method of the present invention can be applied to video codecs supporting various multimedia applications such as over-the-air television broadcasting, cable television transmission, satellite television transmission, streaming video transmission (e.g., via the Internet), video captured by inspection cameras, virtual game videos, encoding of video data stored on data storage media, decoding of video data stored on data storage media, or other applications. In some examples, the video codec system 300 can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.

[0089] It should be noted that the video encoding and decoding system 300 shown in Figure 3 is only an example of a video encoding and decoding system and does not limit the video encoding and decoding system in this application. The inter-frame prediction method provided in this application can also be applied to scenarios where there is no data communication between the encoding device and the decoding device. In other examples, the video data to be encoded or the encoded video data can be retrieved from a local memory or streamed over a network. The video encoding device can encode the video data to be encoded and store the encoded video data in a memory, and the video decoding device can also obtain the encoded video data from the memory and decode the encoded video data.

[0090] 3 , source device 10 includes a video source 101, a video encoder 102, and an output interface 103. In some examples, output interface 103 may include a modem and / or a transmitter. Video source 101 may include a video capture device (e.g., a camera), a video archive containing previously captured video data, a video input interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources of video data.

[0091] Video encoder 102 may encode video data from video source 101. In some examples, source device 10 transmits the encoded video data directly to destination device 20 via output interface 103. In other examples, the encoded video data may also be stored on storage device 40 for later access by destination device 20 for decoding and / or playback.

[0092] In the example of FIG3 , destination device 20 includes a display device 201, a video decoder 202, and an input interface 203. In some examples, input interface 203 includes a receiver and / or a modem. Input interface 203 can receive encoded video data via link 30 and / or from storage device 40. Display device 201 can be integrated with destination device 20 or can be external to destination device 20. Generally, display device 201 displays decoded video data. Display device 201 can include a variety of display devices, such as a liquid crystal display, a plasma display, an organic light emitting diode display, or other types of display devices.

[0093] Alternatively, the video encoder 102 and video decoder 202 may each be integrated with an audio encoder and decoder, and may include appropriate multiplexer-demultiplexer units or other hardware and software to handle the encoding of both audio and video in a common data stream or in separate data streams.

[0094] The video encoder 102 and the video decoder 202 may include at least one microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), discrete logic, hardware, or any combination thereof. If the bidirectional inter-frame prediction method provided in the present application is implemented in software, the instructions for the software may be stored in a suitable non-volatile computer-readable storage medium, and at least one processor may be used to execute the instructions in hardware to implement the present application. Any of the foregoing (including hardware, software, a combination of hardware and software, etc.) may be considered as at least one processor. The video encoder 102 may be included in an encoder, and the video decoder 202 may be included in a decoder. The encoder or decoder may be part of a combined encoder / decoder (encoder / decoder) in the corresponding device.

[0095] The video encoder 102 and the video decoder 202 in this application may operate according to a video compression standard (eg, HEVC), or may operate according to other industry standards, which is not specifically limited in this application.

[0096] The video encoder 102 is configured to determine predicted pixels of a to-be-encoded image block, i.e., a current block, in a to-be-encoded frame according to the inter-frame prediction method provided herein, and encode the current block of the to-be-encoded frame based on the predicted pixels of the current block to obtain a bitstream and send the bitstream to the video decoder 202. The bitstream may include some parameter data and some syntax elements.

[0097] In some examples, video decoder 202 may perform a decoding process that is generally the inverse of the encoding process described with respect to video encoder 102 .

[0098] During the decoding process, the video decoder 202 receives a bitstream from the video encoder 102. Optionally, the video data may be stored in a video data memory (not shown). The video data memory may store video data, such as an encoded bitstream, to be decoded by the components of the video decoder 202. The video data stored in the video data memory may be obtained, for example, from the storage device 40, from a local video source such as a camera, via a wired or wireless network for video data communication, or by accessing physical data storage media.

[0099] The video decoder 202 is used to: obtain a code stream, parse the code stream, determine the predicted pixels of the image block to be decoded, i.e., the current block, in the frame to be decoded according to the inter-frame prediction method provided in the present application, and reconstruct the current block of the frame to be decoded based on the predicted pixels of the current block to obtain a reconstructed image frame.

[0100] The video decoder 202 may entropy decode the bitstream to generate quantized coefficients and some syntax elements. The video decoder 202 may receive the syntax elements at the video slice level and / or the image block level.

[0101] The video encoder 102 and the video decoder 202 of the video coding and decoding system can calculate the target motion information of the current image block according to the inter-frame prediction method proposed in this application, thereby determining the predicted pixels of the current image block according to the target motion information of the current image block.

[0102] The inter-frame prediction method provided by the present application, and the video encoding method and video decoding method using the inter-frame prediction method proposed by the present application are described in detail below.

[0103] Figure 4 is a flow chart of an inter-frame prediction method provided by an embodiment of the present application. The method shown in Figure 4 can be executed by an inter-frame prediction device, a video encoding / decoding device, a video encoder / decoder, and other video processors or processing devices with video encoding / decoding functions.

[0104] As shown in FIG4 , the inter-frame prediction method provided in the embodiment of the present application may include the following steps:

[0105] Step 410: Obtain a mapping relationship between the current frame and a reference frame of the current frame.

[0106] The reference frame may be a frame adjacent to the current frame in the video image sequence, such as the previous frame, the next frame, the second previous frame, the second next frame, etc. In FIG4 , the current frame may be denoted by dst and the reference frame may be denoted by src.

[0107] The image content of the reference frame may have the same or similar parts as the current frame, that is, the image content of the reference frame is reference-based for the current frame. Specifically, the image content in the current frame and the reference frame may have a linear transformation and / or perspective transformation (projection transformation) relationship. Linear transformations may include affine transformations such as rotation, flipping (mirroring), scaling (scaling), shearing (shifting), etc. In some embodiments, the image content in the current frame and the reference frame may be two-dimensional images of the same object / scene viewed from different perspectives.

[0108] The relationship between two frames of images with reference content can be represented by a mapping relationship, which enables a pixel on the current frame to be mapped to a pixel on the reference frame, or vice versa.

[0109] The aforementioned mapping relationship can be for the entire image or for a portion of the image content. For example, two frames of images of the same scene viewed from different perspectives have a mapping relationship for the entire image, or partial objects in the two frames are two-dimensional images of the same object viewed from different perspectives, or two-dimensional images of the same object before and after translation, global motion, or complex irregular motion, and the partial images (parts) of the two frames have a mapping relationship.

[0110] In some embodiments, the current frame and the reference frame may be two frames of images captured of the same object / scene in a camera rotation shooting scenario. The camera rotation shooting scenario may be a real camera rotation shooting (e.g., camera inspection, etc.) or a virtual camera rotation shooting (e.g., a virtual camera rotation shooting in a virtual space such as a game space).

[0111] However, the inter-frame prediction method provided in this application is not limited to video images captured of the same object / scene in a camera rotation shooting scenario, but can also be applied to video images in various other scenarios without any restrictions on the scene of the video image.

[0112] As an example only, the image pair 402 shown in Figure 4 is a schematic diagram of the current frame and the reference frame provided in an embodiment of the present application. As shown in the image pair 402, the current frame dst0 and the reference frame src0 are images captured before and after the real camera is rotated. Among them, the camera rotation can be described as a rotation around the optical center of the camera. In the real world, the position of the object has not changed, but the angle of the camera shooting has changed, so there is a mapping relationship between src0 and dst0 (such as a global projection transformation relationship). The pixels in the images of src0 and dst0 can be converted and calculated with each other through the mapping relationship, for example, the pixel on src0 is mapped to a pixel on dst0 through the mapping relationship, or the pixel on dst0 is mapped to a pixel on src0 through the mapping relationship.

[0113] As another example, the image pair 404 shown in FIG4 is a second schematic diagram of a current frame and a reference frame provided in an embodiment of the present application. As shown in image pair 404, the current frame dst1 (target view) and the reference frame src1 (reference view) are two views correspondingly acquired before and after the virtual camera moves in the virtual space of the game, i.e., views from two perspectives.

[0114] In a game scene, a pixel, such as pixel P, undergoes transformations in multiple coordinate spaces to map from three-dimensional space coordinates to two-dimensional pixel coordinates. For example, a model in world space must first enter clipping space from its own model space, then from clipping space to the virtual camera's camera space, then transform from camera space to the game scene's world space. After being captured by the game's virtual camera, it enters camera space (also known as observation space), undergoes a projection transformation into clipping space, and is clipped by the view frustum before being mapped to screen space, becoming the game view seen by the player on the screen.

[0115] From the model space transformation in the above game scene, it can be seen that when the virtual camera moves, the model in the game scene itself does not change, only the virtual camera is displaced or rotated. Therefore, from the perspective of the virtual camera, since the absolute coordinates of the model in the world space remain unchanged, the pixel coordinate relationship between the corresponding matching points on the two frames of view obtained before and after the virtual camera moves can be used as a medium for mutual mapping between the two frames of view. In this way, if the connection between the two frames of image and the world coordinates is omitted, the mapping relationship between the images can be directly established, that is, the pixels in the images of src1 and dst1, such as pixel P, can be converted and calculated to each other through the mapping relationship.

[0116] In some embodiments, the mapping relationship between two frames of images can be represented by a projection transformation matrix / homography matrix (also known as a global projection matrix).

[0117] A camera without lens distortion captures images of the same planar object from different positions using homographies. Therefore, the corresponding homography matrix H can be obtained between the two frames of view acquired before and after the camera moves. Therefore, the homography matrix can more accurately represent the mapping relationship between the current frame and the reference frame.

[0118] The mapping relationship between the two frames of images can be obtained using various existing calculation methods, or can also be obtained using the simplified method provided in this application. The specific content of the simplified method is described in detail later. Alternatively, the mapping relationship between the two frames of images can be pre-calculated using various existing calculation methods or the simplified method provided in this application. The mapping relationship can be obtained by simply obtaining information containing the mapping relationship.

[0119] As an example, the mapping relationship between two frames of images (such as a homography matrix) can be calculated based on a method of matching feature points of the two frames of images. Specifically, the method of matching feature points may include a scale-invariant feature transform (SIFT) algorithm. The SIFT algorithm has scale and rotation invariance, and the key points found are not affected by changes in lighting, viewing angle, etc. Specifically, the method of matching feature points may include scale space extreme value detection and key point positioning.

[0120] Scale-space extreme point detection, which searches for image positions at all scales, builds a Gaussian pyramid, and then generates a Gaussian difference pyramid (DOG). The image is Gaussian blurred to varying degrees to detect potential points of interest that are invariant to scale and rotation. This is also called local extreme point detection.

[0121] First, the scale space L(x,y,σ) of an image is defined as the convolution of a Gaussian function G(x,y,σ) with a varying scale and the original image I(x,y), i.e., L(x,y,σ)=G(x,y,σ)*I(x,y) (3)

[0122] Where m and n represent the dimensions of the Gaussian template, (x, y) represents the position of the image pixel, and σ represents the scale space factor.

[0123] The scale space is represented by a Gaussian pyramid in the implementation. In order to effectively extract stable key points, Gaussian difference kernels of different scales are used with convolution to generate Gaussian difference pyramid (DOG), where k is the coefficient of σ: D(x,y,σ)=[G(x,y,kσ)-G(x,y,σ)]*I(x,y) D(x,y,σ)=L(x,y,kσ)-L(x,y,σ) (5)

[0124] Using the DOG pyramid to generate a Gaussian difference image reveals changes in image pixel values. Since feature points consist of local extreme points of the DOG function, it is necessary to find the extreme points of the DOG function. Each pixel is compared with its neighbors to see whether it is larger or smaller than its neighbors in the image and scale domains.

[0125] After finding the extreme points, in order to enhance matching stability and improve noise resistance, it is necessary to accurately locate the key points, including using the known discrete space point difference to obtain the continuous space extreme points, that is, the sub-pixel difference; and removing the edge response. Next, the main direction of the key points should be assigned based on the local gradient direction of the image so that their descriptors are rotation invariant. To achieve this goal, the gradient and distribution characteristics of the pixels within the 3σ neighborhood window of the Gaussian pyramid image where the key points are located are first collected and statistically analyzed. The modulus and direction of the gradient are as follows, where m(x, y) is the gradient amplitude and θ(x, y) is the gradient direction:

[0126] After completing the above steps, each keypoint now has information about its position, scale, and orientation. Next, we create a unique descriptor for each keypoint, making it immune to external variations and improving matching accuracy. To ensure rotational invariance, we rotate the original image's x-axis, centered around the keypoint, to align with the principal direction before generating the matching feature points.

[0127] Next, we implement feature point matching. We create feature point descriptor sets for the reference image and the target image respectively, and achieve target recognition by matching the feature point descriptors in the set. The similarity measurement of the feature point descriptors with 128 dimensions uses the Euclidean distance. Let the feature point descriptor set in the reference image be R i , the set of feature point descriptors in the target image is S i , then R i and S i The similarity measure between any two feature point descriptors can be expressed as:

[0128] Where j represents the dimension of the feature point descriptor.

[0129] Finally, a Kd-tree structure is used to match feature points. Using the target image's feature points as a benchmark, the nearest feature points in the original image and the next closest feature points are searched for to achieve the final feature point match. To improve matching accuracy and reduce the potential for errors caused by mismatched points, the resulting matching point pairs are subjected to denoising using the Random Sample Consensus (RANSAC) algorithm.

[0130] The RANSAC algorithm treats each matched feature point pair as a set, randomly selects multiple pairs of feature points from them, and fits them into a model. It then calculates the distance between the remaining point pairs in the set and the model, and counts the number of point pairs that do not exceed a fixed threshold, i.e., the inliers. After multiple iterations, the model with the most inliers is selected as the final fitting result. The model finally fitted by RANSAC is applied to the matrix representing the mapping relationship between the two frames, such as the homography matrix H, which can be expressed as:

[0131] According to the form characteristics of the homography matrix, it includes 8 unknown parameters h1, h2, h3, h4, h5, h6, h7, and h8. To recover the 8 unknown parameters, at least 4 pairs of matching points (x1, y1) and (x1′, y1′), (x2, y2) and (x′2, y ′2 ′), (x3,y3) and (x′3,y ′3 ′3), (x4,y4) and (x′4 ′4 ,y ′4 ′4), the model finally fitted by RANSAC can be:

[0132] Among them, h 11 、h 12 、h 13 、h 21 、h 22 、h 23 、h 31 、h 32 、h 33 Are unknown parameters. By simplifying these 9 unknown parameters in the formula, we can get 8 unknown parameters h1, h2, h3, h4, h5, h6, h7, and h8. The 8 unknown parameters constitute the matrix H representing the mapping relationship between the two frames.

[0133] As another example, the mapping relationship between two frames can be calculated by the following method: based on the motion estimation of the blocks of any frame in the two frames, the matching blocks can be obtained in the other frame, and for the matching image block pairs between the two frames, the upper left corner point or the center point is taken as the "matching point", and then the RANSAC algorithm is used to obtain the matrix (which can be the homography matrix H) used to characterize the mapping relationship between the two frames.

[0134] For situations where the current frame and the reference frame can be two frames of images of the same object / scene captured in a camera rotation scenario, when the camera rotates, the distance from the object to the camera remains unchanged, as the camera itself does not move, only its orientation changes. This means that the focal point of the entire image remains unchanged. Therefore, before and after camera rotation, the presented views are homogeneous because they share the same optical center. Corresponding pixels have a non-singular linear relationship, and the far plane can be assumed to be at infinity. Consequently, depth information can be ignored in the pixel mapping transformation between the two frames.

[0135] Therefore, in addition to the existing calculation method, the embodiment of the present application also provides a simplified method for calculating the mapping relationship between the two frames. The method includes: obtaining the rotation angle between the two frames taken by the camera, and obtaining the intrinsic parameters of the camera, and calculating the matrix used to characterize the mapping relationship between the two frames based on the rotation angle of the camera and the intrinsic parameters of the camera (which can be the homography matrix H).

[0136] Specifically, the rotation matrix R can be obtained according to the rotation angle, and the homography matrix H can be obtained according to the rotation matrix R and the camera intrinsic parameter matrix K by the following formula: H = KRK -1 (11)

[0137] As an example, the rotation matrix R is a 3x3 matrix determined by the rotation angle, and the camera intrinsic parameter matrix K is:

[0138] Where f is the focal length of the camera, c is the principal point offset of the camera, and f x is the horizontal focal length of the camera, f y is the vertical focal length of the camera, c x The horizontal offset of the main point is also called horizontal offset, c y The vertical offset of the main point is also called vertical offset.

[0139] Step 420: Determine initial motion information according to the current image block in the current frame and the mapping relationship.

[0140] It can be understood that, based on the mapping relationship between the current frame and the reference frame, there exists a mapping block in the reference frame that has a mapping relationship with the current block in the current frame. The content of this mapping block is similar to that of the current block, and the mapping block can be considered an image block close to the reference block required for inter-frame prediction. Furthermore, based on the mapping relationship and the current block, motion information between the current block and the mapping block in the reference frame that has a mapping relationship with the current block can be determined, which can be used as the initial motion information of the current block. This initial motion information will also be slightly different from the motion information between the reference block and the current block required for inter-frame prediction.

[0141] Specifically, the initial motion information can be determined by the following method: mapping the original pixel positions of the pixels in the current block according to the mapping relationship to obtain the mapped pixel positions, that is, the pixel positions mapped in the reference frame; and determining the initial motion information based on the position change between the mapped pixel positions and the original pixel positions.

[0142] As an example, FIG5 is a schematic diagram of pixel mapping provided by an embodiment of the present application. For any pixel point in the current frame, its mapped pixel position can be obtained according to the method in FIG5. As shown in FIG5, for any pixel point P in the current frame dst2 dst The normalized coordinates of the original pixel position are P dst =(u1,v1,1) T , according to the following formula, the mapping relationship such as homography matrix H and P dst Substituting the coordinates into the calculation, we can get P dst Pixel point P mapped in reference frame src2 src The normalized coordinate P src =(u0,v0,1) T That is, mapping pixel positions:

[0143] In the embodiments of the present application, for various inter-frame prediction modes, there may be various embodiments for determining initial motion information. The following first describes various embodiments of the inter-frame prediction mode.

[0144] In the first inter-frame prediction mode, the motion information of the block can be represented based on the motion vector of a pixel of the block (for example, the motion vector of the block can be the motion vector (vx, vy) of the upper left vertex pixel / center pixel of the block). The inter-frame prediction mode using the motion information representation method in this embodiment can be, for example, an inter-frame prediction mode such as the AMVP mode used in the H.264 / AVC standard and the H.265 / HEVC standard.

[0145] In the second inter-frame prediction mode, the motion information of the block can be represented based on the motion vectors of multiple pixel points (also called control points) of the block (for example, the motion information of the block can be obtained based on the motion vectors of two control points of a four-parameter affine motion model or three control points of a six-parameter affine motion model). The inter-frame prediction mode using the motion information representation method in this embodiment can be, for example, the block-based affine transform motion compensation prediction mode adopted in the H.266 / VVC standard.

[0146] In the second inter-frame prediction mode, the current block may refer to an image block to be encoded / decoded, or, in order to make motion estimation and motion compensation more refined and the predicted pixels more accurate, the current block may refer to a sub-block of an image block to be encoded / decoded. Figure 6 is a schematic diagram of the image block sub-block division provided by an embodiment of the present application. For example, as shown in Figure 6, the image block S to be encoded / decoded can be divided to obtain multiple sub-blocks 1-16. Therefore, in the second inter-frame prediction mode, the initial motion information can be determined for each sub-block, and the target motion information can be subsequently determined for each sub-block based on its initial motion information, and then the predicted pixels can be determined for each sub-block based on its target motion information.

[0147] The present application also proposes a third inter-frame prediction mode. In the third inter-frame prediction mode, the motion vector of each pixel in the current block can be obtained, and the motion information of the current block (such as motion vector / motion field) is represented by the motion vectors of all pixels in the current block. According to the motion information of the current block, the mapping block in the reference frame that has a mapping relationship with the current block can be warped to obtain a warped reference frame, and then the predicted pixels of the current block are obtained based on the warped reference frame.

[0148] The present application also proposes a fourth inter-frame prediction mode. In the fourth inter-frame prediction mode, the current block can be any image block in the current frame. The motion vector of each pixel of each image block in the current frame, that is, the motion vector of all pixels of the current frame, can be obtained to represent the motion information of the current frame (such as motion vector / motion field), and the mapped image of the entire current frame in the reference frame is distorted (warped) according to the motion information of the current frame to obtain a distorted reference frame, and then motion search and motion compensation are performed based on the distorted reference frame to obtain the predicted pixels of the current block to be encoded / decoded.

[0149] The present application also proposes a fifth inter-frame prediction mode. In the fifth inter-frame prediction mode, a distorted reference frame can be obtained in a manner similar to the third inter-frame prediction mode, and a first predicted pixel of the current block can be obtained based on the distorted reference frame. In addition, target motion information of the current block is obtained, and a reference block matching the current block is determined in the reference frame based on the target motion information. The second predicted pixel of the current block is determined based on the reference block, and finally, the predicted pixel of the current block is obtained based on the first predicted pixel and the second predicted pixel.

[0150] The following describes various embodiments of determining initial motion information that can be used for various inter-frame prediction modes.

[0151] In the first method of determining the initial motion information, for the first inter-frame prediction mode and the fifth inter-frame prediction mode, the motion information of the block can be represented based on the motion vector of a pixel point of the block. The initial motion information determination method includes: according to the aforementioned method, for a pixel in the current block (such as the upper left vertex pixel, the center pixel), the position change between the mapped pixel position and the original pixel position of the pixel is obtained to obtain the initial motion information, that is, the motion vector between the mapped pixel position and the original pixel position, which can be called the initial motion vector WMV.

[0152] In the second method of determining the initial motion information, for the second inter-frame prediction mode and the fifth inter-frame prediction mode, the motion information of the block can be represented based on the motion vectors of multiple pixel points (also called control points) of the block. The initial motion information determination method includes: according to the aforementioned method, for each control point, obtaining the position change between the mapped pixel position and the original pixel position of the control point to obtain the initial motion vector of the control point. The initial motion vectors of the multiple control points can be used as the initial motion information of the block.

[0153] In a third method for determining initial motion information, for the third inter-frame prediction mode, the motion information of the current block is represented based on the motion vectors of all pixels in the current block. The method for determining the initial motion information includes: obtaining, for each pixel in the current block, a position change between a mapped pixel position and an original pixel position of the pixel according to the aforementioned method, obtaining the motion vectors of all pixels in the current block to represent the initial motion information of the current block. The mapped pixels corresponding to all pixels in the current block (i.e., pixels at the mapped pixel positions) may constitute a mapped block.

[0154] In the fourth method of determining the initial motion information, for the fourth inter-frame prediction mode, the motion information of the current frame is represented based on the motion vector of all pixels of the current frame (that is, the motion vector of all pixels in all image blocks of the current frame). The method for determining the initial motion information includes: according to the aforementioned method, for each pixel of each image block in the current frame, obtaining the position change between the mapped pixel position and the original pixel position of the pixel, and obtaining the motion vector of all pixels of the current frame to represent the motion information of the current frame, that is, the initial motion information.

[0155] Step 430: Determine target motion information of the current image block according to the initial motion information.

[0156] In the embodiment of the present application, for various inter-frame prediction modes, there may be various embodiments for determining target motion information based on initial motion information. The target motion information is used to determine predicted pixels of the current block.

[0157] In the first target motion information determination method, for the first inter-frame prediction mode, the second inter-frame prediction mode, the third inter-frame prediction mode, and the fifth inter-frame prediction mode, it can be considered that the initial motion information obtained in some scenarios is already very close to the required target motion information, or is equivalent to the target motion information. According to needs, the initial motion information can be directly used as the target motion information of the current block.

[0158] For example, the initial motion vector WMV in the first initial motion information determination method can be directly used as the motion vector MV of the current block, that is, the target motion information.

[0159] For another example, the initial motion vectors of the multiple control points in the second method for determining initial motion information can be directly used as the motion vectors CPMVs of the multiple control points required for inter-frame prediction, i.e., target motion information. The predicted pixels of the current block can then be derived based on the motion vectors CPMVs of the multiple control points. For a more detailed description of how to derive the predicted pixels of the current block based on the CPMVs of the multiple control points, please refer to the relevant introduction above.

[0160] For another example, in the third method of determining initial motion information, the initial motion information of the current block represented by the motion vectors of all pixels of the current block can be directly used as the motion information of the current block (such as motion vector / motion field), that is, the target motion information.

[0161] In the second target motion information determination method, for the first, second, and fifth inter-frame prediction modes, the initial motion information can be used as the starting point for a motion search, i.e., an initial value, and a motion search can be performed in the reference frame to obtain the target motion information for the current block. The motion search method can utilize any existing inter-frame prediction motion search method, without limitation.

[0162] For example, Figure 7 is a schematic diagram of the motion vector search provided in an embodiment of the present application. As shown in Figure 7, the initial motion vector WMV in the first initial motion information determination method can be used as the starting point of the motion search (WMV points to the position of the block in the reference frame), and then in the reference frame, a motion search is performed within the search range R (i.e., the dotted box area) centered on WMV to obtain the motion vector MV of the current block, i.e., the target motion information.

[0163] For another example, for the initial motion vectors of multiple control points in the second method of determining initial motion information, the initial motion vector of each control point can be used as the starting point of the motion search, and then in the reference frame, a motion search is performed within the search range centered on the initial motion vector of the control point to obtain the motion vector CPMV of each control point, that is, the target motion information.

[0164] In the third method of determining target motion information, for the fourth inter-frame prediction mode, the method of obtaining target motion information includes: performing image warping processing (warp processing) on ​​the mapping block in the reference frame according to the initial motion information to obtain a warped reference frame, and then performing motion search in the warped reference frame according to the current block to determine the motion information of the current block, that is, the target motion information.

[0165] After the pixels in the mapping block are warped based on the initial motion information, a corresponding post-warping pixel can be obtained on the reference frame. For example, in Figure 5 , after the pixel Psrc of the mapping block on the reference frame src is warped, a post-warping pixel Psrc' can be obtained on the reference frame src. This post-warping pixel can be similar to / identical to the pixel in the current block that is mapped to the mapping block pixel Psrc.

[0166] The distortion process can be implemented by various methods such as interpolation, which is not limited in this application. The motion search can be implemented by various existing motion search methods for inter-frame prediction, which is not limited in this application.

[0167] After the method of the present application obtains the motion information between the current block and the mapped block, that is, the initial motion information, according to the mapping relationship, it can directly obtain the accurate target motion information (such as motion vector, motion vector / motion field) of the current block based on the initial motion information, and then determine the predicted pixels of the current block based on the target motion information, so that the method of determining the target motion information is more efficient, less complex, and the obtained motion information is more accurate, thereby making the prediction accuracy of the prediction method of the present application higher, the encoding / decoding efficiency higher, and the amount of encoded data (which can be represented by the encoding bit rate) and the amount of transmitted data smaller.

[0168] Step 440: Determine predicted pixels of the current image block according to the target motion information.

[0169] In the embodiment of the present application, for multiple inter-frame prediction modes, there may be multiple embodiments for determining the predicted pixels of the current block according to the target motion information.

[0170] In the first method of determining predicted pixels, for the first inter-frame prediction mode, a reference block matching the current block can be determined in the reference frame based on the target motion information, i.e., the motion vector of the current block. The positional relationship between the reference block and the current block is the positional relationship indicated by the target motion information, and then the predicted pixels of the current block are determined based on the reference block.

[0171] In the second method of determining predicted pixels, for the second inter-frame prediction mode, the motion vector of the block can be determined based on the target motion information, that is, the motion vectors CPMV of multiple control points of the block, and then the reference block matching the current block is determined in the reference frame based on the motion vector of the block. The positional relationship between the reference block and the current block is the positional relationship indicated by the motion vector of the block, and then the predicted pixels of the current block are determined based on the reference block.

[0172] The predicted pixels of the current block determined by the second initial motion information determination method, the first or second target motion information determination method, and the second predicted pixel determination method for the second inter-frame prediction mode have high accuracy, which can reach the same level of high accuracy as the predicted pixels obtained by refining the sub-block-based affine motion compensation prediction by prediction refinement (PROF) with optical flow, and compared with refining the sub-block-based affine motion compensation prediction by prediction refinement (PROF) with optical flow, the prediction efficiency is higher, saving encoding time.

[0173] In a third method for determining predicted pixels, for the third inter-frame prediction mode, a mapping block in a reference frame that is mapped to the current block may be warped based on motion information of the current block represented by motion vectors of all pixels of the current block (e.g., motion vector / motion field), i.e., target motion information. The predicted pixels of the current block may then be obtained based on the warped mapping block. For example, the pixels of the warped mapping block may be directly used as the predicted pixels of the current block.

[0174] In the fourth method of determining predicted pixels, for the fourth inter-frame prediction mode, a reference block (called a second reference image block / second reference block) that matches the current block can be determined in the distorted reference frame based on the target motion information of the current block. The positional relationship between the current block and the second reference block is the positional relationship indicated by the target motion information, and then the predicted pixels of the current block are determined based on the second reference block.

[0175] The warped image block obtained by warping the image content of the mapping block according to the initial motion information is close to / similar to the content of the current block. Therefore, after warping one or more image blocks of the current frame mapped to one or more mapping blocks in the reference frame, the warped reference frame obtained can be aligned with the current frame, that is, the difference between the warped reference frame and the current frame (such as object displacement) can be small. Therefore, by performing a motion search on the warped reference frame, that is, searching for the target motion information of the current image block, it can be more efficient and the accuracy of the searched target motion information can be improved. Based on the method of performing a motion search on the warped reference frame to obtain the target motion information of the current block, a second reference image block that matches the current block is further found in the warped reference frame based on the target motion information, so that accurate predicted pixels of the current block can be obtained based on the second reference image block.

[0176] In the fifth method for determining predicted pixels, for the fifth inter-frame prediction mode, a warped reference frame can be obtained in a manner similar to the third method for determining predicted pixels, and predicted pixels of the current block can be obtained as first predicted pixels based on the warped reference frame. Furthermore, a reference block (referred to as a first reference image block / first reference block) that matches the current block is determined in the reference frame based on target motion information. The positional relationship between the current block and the first reference block is the positional relationship indicated by the target motion information. Predicted pixels of the current block are then determined as second predicted pixels based on the reference block. Finally, predicted pixels of the current block are obtained based on the first and second predicted pixels. For example, a weighted sum of the first and second predicted pixels is performed to obtain a summed predicted pixel, which is then used as the predicted pixel of the current block. The weights of the first and second predicted pixels can be set based on actual needs and experience.

[0177] The image block obtained by warping the image content of the mapped block based on the initial motion information is close to / similar to the content of the current block. Therefore, the warped mapped image block serves as a reference for the current image block. By determining a first reference image block based on the target motion information and deriving predicted pixels for the current block based on the first reference image block, and then fusing the predicted pixels for the current block derived based on the warped mapped image block, more accurate predicted pixels can be obtained, improving prediction performance and, in turn, enhancing video encoding and decoding performance.

[0178] In the aforementioned embodiment, the process of determining the predicted pixels of the current block based on the reference block can be implemented using existing or future methods in video image encoding / decoding, such as using the pixels / reconstructed pixels of the reference block as the predicted pixels of the current block, etc. This application does not impose any restrictions on this.

[0179] The method provided in the embodiment of the present application obtains the motion information between the current block and the mapped block, i.e., the initial motion information, based on the mapping relationship between the current frame and the reference frame. Then, based on the initial motion information, accurate target motion information of the current block (such as a motion vector, a motion field, which can be used to indicate the image block in the reference frame that matches the current block, i.e., the motion information between the reference block and the current block) is obtained, and then the predicted pixels of the current block are determined based on the target motion information. Compared with the method of using a zero motion vector or a motion vector inherited from an image block surrounding the current block as a starting point and searching for motion vectors in the reference frame with a smaller search range each time, the method of determining target motion information based on the initial motion information in the method of the present application is more efficient, less complex, and the obtained motion information is more accurate, thereby making the prediction accuracy of the prediction method of the present application higher, the encoding / decoding efficiency higher, and the amount of encoded data (which can be represented by the encoding bit rate) and the amount of transmitted data smaller.

[0180] The video encoding and decoding method using the inter-frame prediction method proposed in this application may include the following process. This is described below with reference to the video encoder 102 and video decoder 202 shown in the video encoding and decoding system 300 in FIG3 . The method steps performed by the video encoder 102 may constitute a video encoding method, and the method steps performed by the video decoder 202 may constitute a video decoding method.

[0181] The video encoder 102 obtains an image to be encoded, ie, a current frame and a reference frame of the current frame.

[0182] The video encoder 102 obtains a mapping relationship between the current frame and the reference frame. The mapping relationship can be obtained by the video encoder 102 using various existing calculation methods, or can be obtained using the simplified method provided in the aforementioned embodiments of this application. Alternatively, the mapping relationship can be pre-calculated and stored in a memory using various existing calculation methods or the simplified method provided in this application. The video encoder 102 only needs to obtain information containing the mapping relationship from the memory to obtain the mapping relationship.

[0183] The video encoder 102 determines initial motion information based on the current image block in the current frame and the mapping relationship, where the initial motion information indicates motion information between the current image block and a mapped image block in a reference frame that has a mapping relationship with the current image block; determines target motion information for the current image block based on the initial motion information; and determines predicted pixels for the current image block based on the target motion information. For a detailed description of this method step, see the description of steps 420-440 in FIG. 4 .

[0184] The video encoder 102 encodes the current block of the frame to be encoded according to the predicted pixels of the current block to obtain a code stream.

[0185] In one embodiment, the code stream may include identification information. The video encoder 102 may encode the acquired mapping relationship between the current frame and the reference frame and put it into the code stream as identification information of the current frame. Alternatively, the video encoder 102 may encode the parameters used to calculate the mapping relationship between the current frame and the reference frame and put it into the code stream as identification information of the current frame. Wherein, if the simplified method provided in the present application is used to calculate the mapping relationship, the parameters used to calculate the mapping relationship include: camera motion parameters for shooting between the current frame and the reference frame (including camera rotation angles such as horizontal and vertical rotation angles, camera motion speeds such as horizontal speeds and vertical speeds, camera motion angle ranges, camera focal length ranges, etc.), camera intrinsic parameters (including camera focal lengths such as horizontal focal lengths and vertical focal lengths, camera principal point offsets such as horizontal offsets and vertical offsets, etc.).

[0186] By including the identification information of the current frame in the video bitstream, which includes the mapping relationship between the current frame and the reference frame, or the parameters used to calculate the mapping relationship, the decoder can use the obtained identification information to implement the inter-frame prediction method of this application and then perform image reconstruction. As a result, the video bitstream does not need to transmit the motion information of the current frame, or the transmitted motion information parameters are reduced, thereby reducing the amount of transmitted data.

[0187] In another embodiment, the video encoder 102 may generate indication information (such as a syntax element) and include it in the bitstream. The indication information may include information indicating the inter-frame prediction mode used for encoding, information indicating a method for obtaining initial motion information based on a mapping relationship between the current frame and the reference frame, information indicating a method for obtaining target motion information based on the initial motion information, and information indicating a method for determining predicted pixels of the current block based on the target motion information.

[0188] The video encoder 102 sends a code stream to the video decoder 202 .

[0189] The video decoder 202 receives a code stream, and obtains frames to be decoded, ie, a current frame and a reference frame, according to the code stream, and also obtains a mapping relationship between the current frame and the reference frame.

[0190] In one embodiment, if the bitstream includes identification information containing a mapping relationship between the current frame and the reference frame, the video decoder 202 may directly obtain the mapping relationship between the current frame and the reference frame according to the bitstream.

[0191] In another embodiment, if the bitstream includes identification information containing parameters for calculating the mapping relationship between the current frame and the reference frame, the video decoder 202 can obtain the parameters of the mapping relationship between the current frame and the reference frame and calculate the mapping relationship based on the parameters, for example, using the simplified method provided in this application.

[0192] The video decoder 202 determines initial motion information based on the current image block in the current frame and the mapping relationship, where the initial motion information indicates motion information between the current image block and a mapped image block in a reference frame that has a mapping relationship with the current image block; determines target motion information for the current image block based on the initial motion information; and determines predicted pixels for the current image block based on the target motion information. For a detailed description of this method step, see the description of steps 420-440 in FIG. 4 .

[0193] The video decoder 202 reconstructs the current block of the frame to be decoded according to the predicted pixels of the current block to obtain the reconstructed pixels of the current block, and further obtains the reconstructed image of the frame to be decoded.

[0194] The following also tests the video encoding effect of the inter-frame prediction method provided by the embodiment of the present application under various test environments. Among them, different test environments can refer to different camera shooting environments, such as stable camera shooting, camera shaking during shooting, camera parameters set to the first type during shooting, camera parameters set to the second type during shooting, etc. As shown in the following table, videos shot with various camera motion types such as left and right translation, forward movement, backward movement, and rotation were encoded under the first test environment, the second test environment, and the third test environment, respectively. The video encoding method provided by the present application and the existing encoding method were used during encoding, and the changes in the bit rate (Rate) and distortion (Distortion) (BD-rate) of the two encoding methods were calculated. According to the calculation principle of BD-rate, if the BD-rate of the video encoding method provided by the present application is negative compared to the existing encoding method, the encoding performance is improved, and the smaller the BD-rate value, the better the encoding performance improvement effect. The improved encoding performance can include reduced encoding bit rate, reduced distortion, etc.

[0195] Therefore, according to the test data recorded in Table 1 below, the encoding performance of videos shot under the aforementioned various test environments and camera motion types using the video encoding method provided by this application has been improved to varying degrees compared to encoding using existing encoding methods.

[0196] Table 1

[0197] The present application also provides an inter-frame prediction device 800 , as shown in FIG8 , including an acquisition module 810 and a prediction module 820 .

[0198] The acquisition module 810 is used to acquire a mapping relationship between a current frame and a reference frame of the current frame.

[0199] The prediction module 820 is used to determine initial motion information based on the current image block in the current frame and the mapping relationship. The initial motion information is used to indicate motion information between the current image block and a mapping image block that has a mapping relationship with the current image block in the reference frame.

[0200] The prediction module 820 is further configured to determine target motion information of the current image block according to the initial motion information; the prediction module is further configured to determine predicted pixels of the current image block according to the target motion information.

[0201] In one possible implementation, the prediction module 820 is further used to map the original pixel positions of the pixels in the current block according to the mapping relationship to obtain the mapped pixel positions, and further determine the motion information between the current block and the mapped block, i.e., the initial motion information, based on the position change between the mapped pixel positions and the original pixel positions.

[0202] In another possible implementation, the prediction module 820 is further configured to perform a motion search in a reference frame based on the current image block and using the initial motion information as an initial value for the motion search, to determine target motion information of the current image block.

[0203] In another possible implementation, the prediction module 820 is further used to perform image warping on the mapped image block in the reference frame according to the initial motion information to obtain a warped mapped image block; determine a first predicted pixel of the current image block according to the warped mapped image block; determine a first reference image block that matches the current image block according to the target motion information, wherein the first reference image block is an image block in the reference frame, and the positional relationship between the current image block and the first reference image block is the positional relationship indicated by the target motion information; determine a second predicted pixel of the current image block according to the first reference image block; and determine a predicted pixel of the current image block according to the first predicted pixel and the second predicted pixel.

[0204] In another possible implementation, the prediction module 820 is further configured to perform image warping on the mapped image block in the reference frame according to the initial motion information to obtain a warped reference frame; and to perform motion search in the warped reference frame according to the current image block to determine the target motion information.

[0205] In another possible implementation, the prediction module 820 is further used to determine a second reference image block that matches the current image block based on the target motion information, where the second reference image block is an image block in the distorted reference frame, and the positional relationship between the current image block and the second reference image block is the positional relationship indicated by the target motion information; based on the second reference image block, the predicted pixels of the current image block are determined.

[0206] In another possible implementation, the mapping relationship includes a homography matrix between the current frame and the reference frame.

[0207] In another possible implementation, the acquisition module 810 is further used to obtain camera motion parameters and camera intrinsic parameters between the current frame captured by the camera and the reference frame captured by the camera, and then determine the mapping relationship between the current frame and the reference frame based on the camera motion parameters and camera intrinsic parameters.

[0208] In another possible implementation, in the method, the acquisition module 810 is also used to obtain identification information of the current frame; the mapping relationship is obtained based on the identification information of the current frame, the identification information of the current frame includes the mapping relationship or the camera motion parameters and camera intrinsic parameters between the current frame captured by the camera and the reference frame captured by the camera, and the camera motion parameters and the camera memory are used to determine the mapping relationship.

[0209] The present application also provides a video encoding device 900, as shown in FIG9 , comprising the acquisition module 810 and the prediction module 820 as described in FIG8 , and also comprising an encoding module 910. The encoding module 910 is configured to obtain a code stream of the current image block based on the predicted pixels of the current image block obtained by the prediction module 820.

[0210] The present application also provides a video decoding device 1000, as shown in Figure 10, including the acquisition module 810 and the prediction module 820 as described in Figure 8, and also including a reconstruction module 1010, which is used to determine the reconstructed pixels of the current image block based on the predicted pixels of the current image block obtained by the prediction module 820.

[0211] The acquisition module 810 and prediction module 820, as well as the encoding module 910 and reconstruction module 1010, can all be implemented via software or hardware. For example, the implementation of acquisition module 810 will be described below using acquisition module 810 as an example. Similarly, the implementation of prediction module 820, encoding module 910, and reconstruction module 1010 can refer to the implementation of acquisition module 810.

[0212] As an example of a software functional unit, the module acquisition module 810 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, module A may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0213] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0214] As an example of a hardware functional unit, the acquisition module 810 may include at least one computing device, such as a server. Alternatively, the module A may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0215] The multiple computing devices included in the acquisition module 810 can be distributed in the same region or in different regions. The multiple computing devices included in the A module can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the A module can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0216] It should be noted that, in other embodiments, the acquisition module 810 and the prediction module 820 can be used to execute the steps of the inter-frame prediction method as described in FIG. 4 of the embodiment of the present application, and the steps that the acquisition module 810 and the prediction module 820 are responsible for implementing can be specified as needed. By having the acquisition module 810 and the prediction module 820 respectively implement different steps of the inter-frame prediction method provided in the embodiment of the present application, the full functionality of the inter-frame prediction device is achieved. Furthermore, in other embodiments, the encoding module 910 is used to execute the steps of the video encoding method provided in the embodiment of the present application, and the reconstruction module 1010 is used to execute the steps of the video decoding method provided in the embodiment of the present application.

[0217] The present application also provides a video processor, including a non-volatile storage medium and a central processing unit, the non-volatile storage medium stores an executable program, and the central processing unit is connected to the non-volatile storage medium. When the central processing unit executes the executable program, the video processor executes the inter-frame prediction method, or the video decoding method, or the video encoding method as described in Figure 4 in the embodiment of the present application, or implements the functions of the inter-frame prediction device described in Figure 8, the video encoding device described in Figure 9, and the video decoding device described in Figure 10.

[0218] This application also provides a computing device 1100. As shown in FIG11 , computing device 1100 includes a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. Processor 1104, memory 1106, and communication interface 1108 communicate with each other via bus 1102. Computing device 1100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1100.

[0219] Bus 1102 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG11 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 1102 may include a path for transmitting information between various components of computing device 1100 (e.g., memory 1106, processor 1104, and communication interface 1108).

[0220] The processor 1104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0221] The memory 1106 may include volatile memory, such as random access memory (RAM). The processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0222] The memory 1106 stores executable program code, and the processor 1104 executes the executable program code to implement the functions of the aforementioned acquisition module 810 and prediction module 820, respectively. It can also implement the functions of the encoding module 910 and the reconstruction module 1010, thereby implementing the inter-frame prediction, video decoding method, or video encoding method as described in FIG. 4 of the embodiment of the present application. That is, the memory 1106 stores instructions for executing the inter-frame prediction method, video decoding method, or video encoding method as described in FIG. 4 of the embodiment of the present application.

[0223] Alternatively, the memory 1106 stores executable code, and the processor 1104 executes the executable code to implement the functions of the aforementioned inter-frame prediction device, or video encoding device, or video decoding device, respectively, thereby implementing the inter-frame prediction method, or video decoding method, or video encoding method. That is, the memory 1106 stores instructions for executing the inter-frame prediction method, or video decoding method, or video encoding method as described in FIG. 4 in the embodiments of the present application.

[0224] The communication interface 1108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1100 and other devices or a communication network.

[0225] The present application also provides a computer program product comprising instructions, which may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device is caused to perform the inter-frame prediction method, video decoding method, or video encoding method as described in FIG. 4 of the embodiment of the present application, or to implement the functions of the inter-frame prediction device described in FIG. 8 , the video encoding device described in FIG. 9 , and the video decoding device described in FIG. 10 .

[0226] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device, or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the inter-frame prediction method, video decoding method, or video encoding method described in FIG. 4 of the embodiment of the present application.

[0227] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.

[0228] The terms "first", "second", "third" and "fourth" in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects rather than to limit a specific order.

[0229] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

Claims

1. An inter-frame prediction method, characterized in that: The method comprises: Acquire a mapping relationship between a current frame and a reference frame of the current frame; determining initial motion information according to the current image block in the current frame and the mapping relationship, the initial motion information being used to indicate motion information between the current image block and a mapped image block having a mapping relationship with the current image block in the reference frame; determining target motion information of the current image block according to the initial motion information; Determine predicted pixels of the current image block according to the target motion information.

2. The method according to claim 1, characterized in that The determining of the initial motion information according to the current image block in the current frame and the mapping relationship includes: Mapping original pixel positions of pixels in the current image block according to the mapping relationship to obtain mapped pixel positions; The initial motion information is determined based on a position change between the mapped pixel position and the original pixel position.

3. The method according to claim 1 or 2, characterized in that Determining the target motion information of the current image block according to the initial motion information includes: According to the current image block, a motion search is performed in the reference frame using the initial motion information as an initial value for motion search to determine the target motion information.

4. The method according to any one of claims 1 to 3, characterized in that The determining the predicted pixels of the current image block according to the target motion information includes: performing image warping processing on the mapping image block in the reference frame according to the initial motion information to obtain a warped mapping image block; determining a first predicted pixel of the current image block according to the warped mapped image block; determining, in the reference frame according to the target motion information, a first reference image block that matches the current image block, wherein a positional relationship between the current image block and the first reference image block is a positional relationship indicated by the target motion information; determining a second predicted pixel of the current image block according to the first reference image block; The predicted pixels of the current image block are determined according to the first predicted pixels and the second predicted pixels.

5. The method according to claim 1 or 2, characterized in that Determining the target motion information of the current image block according to the initial motion information includes: performing image warping processing on the mapped image block in the reference frame according to the initial motion information to obtain a warped reference frame; The target motion information is determined by performing motion search on the current image block in the warped reference frame.

6. The method according to claim 5, characterized in that The determining the predicted pixels of the current image block according to the target motion information includes: determining, in the warped reference frame according to the target motion information, a second reference image block that matches the current image block, wherein a positional relationship between the current image block and the second reference image block is a positional relationship indicated by the target motion information; The predicted pixels of the current image block are determined according to the second reference image block.

7. The method according to any one of claims 1 to 6, characterized in that The mapping relationship includes a homography matrix between the current frame and the reference frame.

8. The method according to any one of claims 1 to 7, characterized in that The acquiring of a mapping relationship between the current frame and a reference frame of the current frame includes: Obtaining camera motion parameters and camera intrinsic parameters between the current frame and the reference frame captured by the camera; The mapping relationship is determined according to the camera motion parameters and the camera intrinsic parameters.

9. The method according to any one of claims 1 to 7, characterized in that The acquiring of a mapping relationship between the current frame and a reference frame of the current frame includes: Obtaining identification information of the current frame; The mapping relationship is obtained according to the identification information of the current frame, the identification information of the current frame includes the mapping relationship or the camera motion parameters and camera internal parameters between the current frame and the reference frame taken by the camera, and the camera motion parameters and the camera memory are used to determine the mapping relationship.

10. A coding and decoding device, characterized in that: The encoding and decoding device includes an acquisition module and a prediction module; The acquisition module is used to acquire a mapping relationship between a current frame and a reference frame of the current frame; The prediction module is configured to determine initial motion information according to the current image block in the current frame and the mapping relationship, wherein the initial motion information is configured to indicate motion information between the current image block and a mapped image block having a mapping relationship with the current image block in the reference frame; The prediction module is further configured to determine target motion information of the current image block according to the initial motion information; The prediction module is further configured to determine predicted pixels of the current image block according to the target motion information.

11. A decoder, characterized in that: The decoder comprises at least one processor and a memory, wherein the memory is used to store a computer program, so that when the computer program is executed by the at least one processor, the method according to any one of claims 1 to 7 and 9 is implemented.

12. An encoder, characterized in that The encoder comprises at least one processor and a memory, wherein the memory is used to store a computer program, so that when the computer program is executed by the at least one processor, the method according to any one of claims 1 to 8 is implemented.

13. A coding and decoding system, characterized in that: The encoding and decoding system includes the encoder according to claim 12 and the decoder according to claim 11, the encoder is used to perform the operation steps of the method according to any one of claims 1-8, and the decoder is used to perform the operation steps of the method according to any one of claims 1-7 and 9.

14. A computer program product comprising instructions, characterized in that When the instructions are executed by a computer device, the computer device is caused to perform the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Visual positioning method, visual positioning device, storage medium and electronic equipment

    CN113096185A

  • Real-time video noise reduction method, device, terminal and storage medium

    CN113315884A

  • Methods and Devices For Encoding and Decoding Video Pictures

    US20170238011A1

  • Warped motion compensation with explicitly signaled extended rotations

    WO2023287417A1

Cited By

  • Coding processing method and device

    CN121691687A